📝 docs(text-splitters.mdx): improve formatting and add missing information about LanguageRecursiveTextSplitter and its parameters
🐛 fix(text-splitters.mdx): fix typo in the description of `separators` parameter in `RecursiveCharacterTextSplitter`
This commit is contained in:
parent
40ab6b1e87
commit
383c9dc5ff
1 changed files with 23 additions and 9 deletions
|
|
@ -1,10 +1,12 @@
|
||||||
import Admonition from '@theme/Admonition';
|
import Admonition from "@theme/Admonition";
|
||||||
|
|
||||||
# Text Splitters
|
# Text Splitters
|
||||||
|
|
||||||
<Admonition type="caution" icon="🚧" title="ZONE UNDER CONSTRUCTION">
|
<Admonition type="caution" icon="🚧" title="ZONE UNDER CONSTRUCTION">
|
||||||
<p>
|
<p>
|
||||||
We appreciate your understanding as we polish our documentation – it may contain some rough edges. Share your feedback or report issues to help us improve! 🛠️📝
|
We appreciate your understanding as we polish our documentation – it may
|
||||||
|
contain some rough edges. Share your feedback or report issues to help us
|
||||||
|
improve! 🛠️📝
|
||||||
</p>
|
</p>
|
||||||
</Admonition>
|
</Admonition>
|
||||||
|
|
||||||
|
|
@ -44,6 +46,18 @@ The `RecursiveCharacterTextSplitter` splits the text by trying to keep paragra
|
||||||
|
|
||||||
- **chunk_size:** Determines the maximum number of characters in each chunk when splitting a text. It specifies the size or length of each chunk.
|
- **chunk_size:** Determines the maximum number of characters in each chunk when splitting a text. It specifies the size or length of each chunk.
|
||||||
|
|
||||||
- **separator_type:** The parameter allows the user to split the code with multiple language support. It supports various languages such as Text, Ruby, Python, Solidity, Java, and more. Defaults to `Text`.
|
- **separators:** The `separators` in RecursiveCharacterTextSplitter are the characters used to split the text into chunks. The text splitter tries to create chunks based on splitting on the first character in the list of `separators`. If any chunks are too large, it moves on to the next character in the list and continues splitting. Defaults to ["\n\n", "\n", " ", ""].
|
||||||
|
|
||||||
- **separators:** The `separators` in RecursiveCharacterTextSplitter are the characters used to split the text into chunks. The text splitter tries to create chunks based on splitting on the first character in the list of `separators`. If any chunks are too large, it moves on to the next character in the list and continues splitting. Defaults to `.`
|
### LanguageRecursiveTextSplitter
|
||||||
|
|
||||||
|
The `LanguageRecursiveTextSplitter` is a text splitter that splits the text into smaller chunks based on the (programming) language of the text.
|
||||||
|
|
||||||
|
**Params**
|
||||||
|
|
||||||
|
- **Documents:** Input documents to split.
|
||||||
|
|
||||||
|
- **chunk_overlap:** Determines the number of characters that overlap between consecutive chunks when splitting text. It specifies how much of the previous chunk should be included in the next chunk.
|
||||||
|
|
||||||
|
- **chunk_size:** Determines the maximum number of characters in each chunk when splitting a text. It specifies the size or length of each chunk.
|
||||||
|
|
||||||
|
- **separator_type:** The parameter allows the user to split the code with multiple language support. It supports various languages such as Ruby, Python, Solidity, Java, and more. Defaults to `Python`.
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue