skip to content

How does LangChain's RecursiveCharacterTextSplitter decide where to cut a document?

level: middleimportance: must knowfreq 72%

answer

  1. Ordered separators, tried in turn
  2. Recurse only into pieces still too big
  3. Then merge back up to the ceiling
  4. Default length is characters, not tokens
  5. from_tiktoken_encoder changes the unit

basics

~20 s

It tries a separator list in order — blank lines, then newlines, then spaces, then bare characters — recursing into any piece still over chunk_size, then merges neighbouring pieces up to chunk_size while repeating chunk_overlap units of the previous chunk.

solid answer

~40 s

`RecursiveCharacterTextSplitter` holds an ordered separator list (by default `["\n\n", "\n", " ", ""]`). It splits on the first separator, and for any resulting piece still larger than `chunk_size` it recurses with the next separator down, so it degrades from paragraph boundaries to line, word and finally character boundaries only where it must. The surviving pieces are then greedily merged back up to `chunk_size`, with `chunk_overlap` units of the tail of one chunk repeated at the head of the next. Crucially `chunk_size` is measured by `length_function`, which defaults to `len` — that is **characters, not tokens**. Use `from_tiktoken_encoder(...)` to size by model tokens, or `from_language(...)` for code-aware separators. If an atomic piece still exceeds `chunk_size`, the splitter emits it oversized and logs a warning rather than truncating.

code

python · 9 lines
python
from langchain_text_splitters import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
    chunk_size=800,
    chunk_overlap=100,
    add_start_index=True,
)
chunks = splitter.split_documents(docs)
print(chunks[0].metadata["start_index"], len(chunks[0].page_content))

go deeper

for a junior

Know that chunk_size and chunk_overlap exist, that overlap must be smaller than size, and that the default counting unit is characters.

for a middle

Walk through the separator hierarchy, the recursion into oversized pieces and the merge-back pass, and explain the token-versus-character trap concretely.

for a senior

Reason about the cost side: overlap multiplying index size, oversized-chunk warnings hiding in logs, and reproducible boundaries across re-ingestion.

for a principal

Own chunking as a corpus-wide policy — structure-aware first pass, size ceiling second, offsets recorded — and require it to be re-runnable without invalidating citations.

## The algorithm `RecursiveCharacterTextSplitter` is the default chunker for prose in LangChain, and its whole design is one idea: **prefer semantically meaningful boundaries, fall back only as far as you have to.** It carries an ordered list of separators, by default `["\n\n", "\n", " ", ""]`. Given a text it takes the first separator that appears, splits on it, and inspects each piece. Any piece already within `chunk_size` is kept as an atom. Any piece still too large is re-split using the *next* separator in the list, recursively. The empty-string separator at the end is the escape hatch: it splits between individual characters, which guarantees termination for pathological input such as a single 100k-character line with no whitespace. Once recursion has produced pieces that are individually small enough, a merge pass greedily concatenates neighbouring pieces (rejoined with their separator) until adding the next one would exceed `chunk_size`. That is what stops you from getting a chunk per sentence: the splitter cuts finely and then packs. ## chunk_overlap `chunk_overlap` makes the merge pass carry the tail of the previous chunk into the head of the next. The purpose is to stop a fact that straddles a boundary from being unretrievable in both chunks: a sentence whose subject is at the end of chunk *n* and whose predicate opens chunk *n+1* embeds badly in both. Overlap is measured in the same units as `chunk_size` and must be smaller than it — a common configuration error is copying a snippet with `chunk_size=200, chunk_overlap=400`, which the splitter rejects. Overlap is not free: it multiplies your chunk count, your embedding bill and your index size roughly by `1 / (1 - overlap/size)`. ## The units trap The single most-missed detail: `chunk_size` is counted by `length_function`, which defaults to Python's `len`. That means **characters**. Engineers routinely write `chunk_size=1000` believing they have configured 1000 tokens, then wonder why chunks are roughly a quarter the size they expected and why the context assembled from `k` of them is far under the budget they planned. Two constructors fix this. `RecursiveCharacterTextSplitter.from_tiktoken_encoder(encoding_name=..., chunk_size=..., chunk_overlap=...)` measures length with a tokenizer, so the numbers line up with the model's context accounting. `from_huggingface_tokenizer(...)` does the same for a Hugging Face tokenizer. You can also pass any callable as `length_function` yourself. ## Structure-aware variants `from_language(language=Language.PYTHON)` swaps in a separator list appropriate to a programming language — for Python, class and function boundaries before blank lines — so code chunks stop mid-nothing rather than mid-function. For Markdown, a header-aware splitter that promotes headings into metadata is usually a better first pass, with the recursive splitter run afterwards on any oversized section. The general pattern is: split on the document's real structure first, use the recursive splitter to enforce the size ceiling second. ## add_start_index Passing `add_start_index=True` records each chunk's character offset in the original document under a `start_index` metadata key. That is what lets you highlight the retrieved passage inside the source document later, and what lets you reconstruct neighbouring context at answer time. It costs nothing; turn it on. ## Oversized chunks and the warning If a single atom is still larger than `chunk_size` after every separator has been tried — which happens when you remove the empty-string separator, or with a language whose text has no spaces — the merge step emits the chunk anyway and logs a warning that a chunk was created longer than the specified size. It is a warning, not an exception, so it is easy to miss in production logs while oversized chunks silently blow your context budget. ## What this splitter does not do It has no notion of meaning. It cannot tell that a table's header belongs with its rows, or that a heading belongs with the paragraph beneath it, beyond what the separator hierarchy happens to preserve. Semantic or structure-aware splitting exists precisely because character-recursive splitting is a heuristic. It is, however, the right default: fast, deterministic, dependency-free and predictable, which matters when you re-ingest a corpus and want the chunk boundaries to be reproducible.

  • Why does the default separator list end with an empty string?
    It is the termination guarantee. If a piece is still over chunk_size after splitting on blank lines, newlines and spaces — a minified JSON blob, or a language written without spaces — the empty separator splits between individual characters so the recursion always bottoms out. Remove it and the splitter can only emit that piece oversized, with a warning.
  • How do you make chunk_size mean tokens rather than characters?
    Build the splitter through `RecursiveCharacterTextSplitter.from_tiktoken_encoder(...)` or `from_huggingface_tokenizer(...)`, which install a tokenizer-based `length_function`; you can also pass any counting callable as `length_function` directly. Then `chunk_size=512` really is 512 tokens, and your context-budget arithmetic against the model's window becomes honest.
  • What does raising chunk_overlap from 0 to 200 cost on a 1000-character chunk size?
    Roughly 25 percent more chunks, and with them 25 percent more embedding calls, index storage and candidate noise at query time, since near-duplicate chunks now compete for the same top-k slots. Overlap buys insurance against facts straddling a boundary; it is worth paying when your text has long cross-referencing sentences, not as a reflex.

saying these in an interview costs you the question

  • Thinks chunk_size defaults to tokens
  • Sets chunk_overlap larger than chunk_size
  • Believes the splitter understands sentence meaning
  • Assumes every chunk is exactly chunk_size long
  • Says the loader already chunked the text

context