skip to content

Chunking Strategies

How documents get cut into retrievable pieces — fixed size, sentence, or semantic boundaries, with overlap — and why chunk size trades retrieval precision against how much useful context fits in the prompt. Interviewers ask because bad chunking caps the quality of everything downstream.

on this pageshow

questions

14

How does recursive text splitting choose cut points, unlike a fixed-size window?

level: juniorimportance: must knowfreq 66%

answer

  1. an ordered preference over cut points
  2. cheapest boundary that fits wins
  3. recurse with the next separator down
  4. empty string is the hard fallback
  5. merge small pieces back toward target size

basics

~20 s

Recursive splitting walks an ordered list of separators — paragraph break, line break, sentence end, space, then bare characters — and cuts on the highest-priority one that fits. A fixed-size window ignores the text and cuts when a counter runs out.

solid answer

~50 s

A fixed-size splitter counts to N and cuts, wherever that lands. A recursive splitter is given a priority list of separators, typically blank line, then newline, then sentence end, then space, then the empty string. It splits on the first separator, and for any piece still over the size limit it recurses with the next separator down, repeating until the piece fits or the list runs out. The empty-string entry is the guaranteed fallback: a hard character cut that caps every chunk. Good implementations then merge consecutive small pieces back up toward the target size so you do not end up with a chunk per sentence. The result is variable-length chunks that respect the most natural boundary available while still honouring the size cap — strictly better than a blind window on structured prose, and roughly the same when the text has no separators to find.

code

python · 11 lines
python
def split(text, size, seps):
    if len(text) <= size or not seps:
        return [text]
    sep, rest = seps[0], seps[1:]
    parts = text.split(sep) if sep else list(text)
    out = []
    for p in parts:
        out.extend(split(p, size, rest) if len(p) > size else [p])
    return out

print(split("aa\n\nbbb\ncccccc", 3, ["\n\n", "\n", ""]))

go deeper

for a junior

Be able to state the core idea plainly: a priority list of separators, cut on the best one that fits, fall back down the list when a piece is still too big.

for a middle

Explain the recursion and the merge-back-up pass, and know why the chain must end in a character-level cut to make the size limit a real guarantee.

for a senior

Show you inspect what actually happens on your corpus — which separator fires, how often the chain falls through to a hard cut, and what that implies for chunk shape on formats like chat exports or PDF-extracted text.

for a principal

Own the position that separator ordering is a corpus-specific configuration, not a default. Argue for per-format splitting policies and for instrumenting boundary quality rather than assuming one chain fits an entire heterogeneous document store.

## The baseline it improves on Fixed-size splitting is the simplest possible chunker: advance a counter over the text, emit a chunk whenever it hits N characters or N tokens, repeat. It is fast, deterministic and completely blind. It cuts mid-word, mid-sentence, mid-table-row, because nothing in the algorithm looks at what the text says. Recursive splitting keeps the size cap but adds one idea: a *preference order over cut points*. ## The separator hierarchy You hand the splitter an ordered list of separator strings, most-preferred first. A conventional list for prose runs: 1. A blank line — the paragraph boundary, the strongest natural break in plain text 2. A single newline — a line break 3. A sentence-ending pattern, such as a period followed by a space 4. A single space — a word boundary 5. The empty string — meaning cut anywhere The ordering encodes a claim about how much meaning a cut destroys. Cutting between paragraphs costs almost nothing; cutting between words costs a little; cutting mid-word costs a lot. The list says: pay the cheapest price that gets the piece under the limit. ## The recursion The algorithm is a straightforward tree walk. Split the text on separator 1. For each resulting piece, check its length. If it already fits under the size cap, keep it. If it does not, call the same procedure on that piece with the *remaining* separator list. A piece that is still oversized after every named separator falls through to the empty-string entry and gets a hard character cut. This is why the final fallback matters. Without a character-level entry at the end of the chain, a single 5,000-token paragraph with no internal newlines and no sentence punctuation — a minified blob, a base64 payload, a long URL — will come out as one oversized chunk that your embedding model then silently truncates. The empty-string entry is what turns "try to respect boundaries" into a hard guarantee that no chunk exceeds the limit. ## Merging back up A naive recursion produces a lot of tiny pieces: split a document on paragraphs and you get one chunk per paragraph, many of them two sentences long. Undersized chunks are their own problem — their embeddings are noisy and they carry too little context to answer anything. So practical recursive splitters add a merge pass: walk the pieces in order and greedily concatenate consecutive ones, rejoining them with the separator they were split on, until adding the next piece would exceed the size cap. Overlap, if configured, is applied in this same pass by re-including trailing pieces at the head of the next chunk. The net behaviour is: chunks are variable-length, always under the cap, usually near it, and they begin and end at the most natural boundary available. ## When every separator misses The hierarchy is only as good as the text's actual structure, and this is where the technique quietly degrades. Take a Slack channel export: messages are one per line with no blank lines between them. The chain blank-line, then newline, then sentence, then space effectively skips its first rung — splitting on the blank line returns the entire export as a single piece — and everything is decided by the second separator. That is not a failure, but you should know it happened, because the size and shape of your chunks are now governed by a separator you did not think would be doing the work. The harder case is text where *no* meaningful separator exists: a CSV dump with no line breaks, a document converted from a PDF that lost its newlines, machine-generated JSON on one line. There the chain falls all the way through to the character cut, and recursive splitting has degenerated into fixed-size splitting. The diagnostic habit worth having is to inspect a sample of produced chunks and check which separator actually fired. ## Character-counted vs token-counted variants The same recursion works whether the size cap is measured in characters or tokens; only the length function changes. Token-counted variants call a tokenizer on each candidate piece, which is more accurate against a model's real limit but noticeably slower, since the recursion measures pieces repeatedly. ## What it still breaks Recursive splitting reduces mid-sentence damage; it does not eliminate boundary loss. A cut at a paragraph boundary still separates a paragraph from the heading that gave it meaning, and a table split between rows still leaves the second half without its column headers. Length is the only thing the algorithm optimises. Whenever the interviewer pushes on "and where does that still go wrong?", these are the answers.

  • Why does the separator list usually end with an empty string?
    It is the guaranteed terminator. A piece with no paragraph breaks, no newlines and no spaces — a long base64 blob, a URL, minified JSON — would otherwise survive every named separator and come out oversized, which the embedding model then silently truncates. The empty-string entry means cut anywhere, so the size cap becomes a hard guarantee rather than a preference.
  • After splitting on paragraphs you get hundreds of two-sentence chunks. What is missing?
    The merge pass. Recursion alone produces pieces as small as the separator happens to make them; a practical splitter then greedily rejoins consecutive pieces, with their original separator, until the next one would exceed the cap. Without it you get undersized chunks whose embeddings are noisy and which carry too little context to answer anything.
  • How would you tell whether the separator hierarchy is actually doing any work on a given corpus?
    Sample the produced chunks and look at what they start and end with. If almost every chunk boundary is a mid-word character cut, the named separators never fired and you have effectively deployed fixed-size splitting. Counting how often each separator triggers during indexing is a cheap instrument worth adding.

saying these in an interview costs you the question

  • Thinks recursive splitting produces equal-sized chunks
  • Believes it splits on all separators simultaneously
  • Forgets the character-level fallback, so oversized chunks escape
  • Assumes it understands meaning rather than punctuation
  • Skips the merge step and ships one-sentence chunks

context

open as a page

Why split Markdown docs on their heading hierarchy instead of a fixed character count?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Headings mark where one topic ends and the next begins, so splitting on them yields chunks that each cover one subject. Fixed-length cuts merge unrelated sections and throw away the heading path that says what the chunk is about.

open as a page

In RAG chunking, what does chunk overlap buy you, and what does it cost?

level: middleimportance: must knowfreq 70%

basics

~20 s

Overlap repeats the tail of one chunk at the start of the next, so a fact cut by a boundary still appears whole somewhere. The cost is duplicated text: more chunks to embed and store, and near-duplicate search hits.

open as a page

How does semantic chunking decide where to split a document into chunks?

level: middleimportance: must knowfreq 62%

basics

~20 s

Semantic chunking splits text into sentences, embeds each sentence together with a small window of its neighbours, and cuts wherever similarity between consecutive windows drops sharply — treating that dip as a topic boundary rather than counting characters.

open as a page

In RAG, why embed small child chunks but return their larger parent sections?

level: middleimportance: must knowfreq 66%

basics

~20 s

Small chunks embed cleanly, so matching is precise; but a two-sentence hit often lacks the surrounding argument the model needs to answer. Small-to-big retrieval matches on the child and hands the generator the enclosing parent section instead.

open as a page

In RAG chunking, why can a 1000-character chunk overflow your token budget?

level: middleimportance: should knowfreq 50%

basics

~20 s

Characters and tokens are not proportional. English prose averages roughly four characters per token, but code, dense punctuation and Japanese or Chinese text run far closer to one, so the same character window can produce three or four times as many tokens.

open as a page

How do you chunk API reference docs so code blocks and tables stay whole?

level: middleimportance: should knowfreq 50%

basics

~20 s

Split on the document's own structure, not on length. Treat each fenced code block, each table together with its caption and header row, and each function or class in a source file as an atomic unit the splitter is forbidden to cut.

open as a page

In a RAG pipeline retrieving 8 chunks into an 8k-token context, how do you pick chunk size?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Work backwards from the reader's budget: subtract the system prompt, the question and room for the answer, then divide what remains by k. With 8k and k=8 that lands near 700 tokens per chunk — then raise the floor so the smallest self-contained unit still fits.

open as a page

Semantic chunking yields chunks from 40 to 3,000 tokens — what breaks downstream?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Variable chunk sizes break the assumptions after retrieval: oversized chunks get truncated by the embedding model and swallow the prompt budget, tiny chunks retrieve without enough context to answer, and a fixed top-k returns wildly different amounts of text per query.

open as a page

How do you choose the breakpoint threshold in semantic chunking, and what does it change?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The threshold is the one real knob: it decides which similarity dips count as boundaries. Pick it distributionally — a percentile or a number of standard deviations over the document's own distances — then sweep candidate values and score end-to-end retrieval, never trust a default.

open as a page

In RAG, why prepend an LLM-written situating summary to each chunk?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Chunks lose their document context. A paragraph reading "revenue fell 3% against the prior period" never says which company, filing or quarter. Prepending a short model-written blurb that situates the chunk in its document makes it self-describing, lifting both embedding and keyword retrieval.

open as a page

Why is fixed-size chunking still the baseline any smarter chunker must beat?

level: principalimportance: should knowfreq 33%

basics

~20 s

Because it is deterministic, linear-time, model-free and trivially reproducible, so its cost and behaviour are fully predictable across re-indexes. Any alternative adds inference cost, nondeterminism and an extra dependency, and must earn that with measured end-to-end gains.

open as a page

When does semantic chunking not pay for itself compared with cheaper splitting?

level: principalimportance: should knowfreq 38%

basics

~20 s

Semantic chunking earns its keep on long, unstructured prose with no markers to exploit. It usually does not pay on documents that already carry explicit structure, on short documents, on high-churn corpora where the extra ingest pass recurs, or when a measured sweep shows no retrieval gain.

open as a page

When does a recursive summary tree beat flat parent-child chunking in RAG?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

When a real share of questions span many sections rather than living in one. A recursive tree clusters chunks, summarizes each cluster into a higher-level node, and indexes every level, so broad questions match summaries while specific ones still match leaves.

open as a page