In a RAG pipeline retrieving 8 chunks into an 8k-token context, how do you pick chunk size?
answer
- work backwards from the reader's budget
- k times chunk size plus prompt plus answer
- the floor is the smallest unsplittable unit
- small is precise, large is contextual
- sweep sizes against an eval set
basics
~20 sWork backwards from the reader's budget: subtract the system prompt, the question and room for the answer, then divide what remains by k. With 8k and k=8 that lands near 700 tokens per chunk — then raise the floor so the smallest self-contained unit still fits.
solid answer
~50 sChunk size is a budget allocation, not a round number. Start from the reader's context: reserve tokens for the system prompt, the user question and the answer you want it to write — call it 2k of an 8k window — leaving roughly 6k for retrieved text, so 8 chunks means about 700–750 tokens each, and overlap duplication eats into that further. Then check the floor: sample the corpus for the smallest unit that must stay whole, such as a warehouse SOP step together with its safety caveat, and refuse to go below it, because a chunk that splits an atomic unit produces a fragment that is wrong in isolation. Between floor and ceiling the tradeoff is precision against context: smaller chunks give sharper embeddings and more distinct sources but more mid-sentence cuts; larger chunks carry more context per hit but mix topics, diluting the vector. Sweep a few sizes over an eval set rather than arguing about it.
go deeper
Know that chunk size, the number of retrieved chunks and the model's context window are one shared budget, and that the prompt and the answer need room in it too.
Do the arithmetic out loud, and explain the precision-versus-context tradeoff: small chunks sharpen embeddings and add sources, large chunks carry context but dilute the vector.
Show the corpus side of the decision — measuring the smallest unit that must stay whole, and sweeping candidate sizes on an eval set scored end to end rather than on retrieval recall alone.
Own the framing that chunk size is a budget allocation across reader context, k and index cost, and that when the atomic-unit floor exceeds the ceiling the answer is an architecture change, not another tuning pass.
## The arithmetic ceiling The reader's context window sets a hard ceiling on chunk size once you have fixed k. The budget is: context window = system prompt + question + (k x chunk size) + answer headroom. With an 8,000-token window, a typical allocation reserves 500–1,000 tokens for the system prompt and any tool or format instructions, a couple of hundred for the user's question and conversation state, and 800–1,500 for the answer the model must still be able to write. That leaves roughly 6,000 tokens of retrieved material. Divide by k=8 and you get about 750 tokens per chunk. Configure overlap at 15% and the *distinct* content per chunk drops nearer 640 — the duplicated text still consumes context. This arithmetic is why "just use 1,000 tokens" is not an answer. At 1,000 tokens, k=8 alone is 8,000 tokens and the prompt does not fit. The three quantities — chunk size, k and context window — are one budget, and moving any of them moves the others. ## The floor set by the atomic unit The ceiling comes from the reader; the floor comes from the corpus. Every corpus has a smallest unit of text that is *wrong when split*. In warehouse standard operating procedures it is usually a numbered step plus the safety caveat attached to it: "Lower the mast fully" retrieved without "before reversing on a gradient" is not a partial answer, it is a dangerous one. In a policy manual it is a clause plus its exceptions. In a runbook it is a command plus its preconditions. So the sizing procedure is: sample a few hundred of these units, measure their token lengths, take a high percentile — 90th or 95th — and treat that as the minimum viable chunk size. If that floor exceeds the ceiling the context budget allows, you do not have a chunk-size problem; you have a k problem or a context-window problem, and the honest move is to reduce k, move to a larger-context reader, or accept that some queries will need a second retrieval round. ## The precision/context tradeoff between the bounds Inside the feasible band the tradeoff is real and not merely aesthetic. **Smaller chunks** produce embeddings that represent one idea, so similarity against a specific question is sharp and precision is high. You also get more distinct sources in the top-k, which helps questions that need corroboration. The costs are more chunks in the index, more mid-sentence and mid-unit cuts, and answers that arrive without surrounding context — the model sees the rule but not the scope it applies to. **Larger chunks** carry their own context, so the reader has more to work with per hit, and there are fewer boundaries to damage anything. But a single vector must now represent several topics at once, and averaging over more content pulls the embedding toward the document's general subject and away from any specific claim in it. That is dilution: the chunk that contains your answer scores lower because most of it is about something else. Fewer, longer chunks also mean the same k covers fewer independent sources. ## Where length-driven splitting hits its limit Sizing cannot fix everything. A freight tariff table split at any fixed length will eventually cut a rate row so the number lands in one chunk and its unit in another, and no chunk size makes that safe if the table is longer than the window. The diagnostic is worth knowing: retrieval finds the right document, the answer is confidently wrong, and inspecting the retrieved chunk shows a value with no header or unit near it. At that point you have found the boundary of what length-driven splitting can do, and the remedy lives in a different technique rather than in another round of size tuning. ## Measure, do not argue Chunk size is cheap to sweep and expensive to guess. Build a small eval set of real questions with known answer locations, then run the full pipeline at three or four sizes — say 400, 600, 800, 1,200 — holding everything else fixed, and score both retrieval (does the gold passage appear in the top-k?) and end-to-end answer quality. The two often disagree: a size that maximises retrieval recall can lose on answer quality because the chunks are too fragmentary for the reader. End-to-end quality is the metric that matters. ## Signals that the size is wrong - A high fraction of chunks begin or end mid-sentence — the size is fighting the corpus structure. - Retrieved chunks are topically right but never contain the specific fact — chunks are too large, and the answer is being diluted. - The reader keeps saying it lacks information while the right document was retrieved — chunks are too small, or context that lived elsewhere in the document was lost. - Many chunks are far below target — the merge behaviour of the splitter is producing fragments, not the size setting.
- The atomic unit in your corpus is 1,200 tokens but the budget allows only 700. What do you do?Stop tuning size — the constraint is elsewhere. Reduce k so fewer, larger chunks fit; move to a reader with a bigger context window; or restructure the pipeline so retrieval returns identifiers and a second step fetches the full unit. Shrinking chunks below the atomic unit trades a capacity problem for a correctness problem, which is a worse trade.
- Retrieval recall improves at 400 tokens but answer quality drops. How do you read that?Small chunks give sharper embeddings, so the gold passage ranks higher — but the reader receives fragments without the surrounding conditions and cannot assemble a complete answer. Retrieval recall is an intermediate metric; end-to-end answer quality is the one to optimise. Either raise chunk size to the point where both hold, or keep small chunks for matching and expand them before the reader sees them.
- How does raising k interact with chunk size in the same budget?They multiply. Doubling k halves the chunk size you can afford at a fixed context window. More, smaller chunks widen the net across documents and help questions needing several sources; fewer, larger chunks suit questions answered by one continuous passage. Decide which shape your query distribution has, then split the budget accordingly rather than tuning either number in isolation.
saying these in an interview costs you the question
- Picks a round number like 512 with no reference to the reader budget
- Forgets to reserve context for the answer itself
- Treats chunk size and k as independent settings
- Assumes bigger chunks always improve answer quality
- Tunes on retrieval recall alone and ignores end-to-end quality