skip to content

How do you set chunk size and top-k when retrieved passages share the context window?

level: seniorimportance: should knowfreq 58%

answer

  1. two numbers multiply into one bill
  2. recall curve and accuracy curve differ
  3. the plateau, not the maximum
  4. duplicates and contradictions compete
  5. the window is shared, not yours

basics

~20 s

Chunk size times top-k is the token bill. Pick the smallest k at which end-to-end answer accuracy stops improving — not the k that maximises retrieval recall — and deduplicate near-identical passages before they crowd out the one that matters.

solid answer

~50 s

Treat these as budget decisions, not pure information-retrieval decisions. The retrieved payload is roughly chunk size multiplied by k, and it competes for the window with tool schemas, conversation history, instructions and the output itself. Two curves matter and they are not the same curve: retrieval recall rises monotonically with k, while end-to-end answer accuracy rises, plateaus, and then falls as extra passages bring near-duplicates, superseded versions and contradictory statements the model must adjudicate. So tune on answer accuracy and stop at the plateau. Big chunks carry more surrounding context per hit but spend the budget fast and dilute the vector; small chunks are more precise and need more of them. Shrink the candidate pool with metadata filters before you spend tokens on it, deduplicate near-identical passages, and do not assume a million-token window rescues you — as of mid-2026, benchmarks put reliably usable context well below the advertised figure, with cliff-shaped rather than gradual degradation.

go deeper

for a junior

Know that retrieved passages consume the same context window as the prompt and history, and that chunk size multiplied by top-k is roughly what you spend per query.

for a middle

Explain why raising k helps recall but can hurt answers — duplicates, contradictions and plausible-but-irrelevant passages — and describe the tradeoff between large and small chunks.

for a senior

Show that you tune on end-to-end accuracy with a labelled set, take the smallest k on the plateau, and use metadata filters and deduplication to buy back budget. Be ready to diagnose a retrieved-but-wrong answer.

for a principal

Own the budget across the whole system: how the retrieval allowance is divided against tool schemas, history and output headroom, what token spend per query the product can afford, and how the policy is re-validated when corpus or model changes.

## The two dials are one dial Chunk size and top-k are almost always tuned separately, by different people, against different metrics. In a context-engineering framing they are one decision, because what reaches the model is roughly `chunk_tokens x k`. Move either dial and you move the same number. A pipeline running 200-token chunks at k=20 and one running 800-token chunks at k=5 spend the same budget and behave very differently. ## Why the window is smaller than it looks The retrieved payload is not alone. Tool and function schemas can occupy a large fixed block before the user types anything; conversation history and prior tool results accumulate; the system instructions are there; and the model needs headroom to generate. Whatever is left is the retrieval allowance, and it shrinks as the session runs. And the advertised window is not the usable one. As of mid-2026, long-context evaluation consistently reports effective context materially below the advertised number, with degradation that arrives as a cliff rather than a gentle slope — quality holds, then falls off. Planning to fill a very large window because it exists is the mistake this question is really probing. ## The two curves **Retrieval recall@k** measures whether the passage containing the answer is anywhere in the returned set. It only rises as k rises; you cannot lose a hit by adding more results. **End-to-end answer accuracy** measures whether the final response is correct. It rises with k while additional passages are adding signal, plateaus, and then declines. That decline is the interesting part and it has concrete causes: - **Near-duplicates.** Three revisions of one policy paragraph occupy budget and force the model to pick a version, often silently choosing the wrong one. - **Contradiction.** Two retrieved passages state incompatible facts — an old limit and a new one. The window now contains a clash the model must resolve with no authority signal. - **Plausible irrelevance.** Passages that are topically close but do not answer the question are the hardest distractors, because they look like evidence. - **Crowding.** Every marginal passage displaces budget that could have held history, tool results, or a longer answer. The practical rule: tune k on answer accuracy, not on recall. It is common for accuracy to peak well below the k that recall says you need, and shipping the recall-optimal k is a recognisable mistake. ## Choosing chunk size Larger chunks carry more surrounding context per hit, so the model can read the passage in situ; they also blur the embedding, since one vector must summarise more material, and they spend the budget quickly. Smaller chunks embed more sharply and retrieve more precisely, but arrive stripped of their surroundings, which pushes you to retrieve more of them or to expand them at read time. Structure is usually a better guide than a token count. A chunk that ends mid-procedure is a defect regardless of its length. Where documents have real boundaries — sections, clauses, function definitions, log entries — use them, and let size vary. Where they do not, a fixed window with overlap is the fallback, and the overlap exists specifically to stop boundary-straddling facts from becoming unfindable. ## Spend fewer tokens by narrowing first Metadata filters are a budget instrument as much as a relevance instrument. Restricting to the current product line, the current document version, or the tenant that owns the request shrinks the candidate pool before ranking, which means a smaller k can achieve the same recall — and it removes an entire class of clash, because superseded versions never enter the window at all. The same logic applies to deduplication: collapse near-identical passages after ranking and you often recover several slots' worth of budget with no loss of information. ## Why good context still yields wrong answers This is the version of the question interviewers actually enjoy. The correct passage was retrieved, and the answer was still wrong. Usual causes, in rough order of frequency: a contradicting passage was retrieved alongside it and won; a near-duplicate from an older revision was preferred; the chunk was truncated so the qualifying clause never arrived; the passage was correct but the question required combining two passages and only one was retrieved; or so much marginal material was included that the model latched onto the wrong part of it. Notice that four of the five are *too much retrieval*, not too little — which is why the instinct to raise k when quality is poor so often makes things worse. ## How to actually tune it Build a small labelled set of real questions with known answer passages. Sweep k at a fixed chunk size and plot both curves. Find the accuracy plateau and take the smallest k on it. Then sweep chunk size at that k. Record the resulting token spend per query as a first-class number alongside accuracy, because the next person to raise k will need to see what it costs. Re-run the sweep when the corpus or the model changes; neither curve is stable across either.

  • Retrieval recall@20 is 95% but answer accuracy is best at k=5. What do you ship, and why?
    Ship k=5. Recall measures whether the answer passage is present; accuracy measures whether the system gets the answer right, and that is the objective. The gap tells you the extra fifteen passages are net-harmful — duplicates, superseded versions and plausible irrelevance that the model must adjudicate. Keep the recall number as a diagnostic: if it were low, you would have a retrieval problem rather than a budgeting one.
  • Your window is a million tokens. Does the budgeting argument still hold?
    Yes, for two reasons. Reliably usable context is well below the advertised size, and mid-2026 evidence shows degradation arriving as a cliff rather than a slope, so a full window is not a safe window. And cost and latency scale with what you send. A large window buys headroom for history and tool results; it is not a licence to raise k until recall saturates.
  • How do metadata filters change the chunk-size and k calculation?
    They shrink the candidate pool before ranking, so a smaller k reaches the same recall — you spend fewer tokens for the same coverage. They also remove whole classes of contradiction by keeping superseded versions, other tenants or out-of-scope products out of the window entirely. The risk is an over-narrow filter that silently excludes the answer, so filters need the same evaluation discipline as k.

saying these in an interview costs you the question

  • Tuning top-k on retrieval recall instead of answer accuracy
  • Assuming more retrieved passages can only help
  • Treating a large advertised window as fully usable
  • Ignoring tool schemas and history when budgeting the window
  • Leaving near-duplicate passages in the retrieved set

context