Why evaluate a RAG retriever at a larger k than the k you actually send to the generator?
answer
- two k values, two questions
- one finds the ceiling
- the other fits the budget
- the gap is a ranking problem
- a flat curve means look upstream
basics
~20 sMeasuring at a generous k shows the retriever's ceiling — whether the right chunk is anywhere in the candidate pool. Production k is set by the prompt budget. The gap between the two is a ranking problem, and rankers are cheap to fix.
solid answer
~50 sThe two k values answer different questions. **Measurement k** is deliberately large — you want to know whether the correct chunk is recoverable at all, because that separates a retrieval-input failure from an ordering failure. **Production k** is whatever the context budget, latency target and distraction tolerance allow, usually a handful. Take an internal engineering-docs search where recall@20 is 0.92 but recall@5 is 0.61, and the reader only ever sees 5. The retriever is finding the right document for 92% of queries; 31 points are being lost purely to ordering. That diagnosis points straight at a cross-encoder reranker over the top 20 to choose the final 5, not at a new embedding model or a re-chunk. Had recall@20 also been 0.61, reranking would have been wasted work and the fix would be upstream — chunking, embeddings or query rewriting.
go deeper
Know that every retrieval metric is defined at a cutoff k, and that the k used for evaluation need not be the k sent to the model.
Explain that recall is non-decreasing in k, so measuring at a large k estimates the retriever's ceiling while the shipped k reflects the context budget.
Read the recall curve as a diagnosis: a large gap between a big-k and small-k recall means reranking, a flat curve means fixing chunking, embeddings or query rewriting instead.
Own the evaluation protocol — which cutoffs are reported, when the ceiling is re-measured after a corpus or embedding change, and how production k is traded against cost and latency.
## Two k values, two questions Every retrieval metric is defined at a cutoff, and it is easy to assume that cutoff should match what you ship. It should not, because measurement and production are optimising different things. - **Production k** is a budget decision. How many chunks fit in the context you are willing to pay for, at the latency you promised, before distraction starts costing you accuracy? For most systems that is a single-digit number. - **Measurement k** is a diagnostic decision. You want the largest k that still means something, because recall at a large k is an estimate of the retriever's **ceiling**: the fraction of queries whose answer is recoverable by *any* downstream reordering. Reporting only the production number collapses two distinct failure modes into one and leaves you guessing which one you have. ## The diagnostic curve The practical technique is to compute recall at several cutoffs and read the shape of the curve — say recall@1, @5, @20, @50. Consider an internal engineering-docs assistant. Measured over a labelled query set, recall@20 is 0.92 while recall@5 is 0.61, and only 5 chunks ever reach the prompt. **Interpretation.** For 92% of queries the needed chunk is inside the top 20 candidates. Of that, only 61 points survive the truncation to 5. So 31 percentage points of achievable quality are being destroyed by ordering alone, while just 8 points are genuinely unretrievable. The bottleneck is the ranker, not the recall stage. **What that licenses.** Add a second-stage reranker — typically a cross-encoder that scores each query-chunk pair jointly rather than comparing precomputed vectors — over the top 20 and let it select the 5 that go into the prompt. The candidate pool is unchanged; only the ordering improves. Because the reranker runs over 20 items rather than the whole corpus, the added latency is bounded and predictable. **The opposite shape.** If recall@20 had also been 0.61, the curve is flat: the correct chunk simply is not in the candidate pool, and no reordering can help. The fixes live upstream — chunk boundaries that split an answer across two chunks, an embedding model that does not encode the domain's vocabulary, a query phrased in user language while the documents use internal jargon (a case for query rewriting or a hybrid lexical-plus-dense retriever), or documents missing from the index altogether. That single fork — is the ceiling high or low? — is the main reason to measure at a k you would never ship. ## Choosing the measurement k A few rules of thumb. - Go far enough that the recall curve visibly flattens. The point where extra k stops buying recall is your effective ceiling; measuring past it adds nothing. - Do not go so far that your relevance labels stop being trustworthy. On a hand-built set, documents deep in the ranking are usually unjudged, so a large k inflates apparent false positives and makes precision at that k meaningless. - Keep it stable across runs. Changing the measurement k between experiments makes the numbers incomparable, and it is a common accidental way to "improve" a retriever. ## Choosing the production k Production k is an empirical tradeoff, not a constant to copy from a blog post. Sweep it and look at end-to-end answer quality, not just retrieval metrics: recall keeps rising with k while answer quality typically peaks and then declines as distractors and prompt length take their toll. The peak is your k. Then account for cost — every extra chunk is billed on every request — and for latency, since a longer prompt raises time-to-first-token. A related decision is whether to fix k at all. Score-threshold or dynamic cutoffs return fewer chunks for narrow queries and more for broad ones, which is often better for cost, but they make metrics harder to compare because the denominator moves. If you use them, still report the fixed-k numbers alongside for comparability. ## Reporting discipline Three habits keep this honest: 1. Always print the k with the metric — `recall@20 = 0.92`, never "recall 0.92". 2. Report at least two cutoffs: one at the ceiling and one at the shipped k. The gap between them is your reranking headroom, and it is the single most actionable number in retrieval evaluation. 3. Re-measure the ceiling whenever the corpus, chunker or embedding model changes, since all three move it. Reranker changes move only the shipped-k number.
- Recall@50 and recall@5 are both 0.60 on your query set. What do you do?Stop tuning the ranker — the curve is flat, so extra candidates contain nothing useful and reranking has no headroom. The correct chunk is not in the pool at all for 40% of queries. Investigate upstream: chunk boundaries splitting answers, an embedding model that misses domain vocabulary, a vocabulary mismatch between user phrasing and document language (try hybrid lexical retrieval or query rewriting), or documents simply missing from the index.
- How do you pick the production k in practice?Sweep it and measure end-to-end answer quality rather than retrieval metrics, because recall keeps climbing with k while answer quality usually peaks and then falls as distractors and prompt length take over. Ship the k at the peak, adjusted down if cost or time-to-first-token pushes back. Then hold it fixed so later experiments stay comparable.
- Is there a limit to how large the measurement k should be?Yes — your labels. On a hand-built ground-truth set, documents deep in the ranking are typically unjudged, so relevant-but-unlabelled chunks are counted as false positives and precision at that depth becomes meaningless. Push k far enough that the recall curve flattens, then stop, and keep the value fixed across runs so results stay comparable.
saying these in an interview costs you the question
- Measures retrieval only at the k that ships
- Adds a reranker without checking recall at a larger k
- Reports a recall number with no k attached
- Assumes a bigger production k always improves answers
- Changes the measurement k between experiments