skip to content

Rerank Cascades and Result Selection

Reranking is a budget decision: how many candidates you pull, how many stages you cascade, and which survivors actually reach the prompt — including whether that final set is diverse or five copies of one fact.

on this pageshow

questions

5

In a RAG rerank cascade, how do you choose k retrieved, n reranked and m read?

level: middleimportance: must knowfreq 68%

answer

  1. three widths, three different budgets
  2. the first number sets a ceiling
  3. cheap and generous, then expensive and few
  4. recall knee, latency budget, answer eval
  5. 100 to 25 to 5

basics

~20 s

Pick k by measured recall — large enough that the correct passage is somewhere in the candidate pool. Pick n by what the reranker can score inside the latency budget. Pick m by how many passages the model actually uses well.

solid answer

~40 s

Treat the three numbers as separate, measurable decisions. **k** (candidates pulled from cheap retrieval) is a recall decision: sweep k and measure how often a known-good passage appears anywhere in the pool. Recall is the ceiling — nothing a later stage does can recover a passage k missed. **n** (candidates handed to the expensive reranker) is a cost decision: reranking is roughly linear in n, so n is set by your latency and spend budget, and there is little point making n larger than the k where recall has already flattened. **m** (passages injected into the prompt) is a quality decision, usually small: past a handful, extra passages add tokens and dilution faster than they add answers. A legal-research console might run 100 → 25 → 5, with an explicit millisecond budget per stage.

go deeper

for a junior

Know that RAG usually pulls many candidates cheaply, then narrows them with a more expensive scorer before only a few reach the prompt. Be able to say why all three counts differ.

for a middle

Explain each number's governing constraint: k by candidate-pool recall, n by reranker latency and cost, m by end-to-end answer quality. Give a concrete shape such as 100 → 25 → 5 and say how you would sweep each.

for a senior

Show you have tuned this against measurements on a real corpus — recall knees, per-stage latency budgets, the regression where a bigger m improved ranking scores but degraded answers — and say what you re-tune after an embedding-model or chunking change.

for a principal

Own the cascade as a cost and reliability surface: budget milliseconds and cents per stage, decide when a third stage or per-query-class routing is worth the operational complexity, and set the org policy that pool recall is reported alongside every reranker win.

## The cascade as three separate numbers A retrieval-augmented generation (RAG) pipeline that reranks has at least three tunable widths: **k**, how many candidates the cheap first-stage retriever returns; **n**, how many of those the expensive reranker actually scores; and **m**, how many survivors are written into the prompt the model reads. They are frequently conflated ("we use top-5"), but each is governed by a different constraint, measured with a different signal, and fails in a different way. ## Choosing k — recall is the ceiling First-stage retrieval exists to be cheap and generous. Its job is not to rank well; its job is to not lose the answer. Whatever recall you have at k is a hard ceiling on the whole pipeline: a reranker can only reorder what it was given, so a passage missing from the candidate pool is unrecoverable no matter how good the later stages are. The practical method is a sweep. Take a labelled set of queries with known-relevant passages, and plot the fraction of queries whose relevant passage appears anywhere in the top k, for k = 10, 25, 50, 100, 200. The curve rises steeply then flattens. Pick k somewhere past the knee — the point where doubling k buys a percentage point or less. Typical production values sit between 50 and 200 for a corpus of any size; very small corpora may need far less. ## Choosing n — the reranker's cost curve The second stage is expensive per candidate, so its cost is roughly linear in n. That makes n the budget dial. Two constraints bound it from above: the latency you can afford before the user sees a first token, and the money per query. Two observations bound it from below: n smaller than the k where recall flattened throws away recall you already paid for, and an n so small that the reranker is only reordering near-ties gives you no measurable quality lift. When a single n cannot satisfy both quality and latency, you add a stage rather than compromise: 200 candidates → a cheap intermediate scorer → 25 → an expensive reranker → 5. Each stage should be strictly cheaper per item and strictly higher-recall than the one after it. More than three stages is rare outside web-scale search; each stage adds a place for tuning to rot. ## Choosing m — what the reader can actually use m is not "as many as fit". Long contexts degrade: attention over a large body of loosely relevant text makes the model likelier to lean on a weak passage, to blend two documents into a claim neither made, or to miss the one passage that mattered. Every extra passage also costs tokens on every request forever. In practice m is small — often three to eight chunks — and is chosen by running the end-to-end answer eval at several values and taking the point where answer quality stops improving, not by filling the window. m interacts with chunk size. Twenty small chunks and five large ones can be the same token count with very different behaviour, so tune m against the chunking you actually ship. ## A tuning procedure that works 1. Fix chunking and the retriever. Sweep k, measure candidate-pool recall, choose k past the knee. 2. Fix k. Sweep n against latency and cost per query; confirm end-to-end answer quality is flat above your chosen n. 3. Fix k and n. Sweep m against end-to-end answer quality, not ranking quality, since m is the only number the generator sees. 4. Re-run the sweep whenever chunking, the embedding model, the reranker, or the corpus composition changes. These numbers are not portable across those changes. ## Failure modes to name in an interview - **k too small.** Ceiling failure. Symptom: the reranker looks fine on the candidates it gets, but the system confidently answers from the wrong document. Diagnostic: candidate-pool recall, measured separately from final answer quality. - **n ≈ k with a large k.** You are paying full reranker price for the entire pool; latency climbs linearly while quality is flat past the knee. - **m too large.** Token cost rises, answers get vaguer, and the true passage competes with distractors. This is the most common silent regression, because ranking metrics improve while answers get worse. - **Tuned on the wrong signal.** Optimising k and n on ranking quality alone is fine; optimising m that way is not — only an end-to-end answer eval can tell you whether the reader benefited. - **Never re-tuned.** Numbers chosen for a 10k-document pilot are usually wrong at 10M documents. The interview answer that lands is the one that says these are three decisions with three different measurements, gives a concrete shape such as 100 → 25 → 5, and names recall as the ceiling everything else lives under.

  • If your candidate-pool recall at k=100 is 92%, what does that tell you about the best possible end-to-end accuracy?
    It caps it. In 8% of queries the supporting passage never enters the pipeline, so no reranker, no prompt change and no bigger model can recover those answers. The fix lives upstream — chunking, the query representation, adding a keyword or metadata path, or raising k. Reporting a reranker win without also reporting pool recall hides that ceiling.
  • When would you skip the rerank stage entirely for some queries?
    When the first-stage result is already unambiguous or the query class does not need precision. Common triggers: a very large score gap between the top hit and the rest, an exact-identifier lookup, a cached or previously answered query, or a latency-critical surface such as type-ahead. Route those straight through and spend the reranker budget on the ambiguous natural-language queries where ordering is actually wrong.
  • Why not just make m large now that context windows are huge?
    Because cost and quality both push back. Every injected passage is tokens on every request, and long, loosely relevant context measurably increases blending and distraction — the model can anchor on a plausible but wrong passage. Large windows change what is possible, not what is optimal; m should still be chosen by an end-to-end answer eval, and it usually lands in the single digits.

saying these in an interview costs you the question

  • Believing the reranker can recover a passage retrieval never returned
  • Setting k, n and m to one number because it looked round
  • Treating more injected passages as strictly better for answers
  • Tuning m on ranking metrics instead of end-to-end answer quality
  • Never re-tuning the cascade after the corpus grows

context

open as a page

A media-monitoring RAG's top 10 are copies of one wire story — how do you fix it?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Relevance ranking rewards redundancy, so syndicated copies all score alike. Collapse near-duplicates first with a similarity or hash check, then select the final set with a diversity-aware rule such as maximal marginal relevance, which penalises a candidate for resembling what you already picked.

open as a page

In LLM reranking, how does listwise ranking differ from pointwise scoring?

level: middleimportance: should knowfreq 38%

basics

~20 s

Pointwise scores each passage alone and sorts by the number. Listwise shows the model several passages at once and asks it to output an ordering, so it can compare them directly — better ordering, but order-sensitive, harder to parallelise and it emits ranks rather than comparable scores.

open as a page

When should a RAG reranker return nothing, and how do you pick that score cutoff?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Return nothing when the best candidate's absolute relevance score falls below a calibrated floor — the corpus simply has no answer. Set the floor from labelled in-scope and out-of-scope queries, choosing the point that meets your tolerated false-answer rate.

open as a page

In RAG, when relevance alone picks the wrong passages, how do you design selection policy?

level: principalimportance: should knowfreq 33%

basics

~20 s

Add an explicit policy layer after ranking that turns scores into a chosen set under constraints: recency windows, source authority, per-document caps, required coverage slots and a relevance floor. Then evaluate the policy on answer outcomes, not ranking order.

open as a page