skip to content

In hybrid retrieval, how deep should each leg fetch before the lists are fused?

level: seniorimportance: should knowfreq 38%

answer

  1. fusion reorders, it cannot recover
  2. per-source k is the recall ceiling
  3. final k is a different number
  4. single-leg finds get truncated first
  5. stop where recall@k flattens

basics

~20 s

Each leg's fetch depth sets the recall ceiling: fusion can only reorder what was retrieved, never recover a document both legs missed. Fetch several times deeper than the number of results you finally keep, then trim — bounded by latency, not by the final list size.

solid answer

~60 s

Per-source top-k is the candidate budget, and it is separate from how many results you ultimately keep. If you need 10 final results, fetching 10 from each leg is the classic mistake: a document the lexical leg ranks second but the dense leg misses entirely has only one fusion contribution, so it lands well down the merged list and gets truncated away — even though it was the answer. Fetching 50 to 100 per leg gives fusion room to promote cross-leg agreement from deeper positions while still surfacing strong single-leg finds. Budgets do not have to be symmetric: if one leg is consistently the weaker performer on your traffic, give it less depth. The costs of going deeper are ANN search latency and index scan work, deduplication across the two lists, and a merged tail that becomes almost entirely single-leg documents contributing one small 1/(k+rank) term each — so beyond some depth you are adding noise and milliseconds, not recall. Measure recall@depth per leg and stop where the curve flattens.

go deeper

for a junior

Know that each retriever returns its own candidate list before merging, and that how many you keep at the end is a separate number from how many each leg fetches.

for a middle

Explain that per-source depth sets the recall ceiling because fusion can only reorder what it was given, and that shallow budgets truncate away single-leg finds.

for a senior

Show you size depth from measured per-leg recall curves, allow asymmetric budgets, and can name the costs — ANN latency, deduplication, longer downstream payload — and where the curve stops paying.

for a principal

Own the depth-versus-latency budget across the platform: where the recall ceiling should sit for each class of corpus, how it is re-derived as indexes grow, and who is accountable for noticing when it silently stops being enough.

## Two different k's Hybrid pipelines carry two numbers that people routinely conflate: - **Per-source top-k** — how many candidates each retriever returns before fusion. - **Final k** — how many fused results the caller keeps. They are not the same parameter and should not be given the same value. Per-source k is a *recall* budget; final k is a *precision and payload* budget. ## Fusion cannot invent candidates The hard constraint: a merge step can only reorder documents it was given. If a relevant document sits at rank 63 in the dense list and rank 9 in the lexical list, and you fetched 10 from each, only the lexical appearance exists and the fused evidence is one term. Fetch 100 from each and the same document accumulates two contributions and rises. Whatever neither leg returned is unrecoverable at any later stage. That makes per-source k the ceiling on the whole retrieval pipeline's recall — the most consequential number in the retrieval stage, and the cheapest one to get wrong invisibly. ## The truncation trap Setting per-source k equal to final k has a specific, damaging effect under rank fusion. Fused scores are dominated by cross-list agreement: a document in both lists collects two terms, one in a single list collects one. Under RRF with k=60, a document at rank 20 in both lists (2/80 = 0.025) outranks a document at rank 1 in only one list (1/61 ≈ 0.0164). That is usually the behaviour you want — consensus is evidence — but it means single-leg finds are systematically pushed down. When the fetch window is shallow, those single-leg finds fall off the end and you lose exactly the documents hybrid retrieval was adopted to catch: the identifier match only BM25 found, the paraphrase only the embedding found. Fetch deeply and truncate late, so demotion is a ranking decision rather than a window artifact. ## Sizing it empirically The measurement is straightforward. On a labelled query set, compute per-leg recall@k as k grows: 10, 20, 50, 100, 200. Each leg's curve rises and then flattens; the useful depth is where flattening begins, per leg, per query slice. Then check fused quality at the final cut, since more candidates is not monotonically better downstream — a longer candidate list means whatever consumes the fused output sees more marginal documents. Typical outcomes: the leg that is strong on your traffic saturates early (its relevant documents are near the top anyway), and the weaker leg keeps contributing further down, which is an argument for asymmetric budgets rather than one shared number. Asymmetry is also how you express a preference without touching a fusion weight — depth is a blunt but very stable knob. ## What the merged tail looks like A useful diagnostic when you extend depth: inspect the composition of the fused list by position. Near the top, documents tend to appear in both source lists — agreement is concentrated there. Deep in the tail, nearly every entry appeared in exactly one list and carries a single small reciprocal term, so those positions are close to arbitrary ordering between near-tied scores. Once your tail is entirely singletons at essentially tied fused scores, extra depth is buying candidates that fusion cannot meaningfully rank. That is the empirical stopping signal, and it usually arrives well before latency becomes the binding constraint. ## Costs of depth - **ANN latency.** Retrieving more neighbours means exploring more of the graph or scanning more clusters; the cost grows with depth, gently at first and then not. - **Lexical scan work.** Deeper top-k means maintaining a larger heap over longer posting lists. - **Deduplication.** The two lists overlap and may key documents differently — chunk id versus document id — so the merge has to canonicalize. Deeper lists make identity bugs more visible and more expensive. - **Downstream payload.** Everything after fusion pays for a longer list. Because both legs are usually issued in parallel, the added latency is the slower leg's increase, not the sum — which is why moderate depth is cheap and why shallow per-source budgets are rarely justified on latency grounds alone. ## The defensible answer Fetch several times the final list size from each leg, set the depth from per-leg recall curves rather than from the final k, allow the two budgets to differ, and re-check the numbers whenever the corpus grows substantially — recall@50 on a 50k-chunk index is not recall@50 on a 5M-chunk index.

  • Why is fetching exactly the final list size from each leg a particularly bad default under rank fusion?
    Because fusion rewards cross-list agreement, single-leg documents are systematically pushed down the merged order. With a shallow window they fall off the end entirely — and single-leg finds are precisely what hybrid retrieval exists to catch: the identifier only BM25 found, the paraphrase only the embedding found. Fetch deep, truncate late, so demotion is a ranking outcome rather than a window artifact.
  • Would you ever give the two legs different per-source budgets?
    Yes, and it is often the cleanest knob. If one leg's recall curve saturates by rank 20 on your traffic while the other keeps contributing to 100, matching budgets wastes work on one side and starves the other. Depth is also more stable than a score weight — it does not need recalibrating when you swap embedding models — so expressing a mild source preference through depth is usually more durable.
  • How does corpus growth affect a per-source k you tuned last year?
    It erodes it. Recall@50 measured on a small index does not transfer to one ten times larger: there are more competing near-duplicates and more distractors above the relevant document in each leg. Re-measure the per-leg recall curves after any significant corpus growth or re-chunking, and treat the depth setting as a periodically re-derived number rather than a constant.

saying these in an interview costs you the question

  • Sets per-source top-k equal to the final result count
  • Believes fusion can recover a document neither leg returned
  • Assumes deeper fetch always improves final quality
  • Uses identical budgets for both legs without measuring
  • Never re-measures fetch depth as the corpus grows

context