After a RAG multi-query fan-out returns four ranked lists, what breaks if you just concatenate them?
answer
- paraphrases overlap on purpose
- forty slots, far fewer distinct chunks
- repetition reads as corroboration
- same id and same text are different tests
- truncate to budget, not to paraphrase count
basics
~20 sConcatenation repeats chunks that several paraphrases found, so the context fills with duplicates instead of distinct evidence, and the effective top-k collapses. Merge on a document or chunk identifier first, keep each chunk once, then cut to the context budget.
solid answer
~50 sFour paraphrases of the same question overlap heavily by design — that is the signal that a chunk is genuinely relevant — so the concatenated 40 slots may contain only 12 distinct chunks. Passing that to the generator wastes most of the context budget on repetition and, worse, repeated passages read as corroboration to the model. So the merge step has to (1) key on a stable chunk or document id and keep one copy, (2) order the survivors in a way that respects agreement across lists rather than raw per-query similarity, and (3) truncate to a fixed budget. Reciprocal rank fusion is the usual choice for step two because it needs only positions, not scores from different query vectors. Dedup belongs here, at the merge, not in the prompt and not after generation — and duplicate *content* under different ids (near-identical chunks from overlapping windows) survives id-based dedup, so near-duplicate collapsing may be needed too.
code
python · 16 linesdef merge_fanout(lists, top_n=6):
best_rank = {}
for ranked in lists:
for rank, chunk_id in enumerate(ranked, start=1):
if chunk_id not in best_rank or rank < best_rank[chunk_id]:
best_rank[chunk_id] = rank
return sorted(best_rank, key=best_rank.get)[:top_n]
fanout = [
["c1", "c2", "c3"],
["c2", "c4", "c1"],
["c5", "c2", "c3"],
["c1", "c5", "c6"],
]
print(len([c for lst in fanout for c in lst]), "slots")
print(merge_fanout(fanout), "distinct, budgeted")go deeper
Know that fanning out several queries returns several result lists, and that the same document usually appears in more than one of them, so duplicates must be removed before building the prompt.
Explain the mechanics: dedup on a chunk identifier, order by agreement across lists rather than by per-query score, and truncate to a fixed context budget.
Show the operational depth — near-duplicate text under distinct ids, dedup before fusion so nothing double-counts, and the distinct-chunk ratio as the signal that a fan-out is redundant or has drifted.
Frame it as a budget decision: fan-out should change which evidence reaches the model, not how much. Decide where dedup lives in the pipeline, what the context budget is, and how the win is measured before shipping it broadly.
## What fan-out actually hands you Run four paraphrases of a terse bug-tracker title through the retriever and you get four ranked lists, each of, say, ten chunks. Naive pipelines concatenate them and pass the result to the generator. That is wrong in three separate ways, and each one has a different fix. ## Problem one: the lists overlap, and that is the point Paraphrases are deliberately close to each other, so their result sets intersect heavily. If the same chunk appears in all four lists, forty slots may hold only ten to fifteen distinct chunks. Two consequences follow. The budget consequence is obvious: you are paying tokens to send the same passage four times, and the distinct evidence you actually gained from fanning out — the chunks only one paraphrase found — gets pushed out by the repeats. The subtler consequence is that repetition in a prompt is not neutral. A passage stated four times reads as four sources agreeing, and models weight it accordingly. If the repeated chunk is a plausible-but-wrong match that every paraphrase happened to hit, fan-out has amplified an error rather than corrected it. The fix is to key on a stable identifier — chunk id, or document id plus offset — and keep one copy. Do it at the merge step. Doing it in the prompt ("ignore repeated passages") wastes tokens and trusts the model with bookkeeping; doing it after generation is too late, because the answer is already skewed. ## Problem two: id-based dedup is not content dedup Overlapping chunk windows, documents that exist in more than one source system, boilerplate headers and near-identical revisions of the same policy all produce chunks with different ids and nearly identical text. Id-based dedup lets every one of them through. Fan-out makes this worse than single-query retrieval, because different paraphrases tend to surface different near-duplicates of the same passage. If your corpus has this shape — and most enterprise corpora do — you need a second pass: collapse chunks whose text is near-identical, keeping the highest-ranked representative and, ideally, remembering that the others existed so citation still works. This is cheap to do with a hash of normalized text for exact repeats, and more expensive but more effective with a similarity threshold for near-repeats. ## Problem three: order Once you have distinct chunks, in what order do you keep them? Concatenation gives you the first list's order followed by the second's, which arbitrarily privileges whichever paraphrase ran first. The property you actually want is *agreement*: a chunk retrieved near the top by three paraphrases is a stronger candidate than one retrieved at rank one by a single paraphrase. Reciprocal rank fusion is the standard way to express that, and its practical virtue in this setting is that it consumes only rank positions. Every list came from a different query vector, so per-query similarity numbers are not a common currency you can average; positions are. That property is why RRF is reached for at the query-merge layer specifically, rather than any score-blending scheme. However it is ordered, the merged list must then be truncated to a fixed budget — the same number of chunks you would have passed without fan-out, or a modest multiple. Fan-out is meant to improve *which* chunks reach the generator, not to inflate how many do. Letting the context grow with the number of paraphrases is how teams accidentally convert a retrieval improvement into a context-dilution regression. ## Ordering the pipeline A workable order is: fan out, retrieve per query, dedup by id, collapse near-duplicates, fuse by rank, truncate to budget, and only then hand off to whatever downstream stage assembles or re-scores the context. Dedup before fusion matters: if the same chunk is still present under two ids when you fuse, it accumulates score twice and rises for the wrong reason. ## Diagnosing it in production The metric worth logging per request is the ratio of distinct chunks to total retrieved slots. If four paraphrases yield forty slots and eight distinct chunks, your paraphrases are too similar to be earning their cost — you are paying four retrievals and a rewrite call for what one query nearly gave you. If they yield thirty-eight distinct chunks, the opposite worry applies: the paraphrases have drifted so far apart that they are searching unrelated regions, and the merged context will be incoherent. A healthy fan-out sits in between, with meaningful overlap at the top of the lists and genuine novelty in the tail. That single ratio is the fastest way to tell whether a fan-out is doing anything at all, and it is worth putting on a dashboard before you start tuning the number of paraphrases.
- Why not simply average the similarity scores from each paraphrase's result list?Each score is computed against a different query vector, so the numbers are not a shared currency — 0.82 for one paraphrase and 0.82 for another need not mean equal relevance, and score distributions vary by query. Rank position is the one comparable quantity across lists, which is why rank-based fusion is the default at this layer. Averaging scores also lets one atypically-scaled query dominate the merged order.
- How do you decide how many chunks to keep after the merge?Keep the budget you would have used without fan-out, or a small multiple of it, and tune it on an eval set. Fan-out is meant to improve which chunks reach the generator, not how many. Growing the context linearly with paraphrase count dilutes attention and raises cost, and it usually masks whether the fan-out itself helped — you can no longer tell recall gains from context-size effects.
- What would you log to tell whether a fan-out is worth its cost?Per request, log the number of retrieved slots, the number of distinct chunks after dedup, and how many chunks in the final context were found by only one paraphrase. Near-total overlap means the paraphrases are redundant and you are paying for nothing; near-zero overlap means they have drifted apart and are searching unrelated regions. Healthy fan-out shows agreement at the top and novelty in the tail.
saying these in an interview costs you the question
- Passing the concatenated lists straight to the generator
- Assuming four lists of ten give forty distinct chunks
- Asking the model in the prompt to ignore duplicates
- Treating id-based dedup as sufficient for near-identical text
- Letting the context size grow with the number of paraphrases