skip to content

Your RAG retriever returns three near-identical copies of one policy paragraph — what do you do?

level: seniorimportance: nice to knowfreq 30%

answer

  1. mirrored sources, one paragraph, three copies
  2. duplicates cost budget, not just tidiness
  3. repetition reads as corroboration
  4. normalize, hash, then near-match
  5. a differing number is a conflict

basics

~20 s

Collapse them before packing. Near-duplicates from mirrored sources spend the token budget three times on one fact and crowd out complementary evidence — unless the copies actually disagree, in which case keep them and make the conflict visible.

solid answer

~60 s

Mirrored corpora produce this constantly: the same expense-policy paragraph lives in Confluence, in an exported PDF and on the wiki, and all three land in the top ten. The cost is not just noise. Each copy pays full token price, so three copies of one fact push three genuinely different pieces of evidence out of the slot, and the repetition can read as corroboration of a claim that has exactly one source. Collapse at injection time: normalize whitespace and case, hash for exact matches, and use shingled overlap or a high pairwise similarity threshold for near-matches — k is small, so an O(k²) comparison over the retrieved set is cheap. Keep the copy with the best provenance (most recent, most authoritative source) and note the others as corroborating sources on the kept block. The important exception: if two copies differ on something material — a number, a date, an exception clause — that is a version conflict, not a duplicate, and collapsing it silently hides the one thing the model most needed to see.

go deeper

for a junior

Know that the same passage can be indexed from several places and come back several times, and that duplicated chunks take up budget that other evidence needed.

for a middle

Explain the detection steps — normalize, hash, then near-match by shingled overlap — and why this is cheap at injection time when only a handful of chunks are in play.

for a senior

Show the judgment: pick the survivor by provenance and recency, record what was collapsed, and refuse to collapse copies that disagree on a material value, because that conflict is the most important thing in the retrieved set.

for a principal

Frame it as an ingest-topology problem surfacing at query time. Argue for fixing mirrored sources at index time, treat a rising duplicate rate as corpus-health telemetry, and be explicit that repetition silently reweights evidence in ways your retrieval design never intended.

## Why duplicates arrive at all Enterprise corpora are mirrored by construction. A policy paragraph is drafted in a wiki, pasted into a Confluence page, exported to a PDF for distribution, and quoted in an onboarding deck. Ingest all four sources and the same 200 words exist four times with different surrounding context. A semantically-scored retriever does exactly what it should: all four are highly relevant to the query, so all four rank near the top. Hybrid retrieval and fusion across multiple retrievers make it worse, because each retriever independently surfaces its own copy. ## The three costs **Budget.** This is the concrete one. If the slot holds roughly ten chunks and three of them are the same paragraph, you shipped eight distinct facts instead of ten. The evidence that would have answered the second half of the user's question was displaced by a copy of the evidence answering the first half. **False corroboration.** Repetition reads as agreement. A model seeing the same claim stated three times in its context is being given an emphasis signal that reflects your ingest topology, not the state of the world. One source that happens to be mirrored three times can outweigh two independent sources that each appear once — which inverts the evidence weighting you actually wanted. **Reader load and dilution.** Long, repetitive context is worse context. The distinct material is now spread further apart across the block, which interacts badly with the fact that material in the middle of a long block is used least reliably. ## Detecting near-duplicates cheaply At injection time you are comparing a handful of chunks, so brute force is fine and exactness is not required: 1. **Normalize** — collapse whitespace, lowercase, strip boilerplate headers and footers, drop punctuation-only differences. This alone catches most PDF-vs-wiki pairs. 2. **Hash** the normalized text to catch exact matches in one pass. 3. **Near-match** the rest with character n-gram shingles and a Jaccard threshold, or MinHash if you want it sublinear. Comparing every pair among ten to twenty chunks is trivially cheap regardless of method. 4. Optionally use the embeddings you already have: very high pairwise cosine within the retrieved set is a strong duplicate hint, though it also fires on genuinely different chunks that merely say similar things — so treat it as a candidate filter, not a verdict. Calibrate the threshold against your corpus and err loose. Collapsing two chunks that were not duplicates is a real information loss; keeping one redundant chunk costs only tokens. ## Choosing the survivor When you collapse, you are choosing a representative. Prefer, roughly in order: the most recently updated version, the most authoritative source (the system of record over an exported copy over a slide deck), and the one whose surrounding context is most complete — a chunk that includes the section heading and the following exception clause is more useful than a bare paragraph. Record the collapsed siblings' identifiers alongside the survivor rather than discarding them, both so the response can list where the fact appears and so debugging shows what was suppressed. ## The exception that matters most Near-duplicates that differ in a material detail are not duplicates. Two copies of a travel policy that are 95% identical but state a $75 and a $100 per-diem limit are a version conflict, and it is the single most important thing in the retrieved set. Collapse them and you have silently picked a winner by whatever incidental rule your dedup used — rank, length, recency of the file's mtime — and the model will state that number with full confidence and no hedge. So make the dedup step diff-aware: when two candidates pass the similarity threshold but differ on tokens that look material (numbers, dates, named entities, negations), do not collapse. Keep both, keep their source and date headers, and let the prompt instruct the model to surface disagreement rather than choose. Detecting the disagreement is also a valuable corpus-health signal in its own right — it usually means a stale mirror nobody has retired. ## Where in the pipeline Cheap exact-hash collapse belongs early, at the point where results from multiple retrievers are merged, so downstream stages do not waste effort on copies. But keep a final guard at injection time: fusion, metadata filters and rerank cascades all reshuffle the candidate set, and the last place you can be sure about what actually enters the prompt is the place where the prompt is assembled. Log the collapse count per request; a rising duplicate rate is an ingest problem — a mirror that should have been deduplicated at index time, not at query time.

  • Where in the pipeline should this collapse happen — retrieval, merge, or injection?
    Cheap exact-hash collapse belongs early, where multiple retrievers' results are merged, so nothing downstream wastes effort on copies. But keep a final guard at injection: fusion, metadata filters and rerank stages all reshuffle the candidate set, and prompt assembly is the last point where you can be certain what actually enters the context.
  • How do you detect near-duplicates cheaply at injection time?
    Normalize first — collapse whitespace, lowercase, strip boilerplate — then hash for exact matches and use character n-gram shingles with a Jaccard threshold for the rest. Only a handful of chunks are in play, so comparing every pair is trivially cheap. Err toward a loose threshold: collapsing a non-duplicate loses information, while keeping one redundant chunk costs only tokens.
  • What does a rising duplicate rate in your logs actually tell you?
    That it is an ingest problem, not a query-time one. Persistent copies mean a mirrored source — an exported PDF, a stale wiki page — is being indexed alongside the system of record. Fixing it at index time removes the cost entirely, whereas collapsing at query time pays the retrieval and rerank cost on every request forever.

saying these in an interview costs you the question

  • Assumes vector search already removes duplicate passages
  • Treats repetition as harmless because it only wastes tokens
  • Collapses two copies that differ on a material number or date
  • Deduplicates on exact string equality only, missing near-duplicates
  • Discards the collapsed copies without recording that they existed

context