skip to content

A media-monitoring RAG's top 10 are copies of one wire story — how do you fix it?

level: seniorimportance: must knowfreq 52%

answer

  1. scoring each candidate independently is set-blind
  2. the fix belongs in selection, not scoring
  3. penalise similarity to what you already picked
  4. lambda trades relevance against coverage
  5. exact copies first, same-point-different-words second

basics

~20 s

Relevance ranking rewards redundancy, so syndicated copies all score alike. Collapse near-duplicates first with a similarity or hash check, then select the final set with a diversity-aware rule such as maximal marginal relevance, which penalises a candidate for resembling what you already picked.

solid answer

~50 s

Ranking by relevance alone has no notion of what you have already selected, so ten copies of the same wire story each score highly and all ten win. Fix it in the selection stage, in two steps. First, **collapse near-duplicates**: hash or shingle the text, or drop any candidate whose similarity to an already-kept candidate exceeds a high threshold — syndications differ only in the byline. Second, apply a **diversity-aware selection rule**, most commonly maximal marginal relevance (MMR), which at each step picks the candidate maximising `lambda * relevance - (1 - lambda) * max similarity to already-selected`. Lambda near 1 is pure relevance; around 0.5 buys real coverage; too low starts injecting off-topic material. Per-source or per-document caps are a cheap deterministic alternative. Judge the result on whether answers cover distinct facts, not on ranking scores.

code

python · 23 lines
python
def mmr(query_vec, doc_vecs, lam=0.5, m=3):
    def dot(a, b):
        return sum(x * y for x, y in zip(a, b))

    selected = []
    candidates = list(range(len(doc_vecs)))
    while candidates and len(selected) < m:
        best, best_score = None, None
        for i in candidates:
            rel = dot(query_vec, doc_vecs[i])
            red = max((dot(doc_vecs[i], doc_vecs[j]) for j in selected), default=0.0)
            score = lam * rel - (1 - lam) * red
            if best_score is None or score > best_score:
                best, best_score = i, score
        selected.append(best)
        candidates.remove(best)
    return selected


# vectors assumed unit-normalised; docs 0 and 1 are near-duplicates
docs = [[1.0, 0.0], [0.99, 0.14], [0.60, 0.80]]
print(mmr([1.0, 0.0], docs, lam=1.0, m=2))  # [0, 1] - pure relevance
print(mmr([1.0, 0.0], docs, lam=0.5, m=2))  # [0, 2] - diversity wins

go deeper

for a junior

Know that ranking scores each passage on its own, so identical passages all rank highly and can fill the final set. Be able to say that duplicates should be removed before anything reaches the prompt.

for a middle

Explain maximal marginal relevance concretely — the greedy loop, the relevance-minus-redundancy trade, what lambda does at 1.0, 0.5 and 0.2 — and distinguish exact-duplicate collapse from semantic redundancy.

for a senior

Show operational judgment: pick between MMR, per-source caps and index-time deduplication for a given corpus, handle chunk-overlap duplicates by merging spans, and say how you monitor duplicate rate in the injected set after a chunker or corpus change.

for a principal

Own coverage as a product property, not a ranking tweak: decide whether the system's job is the single best answer or a representative survey, set the evaluation that measures distinct-fact coverage, and accept the ranking-metric regression that a diversity policy will cause.

## Why good ranking produces a bad set A reranker scores each candidate independently against the query. That is exactly what you asked for, and exactly why the top of the list can be useless: if a corpus contains twelve near-identical documents, the scoring function has no mechanism to notice that it already picked one. In a media-monitoring feed this is the default state, because a single wire story is republished verbatim by dozens of outlets. The same pathology appears anywhere content is duplicated: mirrored documentation, quoted email threads, versioned policy PDFs, and — very commonly — overlapping chunks of one source document produced by a sliding-window chunker. The consequence is a coverage failure. If the user asks "what has been reported about this outage?", ten copies of one report answer a narrower question than two copies plus three independent accounts. Worse, redundancy is quietly persuasive: a model shown the same claim five times treats it as more established than a claim it sees once, so duplication can bias the answer as well as starve it. ## Two distinct problems, two distinct fixes It helps to separate them explicitly in an interview. **Duplicates** are candidates that are substantially the same text. The fix is deterministic collapse, and it belongs early — right after reranking, before selection. Techniques, cheapest first: exact hashing of normalised text; near-duplicate detection with shingling and MinHash or SimHash; or a similarity check against already-kept candidates with a high cutoff (embedding cosine around 0.95, tuned on your corpus). When you collapse, keep the highest-scored representative and, if provenance matters, record how many copies it stood for — "reported by 12 outlets" is itself useful signal. Chunk-overlap duplicates deserve special handling: adjacent chunks from the same document are often better merged into one longer passage than deduplicated away. **Redundancy** is subtler: candidates that are different texts making the same point. Hashing will not catch it. This is what diversity-aware selection is for. ## Maximal marginal relevance MMR is the standard, and interviewers expect you to be able to state it. It builds the final set greedily. Starting from the highest-scored candidate, at each step it picks the candidate that maximises `lambda * rel(query, d) - (1 - lambda) * max_{s in selected} sim(d, s)` where `rel` is the reranker score and `sim` is similarity between candidates (usually cosine over the same embeddings retrieval used). The second term is a redundancy penalty against the *already-selected* set, which is the whole trick: selection becomes order-dependent and set-aware rather than a pure top-m cut. **Lambda** is the dial. At 1.0 MMR degenerates to plain relevance ranking. Around 0.7 it lightly breaks up clusters. Around 0.5 it visibly buys coverage and is a common starting point. Below roughly 0.3 the penalty dominates and you start promoting weakly relevant material purely because it is different, which is a real and embarrassing failure — the answer cites something off-topic. Lambda is corpus- and task-dependent: a broad "what is being said about X" query wants more diversity than a precise factual lookup, so some systems set lambda per query class rather than globally. Normalisation matters. The relevance and similarity terms must be on comparable scales or lambda means nothing; if reranker scores are unbounded logits, squash or rank-normalise them before mixing. ## Cheaper and complementary mechanisms MMR is not the only tool, and often not the first one you should reach for. - **Per-source or per-document caps.** "At most two chunks from any one document, at most one item per outlet." Deterministic, trivially explainable to stakeholders, and in a syndication scenario it solves most of the problem on its own. - **Clustering the candidates** and taking the top item from each cluster. Useful when you can afford the extra pass and want explicit topical coverage. - **Metadata quotas.** Require the final set to span a time range, or to include at least one item from a designated authoritative source. - **Upstream deduplication at index time.** If the corpus itself is full of syndicated copies, collapsing them during ingestion is cheaper and more permanent than fixing it per query — at the cost of losing the "how widely was this carried" signal unless you store it. ## Knowing whether it worked Ranking metrics will not show the improvement and may show a regression, because you deliberately moved a high-scoring candidate down. Evaluate on the outcome: on a set of broad queries, how many *distinct* facts or sources does the final answer cover, and does a human rater prefer the answer? Track the duplicate rate in the injected set as an operational metric — the share of final passages exceeding the similarity threshold against another selected passage. A rise in that number after a chunker change or a corpus import is an early warning that selection has silently collapsed onto one document again. ## What a strong answer sounds like Name the cause (per-candidate scoring is set-blind), separate duplicates from redundancy, give the MMR formula and the meaning of lambda with a concrete value, mention the cheap deterministic alternative, and finish on how you would measure coverage rather than ranking.

  • What happens if you set lambda too low in MMR?
    Diversity dominates relevance and the selector starts promoting candidates chiefly because they are unlike the ones already chosen. The final set drifts off-topic, the model cites weakly related passages, and answer quality falls even though the set looks pleasingly varied. Below roughly 0.3 this is common. Tune lambda on end-to-end answer quality across both narrow factual and broad survey queries, not on how diverse the set looks.
  • Your chunker uses a sliding window with 20% overlap and the final set contains three overlapping chunks of one page. Is that a diversity problem?
    It is a duplication problem, and MMR is the wrong first tool. Overlapping chunks from one document are better handled structurally: detect adjacency by document id and offset, then merge the span into a single passage rather than dropping two of them, so no boundary-straddling sentence is lost. Apply diversity selection afterwards, over documents rather than raw chunks.
  • Why might you keep a count of collapsed duplicates rather than discarding them silently?
    The multiplicity is signal. In a media feed, how many outlets carried a story indicates reach; in a support corpus, how many tickets repeat a complaint indicates prevalence. Storing the count lets the answer say "reported by twelve outlets" from one injected passage, which is both cheaper in tokens and more informative than injecting twelve copies.

saying these in an interview costs you the question

  • Assuming a better reranker will fix redundancy on its own
  • Treating exact duplicates and same-fact redundancy as one problem
  • Setting lambda near zero to force maximum variety
  • Mixing raw reranker scores and cosine similarity without normalising
  • Judging diversity changes by ranking metrics rather than answer coverage

context