skip to content

How do you localize a wrong RAG answer to the retriever or the generator?

level: seniorimportance: must knowfreq 62%

answer

  1. One number for a chain hides the stage
  2. Two booleans, four cells
  3. Fix one stage, hold the other still
  4. Perfect retrieval does not exonerate retrieval's neighbours
  5. The right chunk can still be drowned

basics

~20 s

Check two things per failing query: was the needed evidence present in the retrieved context, and was the answer faithful to that context. Evidence missing points at retrieval; evidence present but the answer ungrounded points at generation.

solid answer

~60 s

End-to-end scores tell you a system is bad, never which half. Localize by measuring the stages separately on the same failing queries. First, ask whether the passage containing the needed fact reached the generator at all. If not, the defect is upstream — chunking, embeddings, query rewriting, top-k — and no prompt change will fix it. If it did, score faithfulness on the answer: if the answer states something the context does not support, the generator is the culprit. That second quadrant is the one people miss. In a run where every gold chunk was retrieved and the answer still asserted a $500 deductible that appears nowhere in the policy text, retrieval was blameless. Confirm it with an **oracle-context experiment**: feed the generator hand-picked correct passages and re-run. If the answer is still wrong, the reader is at fault; if it is now right, retrieval was the bottleneck all along. Also watch the middle case — right chunk present but buried among distractors — which behaves like a generator failure but is fixed upstream.

go deeper

for a junior

Know that a RAG answer can fail for two different reasons — the right passage never arrived, or the model ignored it — and that you check for the passage before blaming the model.

for a middle

Be able to describe the two measurements you take per failing query and how the four combinations map to upstream versus downstream fixes.

for a senior

Demonstrate the ablations: oracle context to bound the reader, fixed context to isolate prompt or model changes, and one variable at a time. Name the distractor case as the one that misdirects teams.

for a principal

Own the observability that makes this possible at all — logging retrieved passages and ids, stage signals reported side by side, stratified eval slices — and defend the cost of that instrumentation against the cost of misattributed engineering effort.

## Why end-to-end scoring alone is a dead end A RAG pipeline is a chain: query rewriting, retrieval, optional reranking, context assembly, generation. An end-to-end quality score aggregates every stage's failures into one number. When it drops, you know something broke and nothing about what. Localization — component-level scoring — is the discipline of attributing each bad answer to a stage, and it is the single most useful evaluation habit in production RAG. ## The 2x2 that does most of the work For each failing query, record two independent booleans: 1. **Was the needed evidence in the context handed to the generator?** 2. **Is the answer faithful to that context?** That gives four cells: - **Evidence absent, answer ungrounded.** The classic retrieval miss. The model had nothing to work with and improvised. Fix upstream: chunking, embedding model, query expansion, k, hybrid search, reranking. Prompt work here is wasted. - **Evidence absent, answer faithful.** The system correctly abstained or answered only what the context supported. Behaviourally correct generator, failed retrieval. This cell is a success for the reader and should not be counted against it. - **Evidence present, answer ungrounded.** A generator failure — the fact was on the table and the model contradicted or ignored it. Fix in the reader: prompt, model choice, context ordering, decoding settings, requiring quoted support. - **Evidence present, answer faithful, user still unhappy.** Neither stage hallucinated; the problem is elsewhere — the corpus is wrong or stale, the question was ambiguous, or the answer was incomplete. A run where every gold chunk was retrieved and the answer still invented a $500 deductible sits squarely in cell three. Teams routinely respond to it by tuning the retriever, because retrieval tuning is the reflex, and see no improvement — retrieval was already perfect. ## Interventions that turn correlation into attribution The 2x2 is observational. Two ablations make it causal. **Oracle context (fix retrieval, test the reader).** Replace the retrieved passages with hand-selected passages known to contain the answer, and re-run generation. This is the ceiling of your reader. If quality jumps, retrieval is your bottleneck and reader work is premature. If quality barely moves, the reader is the bottleneck and better retrieval will not save you. Keeping a small oracle-context set in CI gives you a stable reader-only signal that is immune to retriever churn. **Fixed context, swapped reader.** Hold the retrieved context constant and vary prompt or model. Any difference is attributable to the reader alone, because the input was identical. **Fixed reader, swapped retriever.** The mirror image: same prompt and model, different retrieval configuration. Differences are attributable upstream. Run these one at a time — changing both and reading a single end-to-end delta tells you nothing. ## The distractor case people misdiagnose There is an important middle state: the needed chunk *is* in the context, but so are several plausible near-misses — an older policy version, a different product line, a similar clause from another jurisdiction. The generator then produces a confidently wrong answer that is arguably grounded in *something* it was given. Naive faithfulness scoring may even pass it, because the claim is entailed by one of the passages. This looks like a generator failure and is usually fixed upstream, by reducing distractors: reranking, tighter filtering, deduplicating near-identical versions, or adding metadata (effective date, product, jurisdiction) so the passages are distinguishable. The diagnostic tell is that the wrong answer traces to a specific retrieved passage rather than to nothing at all — which is why per-claim verdicts that name their supporting passage are worth the extra bookkeeping. ## Making localization routine Three practices keep this cheap. First, **log the context**, not just the answer: an evaluation you cannot re-run against the exact passages the generator saw cannot attribute anything after the fact. Second, **report stage signals side by side** on every eval run — evidence-present rate and faithfulness — so a regression immediately shows its shape. Third, **stratify by query type**; multi-hop questions, questions whose answer spans documents, and questions the corpus simply does not cover fail in different stages, and a global average buries that. ## Reading the shapes Faithfulness flat, evidence-present rate down: retrieval regressed — an index rebuild, an embedding upgrade, a chunking change. Evidence-present rate flat, faithfulness down: the reader regressed — prompt edit, model version, longer contexts pushing evidence into positions the model attends to less. Both down: something shared moved, typically the ingestion pipeline or a context-assembly change that truncates passages. That decision tree turns a vague "quality is down" alert into a first place to look, which is the entire point of scoring components instead of only the end.

  • What exactly does an oracle-context experiment tell you, and what does it not?
    Feeding hand-picked correct passages to the generator measures the ceiling of your reader with retrieval removed as a variable. If answers become correct, retrieval is the bottleneck; if they stay wrong, the reader is. What it does not tell you is how the reader behaves on realistic context — oracle passages are clean, ordered and distractor-free, so reader quality measured this way is optimistic and must not be quoted as production quality.
  • The gold chunk is retrieved but sits among five near-duplicate passages from an older policy version, and the answer is wrong. Which stage do you fix?
    Diagnose it as a generation symptom with an upstream cause. The reader is being asked to disambiguate versions it has no basis to disambiguate, so the durable fixes are upstream: rerank, filter by effective date or product metadata, deduplicate near-identical chunks, or reduce k. A prompt instruction to prefer the most recent document helps only if the passages actually carry a date the model can see.
  • Why must you log the retrieved context, not just the query and answer?
    Because attribution compares the answer to the exact passages the generator received. Without that record, you cannot tell after the fact whether the evidence was present, cannot recompute faithfulness, and cannot reproduce the failure — retrieval is not deterministic across index rebuilds or embedding upgrades. Logging context, passage ids and their order turns every production failure into a replayable evaluation case.

saying these in an interview costs you the question

  • Tuning the retriever whenever any answer comes out wrong
  • Reading one end-to-end score and calling it a diagnosis
  • Changing prompt and retrieval settings in the same experiment
  • Counting an honest abstention on missing evidence as a generator failure
  • Never storing the context, so failures cannot be replayed

context