In RAG evaluation, why can context precision look high while answers still miss facts?
answer
- two retriever-side halves
- one is completeness, one is noise
- reference answers, not chunk labels
- claims attributable to retrieved context
- ranking is folded into precision
basics
~20 sContext precision only asks whether the chunks you retrieved were useful and well ranked. Context recall asks whether everything the correct answer needs was retrieved at all. A tidy, high-precision context can still be missing half the required facts.
solid answer
~40 sThese are the two retriever-side halves of RAGAS-style evaluation, and they fail independently. **Context recall** is reference-based: take the ground-truth answer, break it into claims, and ask what fraction of those claims can be attributed to the retrieved context. **Context precision** is signal-to-noise plus ordering: of the chunks retrieved, how many were actually useful and were they ranked near the top — it is an average-precision-style score over the retrieved list. A contract-review assistant asked whether a termination clause survives assignment can retrieve three chunks, all genuinely about termination, all ranked well — precision near 1 — while the assignment carve-out sits in a schedule that was never retrieved. Recall exposes that; precision cannot. So report both, and treat recall as the one that bounds answer completeness. Reflects RAGAS-style definitions as of mid-2026.
go deeper
Know the split: context recall asks whether the needed information was retrieved, context precision asks how much of what was retrieved was actually useful.
Explain that context recall is measured against a reference answer's claims while context precision is a rank-weighted judgment of the retrieved chunks, and why the two fail independently.
Use the pair to localise a bad answer — low recall means fix the index or chunking, low precision means rerank or cut k, both high means the defect is in generation.
Own the evaluation design: which annotation you buy (reference answers versus chunk labels), how you keep model-computed judgments stable enough to act on, and what movement counts as signal.
## Why the retriever gets two metrics Stage-wise RAG evaluation splits the retriever's job in two, because there are two distinct ways for retrieved context to be wrong: it can be **missing something** or it can be **full of noise**. One number cannot express both, and optimising either alone produces a characteristic pathology. ## Context recall Context recall asks: *does the retrieved context contain everything needed to produce the correct answer?* The usual implementation is reference-based rather than label-based. You supply a ground-truth answer for the query, decompose it into atomic claims, and for each claim ask whether it can be attributed to any of the retrieved chunks. Context recall is the fraction of reference claims that can. A model typically performs the attribution judgment, which is what makes the metric cheap enough to run at scale. The practical consequence is that you need *reference answers*, not per-chunk relevance labels — a materially different annotation cost from classical recall@k, and often an easier one, because writing the correct answer to a question is more natural for a domain expert than judging fifty documents. Low context recall means the pipeline is information-starved. No prompt engineering, no stronger generator and no reranking can fix it, because the facts are not in the window. ## Context precision Context precision asks: *of what was retrieved, how much was useful, and was the useful part near the top?* It is computed per retrieved chunk — is this chunk relevant to answering the question, judged either against the reference answer or by a model — and then combined as a rank-weighted average, which is an average-precision-style aggregation. Two things therefore move it: the proportion of useful chunks, and their positions. Ten chunks of which two are useful score badly; the same ten score better when those two are ranked first. Low context precision means you are paying tokens for noise and giving the generator plausible material to be misled by. It rarely makes an answer impossible, which is exactly why it is the less alarming of the two. ## The asymmetry, concretely A legal contract-review assistant is asked: "does the termination-for-convenience clause survive assignment of the agreement?" Retrieval returns three chunks: the termination-for-convenience clause, a definitions chunk covering "Assignment", and a governing-law clause. Two of three are on point and ranked first and second, so context precision is high. But the carve-out stating that termination rights do not transfer on assignment lives in a schedule at the back of the contract that was never indexed as related. The reference answer contains a claim that no retrieved chunk supports, so context recall drops — and the generated answer will be confidently incomplete. The reverse failure is just as real: retrieve thirty chunks and you will probably cover every claim (high recall) while burying them among irrelevant text (low precision), producing a bloated, expensive prompt and a higher chance of the model anchoring on the wrong passage. ## Reading them together - **Low recall, any precision** — the ceiling problem. Look at chunking, coverage of the index, embedding fit, query rewriting, or the k you retrieve at. Fix this first; nothing downstream matters until it is fixed. - **High recall, low precision** — a ranking and cutoff problem. Rerank, tighten the shipped k, or filter by score. Cheap to fix and mostly a cost issue. - **Both high, answers still wrong** — retrieval has done its job and the defect is on the generation side. That is the whole point of splitting the metrics: it tells you to stop debugging the retriever. ## Caveats worth stating in an interview These metrics are model-computed, which buys scale at the cost of some noise: attribution judgments vary run to run, and a claim that is *implied* rather than stated in a chunk sits in a genuine grey zone. Pin the judging setup and treat small movements as noise rather than progress. They also depend entirely on the quality of the reference answers — an incomplete reference makes context recall look better than the system deserves, because the claims it omits are never checked. Finally, note the relationship to classical IR metrics. Context recall and recall@k measure the same intuition from different directions: recall@k needs labelled relevant chunks, context recall needs a reference answer. Teams often run classical metrics during retriever development, where chunk-level labels exist, and the reference-based pair in continuous evaluation, where writing answers is the cheaper annotation.
- What annotation do you need for context recall that you do not need for context precision?A reference answer. Context recall decomposes the ground-truth answer into claims and checks each against the retrieved context, so without a correct answer to compare against there is nothing to measure. Context precision only needs a judgment of whether each retrieved chunk is useful for the question, which a model can make from the query alone. That difference often decides which metric a team can afford first.
- Both context recall and context precision are high, but users still report wrong answers. Where do you look?At generation. Retrieval has supplied complete, well-ranked context, so the defect is downstream: the model contradicting or over-extending the retrieved material, or the prompt failing to constrain it to the context. Move to answer-stage measurement and inspect whether the claims in the answer are actually supported by the chunks that were provided. This localisation is exactly what stage-wise evaluation buys you.
- These metrics are computed by a model. What does that cost you?Reproducibility and some validity. Attribution judgments vary between runs and across judge versions, and claims implied rather than stated sit in a grey zone the judge resolves inconsistently. Pin the judging configuration, treat small deltas as noise, and spot-check a sample by hand periodically. The metrics also inherit any gaps in your reference answers — an incomplete reference silently inflates context recall.
saying these in an interview costs you the question
- Treats context precision as evidence that retrieval succeeded
- Thinks context recall needs per-chunk relevance labels
- Ignores that ranking is folded into context precision
- Believes better prompting can recover facts never retrieved
- Reads small run-to-run metric changes as real improvement