skip to content

RAG Evaluation

Measuring a RAG system stage by stage: did retrieval find the right context, and did the answer stay faithful to it — the RAGAS-style split into context recall and precision, faithfulness, and answer relevance. Interviewers ask it to see whether you can localize a bad answer to retrieval or to generation.

on this pageshow

explore

questions

10

In RAG evaluation, how does faithfulness differ from answer relevance?

level: juniorimportance: must knowfreq 72%

answer

  1. Two different comparisons, not one score
  2. One looks at the context, one at the question
  3. Independent axes can move opposite ways
  4. Refusals score well on one, badly on the other

basics

~20 s

Faithfulness asks whether every claim in the answer is supported by the retrieved context. Answer relevance asks whether the answer addresses the question that was asked. They are independent axes, so a fully grounded answer can still be off-topic.

solid answer

~50 s

They measure two different things about the same generated answer. **Faithfulness** (also called groundedness) compares the answer to the retrieved context: is each statement entailed by the passages the generator was given? It says nothing about the user's question. **Answer relevance** compares the answer to the question: does it actually respond to what was asked, without wandering or padding? It says nothing about grounding. Because they look at different sides, they move independently. An answer that copies a retrieved paragraph verbatim but never states the requested figure is highly faithful and barely relevant. An answer that confidently gives exactly the requested figure from the model's own memory, with nothing in the context to support it, is highly relevant and unfaithful. Reporting a single blended "answer quality" number hides which of the two failed, so keep them as separate axes.

go deeper

for a junior

Be able to state plainly that faithfulness compares the answer to the retrieved passages while answer relevance compares it to the question, and give one example of an answer that is high on one and low on the other.

for a middle

Explain how each is actually computed — claim extraction plus entailment for faithfulness, question-alignment scoring for relevance — and why refusals sit awkwardly in both.

for a senior

Show you read the pair together when diagnosing a regression, and that you track abstention rate as the shared confound before blaming the model or the prompt.

for a principal

Own the reporting policy: which axes the org gates releases on, why they are never collapsed into one number, and how the thresholds trade usefulness against grounding for this product's risk profile.

## Why the generation half of RAG needs two metrics A retrieval-augmented generation system does two jobs in sequence: fetch passages, then write an answer using them. Evaluating the second job with one number is tempting and almost always misleading, because there are two independent ways for the writing step to fail. The answer can say things the passages do not support, and the answer can fail to respond to the user. Those are different defects with different fixes, so the generation half of RAG evaluation is conventionally split into faithfulness and answer relevance. ## Faithfulness (groundedness) Faithfulness asks: *is every claim in the answer supported by the retrieved context that was passed to the generator?* The comparison is answer-against-context. Note the two things it does not compare against. It does not compare against a human-written reference answer — that is answer correctness, a separate metric. It does not compare against the world — an answer grounded in a stale or wrong passage is still faithful. The usual measurement recipe decomposes the answer into atomic factual claims, checks each one for entailment against the context, and reports the supported fraction. Some pipelines skip decomposition and ask a judge for a single holistic verdict, which is cheaper but far less diagnostic: it tells you the answer is ungrounded without telling you which sentence broke. A subtlety worth knowing at any level: an answer with no factual claims — "the provided documents do not cover this" — has nothing to be unfaithful about. Frameworks either score it 1.0 or leave it undefined. Either way, a system that refuses more often will drift upward on faithfulness without getting better, which is exactly why the second axis has to be reported next to it. ## Answer relevance Answer relevance asks: *does the answer address the question that was asked?* The comparison is answer-against-question, and the context does not enter into it. Low relevance shows up as answers that restate the question, answer a neighbouring question, bury the requested value in unrequested background, or hedge so heavily that no response survives. One common automated approach reverse-engineers the answer: ask a model to generate the questions this answer would be a good response to, embed them, and measure similarity to the original question. High similarity means the answer is on target; low similarity means it drifted. Others score relevance directly with a rubric. Both are approximations, and both are blind to grounding. ## Why they are genuinely independent Picture the four quadrants: - **Faithful and relevant.** The intended outcome: the answer responds to the question using only what the passages support. - **Faithful, not relevant.** The generator plays it safe, paraphrasing a retrieved paragraph that is near the topic but never states the requested figure. Grounding is perfect; the user leaves empty-handed. - **Relevant, not faithful.** The generator gives a crisp, direct, confident answer drawn from pretraining rather than from the passages. This is the dangerous quadrant, because the answer *reads* best. - **Neither.** Usually a retrieval failure upstream: nothing useful arrived, and the model improvised something off-topic. Because the second and third quadrants exist, the two scores can move in opposite directions when you change the system. Tightening the prompt to "answer only from the provided passages" typically raises faithfulness and lowers relevance, because the model starts refusing or hedging on questions the corpus only partly covers. Loosening it does the reverse. If you track only a blended average you will see "no change" while both behaviours shift underneath. ## How this shapes an evaluation dashboard Report the two axes side by side, per query and aggregated, and never collapse them into one figure. Add the refusal or abstention rate next to them, because it is the shared confound: refusals inflate faithfulness and depress relevance simultaneously, and a movement in both metrics with a flat refusal rate means something different than the same movement with a doubled refusal rate. When you regress, read the pair together. Faithfulness down, relevance flat usually means a prompt or model change that loosened grounding discipline. Relevance down, faithfulness up usually means the opposite — a stricter instruction, a smaller context window, or a model that now abstains sooner. Both down usually points at the retrieval stage rather than the generator, which is why the generation metrics are always interpreted alongside a signal for whether the needed evidence was in the context at all. Finally, remember what neither axis covers: truth. An answer can be perfectly grounded in the retrieved passages and perfectly on-topic while being wrong, because the passages themselves were wrong or out of date. That is a third, separate axis and it needs a reference answer to measure.

  • If you tighten the prompt to "answer only from the provided passages", what do you expect to happen to each score?
    Faithfulness typically rises and answer relevance typically falls. The model stops filling gaps from pretraining, which removes ungrounded claims, but it also starts abstaining or hedging on questions the corpus covers only partly, so fewer answers actually respond to the user. Watch the refusal rate alongside both numbers — if it jumped, the faithfulness gain is largely bookkeeping rather than a real improvement in grounding on answered questions.
  • Why is a single blended "answer quality" score a bad idea for a RAG dashboard?
    Because the two underlying axes can move in opposite directions and cancel out. A change that makes the system refuse more looks flat in the blend while grounding rises and usefulness collapses. The blend also destroys diagnosis: you cannot tell whether to fix the prompt, the model, or the retriever from one number. Keep the axes separate and report abstention rate beside them.
  • How would you score an answer that says "the retrieved documents do not contain this information"?
    It makes no factual claims, so faithfulness is either 1.0 or undefined depending on the framework — treat it as not informative rather than as a win. Answer relevance should be low, since the user's question went unanswered. The right handling is to bucket abstentions separately and judge them by whether the context genuinely lacked the evidence, which is a retrieval question, not a generation one.

A witness who only repeats what is in the case file is faithful; if they answer a question nobody asked, they are still useless. Grounded and responsive are separate virtues.

saying these in an interview costs you the question

  • Treating faithfulness and answer relevance as the same measurement
  • Assuming a grounded answer is automatically a useful one
  • Believing faithfulness compares the answer to a gold reference
  • Averaging both into one quality score and losing the diagnosis
  • Ignoring that refusals inflate faithfulness for free

context

open as a page

How does claim-level entailment scoring measure a RAG answer's faithfulness?

level: middleimportance: must knowfreq 58%

basics

~20 s

Split the answer into atomic factual claims, check each claim for entailment against the retrieved context, and report the supported fraction. Claim-level scoring localizes exactly which statement is ungrounded, which a single whole-answer verdict cannot do.

open as a page

In RAG retrieval evaluation, how do recall@k and precision@k differ, and which one caps answer quality?

level: middleimportance: must knowfreq 75%

basics

~20 s

Recall@k is the share of the relevant chunks that appear in the top k results; precision@k is the share of those k results that are relevant. Recall sets the ceiling — what retrieval misses, the generator can never cite.

open as a page

How do you localize a wrong RAG answer to the retriever or the generator?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Check two things per failing query: was the needed evidence present in the retrieved context, and was the answer faithful to that context. Evidence missing points at retrieval; evidence present but the answer ungrounded points at generation.

open as a page

When would you use nDCG instead of MRR to score a RAG retriever's ranking?

level: middleimportance: should knowfreq 58%

basics

~20 s

MRR records only where the first relevant result landed, so it fits lookup queries with one right answer. nDCG scores the whole top-k list with graded relevance, so it fits queries where several results matter by different amounts.

open as a page

Why can a RAG answer score perfect faithfulness and still be wrong?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Faithfulness measures agreement with the retrieved context, not with reality. If the retrieved passage is stale, wrong, or the wrong version, an answer that reproduces it faithfully scores 1.0 while telling the user something false.

open as a page

In RAG evaluation, why can context precision look high while answers still miss facts?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Context precision only asks whether the chunks you retrieved were useful and well ranked. Context recall asks whether everything the correct answer needs was retrieved at all. A tidy, high-precision context can still be missing half the required facts.

open as a page

Why evaluate a RAG retriever at a larger k than the k you actually send to the generator?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Measuring at a generous k shows the retriever's ceiling — whether the right chunk is anywhere in the candidate pool. Production k is set by the prompt budget. The gap between the two is a ranking problem, and rankers are cheap to fix.

open as a page

How would you detect unsupported claims across production RAG traffic on a budget?

level: principalimportance: should knowfreq 40%

basics

~20 s

Run a cascade: cheap deterministic checks on every answer, a small entailment model on the suspicious ones, and an expensive judge or human review on a sample. Match spend to blast radius rather than scoring everything the same way.

open as a page

How would you bootstrap and maintain a labelled query-document set for retrieval evaluation?

level: principalimportance: should knowfreq 42%

basics

~20 s

Bootstrap with an LLM generating one query per chunk, then hand-audit a sample and expect to discard a large share. Blend in real logged queries for realistic distribution, and re-audit whenever the corpus or chunker changes.

open as a page