In RAG evaluation, how does faithfulness differ from answer relevance?
answer
- Two different comparisons, not one score
- One looks at the context, one at the question
- Independent axes can move opposite ways
- Refusals score well on one, badly on the other
basics
~20 sFaithfulness asks whether every claim in the answer is supported by the retrieved context. Answer relevance asks whether the answer addresses the question that was asked. They are independent axes, so a fully grounded answer can still be off-topic.
solid answer
~50 sThey measure two different things about the same generated answer. **Faithfulness** (also called groundedness) compares the answer to the retrieved context: is each statement entailed by the passages the generator was given? It says nothing about the user's question. **Answer relevance** compares the answer to the question: does it actually respond to what was asked, without wandering or padding? It says nothing about grounding. Because they look at different sides, they move independently. An answer that copies a retrieved paragraph verbatim but never states the requested figure is highly faithful and barely relevant. An answer that confidently gives exactly the requested figure from the model's own memory, with nothing in the context to support it, is highly relevant and unfaithful. Reporting a single blended "answer quality" number hides which of the two failed, so keep them as separate axes.
go deeper
Be able to state plainly that faithfulness compares the answer to the retrieved passages while answer relevance compares it to the question, and give one example of an answer that is high on one and low on the other.
Explain how each is actually computed — claim extraction plus entailment for faithfulness, question-alignment scoring for relevance — and why refusals sit awkwardly in both.
Show you read the pair together when diagnosing a regression, and that you track abstention rate as the shared confound before blaming the model or the prompt.
Own the reporting policy: which axes the org gates releases on, why they are never collapsed into one number, and how the thresholds trade usefulness against grounding for this product's risk profile.
## Why the generation half of RAG needs two metrics A retrieval-augmented generation system does two jobs in sequence: fetch passages, then write an answer using them. Evaluating the second job with one number is tempting and almost always misleading, because there are two independent ways for the writing step to fail. The answer can say things the passages do not support, and the answer can fail to respond to the user. Those are different defects with different fixes, so the generation half of RAG evaluation is conventionally split into faithfulness and answer relevance. ## Faithfulness (groundedness) Faithfulness asks: *is every claim in the answer supported by the retrieved context that was passed to the generator?* The comparison is answer-against-context. Note the two things it does not compare against. It does not compare against a human-written reference answer — that is answer correctness, a separate metric. It does not compare against the world — an answer grounded in a stale or wrong passage is still faithful. The usual measurement recipe decomposes the answer into atomic factual claims, checks each one for entailment against the context, and reports the supported fraction. Some pipelines skip decomposition and ask a judge for a single holistic verdict, which is cheaper but far less diagnostic: it tells you the answer is ungrounded without telling you which sentence broke. A subtlety worth knowing at any level: an answer with no factual claims — "the provided documents do not cover this" — has nothing to be unfaithful about. Frameworks either score it 1.0 or leave it undefined. Either way, a system that refuses more often will drift upward on faithfulness without getting better, which is exactly why the second axis has to be reported next to it. ## Answer relevance Answer relevance asks: *does the answer address the question that was asked?* The comparison is answer-against-question, and the context does not enter into it. Low relevance shows up as answers that restate the question, answer a neighbouring question, bury the requested value in unrequested background, or hedge so heavily that no response survives. One common automated approach reverse-engineers the answer: ask a model to generate the questions this answer would be a good response to, embed them, and measure similarity to the original question. High similarity means the answer is on target; low similarity means it drifted. Others score relevance directly with a rubric. Both are approximations, and both are blind to grounding. ## Why they are genuinely independent Picture the four quadrants: - **Faithful and relevant.** The intended outcome: the answer responds to the question using only what the passages support. - **Faithful, not relevant.** The generator plays it safe, paraphrasing a retrieved paragraph that is near the topic but never states the requested figure. Grounding is perfect; the user leaves empty-handed. - **Relevant, not faithful.** The generator gives a crisp, direct, confident answer drawn from pretraining rather than from the passages. This is the dangerous quadrant, because the answer *reads* best. - **Neither.** Usually a retrieval failure upstream: nothing useful arrived, and the model improvised something off-topic. Because the second and third quadrants exist, the two scores can move in opposite directions when you change the system. Tightening the prompt to "answer only from the provided passages" typically raises faithfulness and lowers relevance, because the model starts refusing or hedging on questions the corpus only partly covers. Loosening it does the reverse. If you track only a blended average you will see "no change" while both behaviours shift underneath. ## How this shapes an evaluation dashboard Report the two axes side by side, per query and aggregated, and never collapse them into one figure. Add the refusal or abstention rate next to them, because it is the shared confound: refusals inflate faithfulness and depress relevance simultaneously, and a movement in both metrics with a flat refusal rate means something different than the same movement with a doubled refusal rate. When you regress, read the pair together. Faithfulness down, relevance flat usually means a prompt or model change that loosened grounding discipline. Relevance down, faithfulness up usually means the opposite — a stricter instruction, a smaller context window, or a model that now abstains sooner. Both down usually points at the retrieval stage rather than the generator, which is why the generation metrics are always interpreted alongside a signal for whether the needed evidence was in the context at all. Finally, remember what neither axis covers: truth. An answer can be perfectly grounded in the retrieved passages and perfectly on-topic while being wrong, because the passages themselves were wrong or out of date. That is a third, separate axis and it needs a reference answer to measure.
- If you tighten the prompt to "answer only from the provided passages", what do you expect to happen to each score?Faithfulness typically rises and answer relevance typically falls. The model stops filling gaps from pretraining, which removes ungrounded claims, but it also starts abstaining or hedging on questions the corpus covers only partly, so fewer answers actually respond to the user. Watch the refusal rate alongside both numbers — if it jumped, the faithfulness gain is largely bookkeeping rather than a real improvement in grounding on answered questions.
- Why is a single blended "answer quality" score a bad idea for a RAG dashboard?Because the two underlying axes can move in opposite directions and cancel out. A change that makes the system refuse more looks flat in the blend while grounding rises and usefulness collapses. The blend also destroys diagnosis: you cannot tell whether to fix the prompt, the model, or the retriever from one number. Keep the axes separate and report abstention rate beside them.
- How would you score an answer that says "the retrieved documents do not contain this information"?It makes no factual claims, so faithfulness is either 1.0 or undefined depending on the framework — treat it as not informative rather than as a win. Answer relevance should be low, since the user's question went unanswered. The right handling is to bucket abstentions separately and judge them by whether the context genuinely lacked the evidence, which is a retrieval question, not a generation one.
A witness who only repeats what is in the case file is faithful; if they answer a question nobody asked, they are still useless. Grounded and responsive are separate virtues.
saying these in an interview costs you the question
- Treating faithfulness and answer relevance as the same measurement
- Assuming a grounded answer is automatically a useful one
- Believing faithfulness compares the answer to a gold reference
- Averaging both into one quality score and losing the diagnosis
- Ignoring that refusals inflate faithfulness for free