Why can a RAG answer score perfect faithfulness and still be wrong?
answer
- Consistent with what, exactly?
- Green dashboard, wrong answer
- A stale page paraphrased perfectly
- Two passages disagree, both are grounded
- Saying nothing is trivially faithful
basics
~20 sFaithfulness measures agreement with the retrieved context, not with reality. If the retrieved passage is stale, wrong, or the wrong version, an answer that reproduces it faithfully scores 1.0 while telling the user something false.
solid answer
~50 sFaithfulness is a *relative* metric: its reference point is the retrieved context, not the truth. That makes it blind to three real failure modes. **Bad source.** A superseded policy page, a draft that was never published, or a customer-authored forum post gets retrieved and paraphrased accurately. Perfect grounding, wrong answer. **Conflicting sources.** Two retrieved passages disagree; the model picks one. Whichever it picks is entailed by the context, so faithfulness passes regardless of which is current. **Degenerate compliance.** An answer that hedges, restates the question, or refuses makes few or no claims and is trivially faithful — a system that abstains more trends upward on this metric without getting better. So faithfulness must be reported next to answer correctness, which compares against a gold answer, and next to abstention rate. Corpus hygiene — freshness, versioning, deduplication, authority ranking — is the fix for the first two, and no generation-side metric will surface them.
go deeper
Remember the one-line reason: faithfulness checks the answer against the retrieved passages, not against reality, so a wrong passage produces a wrong answer with a perfect score.
Explain the specific cases — stale or superseded sources, conflicting passages, and claim-free answers — and name answer correctness as the metric that covers the gap.
Show you would instrument for it: correctness on a curated reference set, chunk-age distribution, contradiction flags on the retrieved set, and abstention rate reported beside faithfulness.
Own the corpus as a product surface — versioning, authority ranking, deletion policy, ongoing ingestion hygiene — and argue why that investment, not generator tuning, is what moves a green-dashboard-wrong-answer system.
## The metric's reference point Every evaluation metric answers "consistent with what?" For faithfulness the answer is: the retrieved context passed to the generator. That choice is deliberate and useful — it isolates the generator's behaviour from the quality of the corpus — but it means the metric inherits whatever the corpus says. Truth is never consulted. A system whose corpus is wrong will have a perfectly grounded, perfectly wrong product, and the faithfulness dashboard will be green throughout. ## Three ways a fully faithful answer is wrong ### The source is wrong or out of date The most common case in enterprise RAG. Corpora accumulate superseded versions, drafts, region-specific variants, deprecated runbooks, and content written by users rather than by the organization. Retrieval finds the passage that best matches the query, not the one that is currently authoritative — an old version often matches *better*, because deprecated documents tend to state things plainly while the current one is hedged. The generator paraphrases it accurately. Every claim is entailed. The score is 1.0 and the user acts on a rule that changed last quarter. ### The sources conflict When the context contains two passages that disagree, any answer the model produces is entailed by *some* retrieved passage, so a naive per-claim entailment check passes whichever branch it takes. Worse, the choice is often arbitrary — driven by position in the context or by phrasing similarity rather than by recency or authority. The metric cannot distinguish "resolved the conflict correctly" from "picked one at random", because both are grounded. Detecting this needs a separate signal: flag contexts whose passages contradict each other, and treat unacknowledged conflict as a defect regardless of the faithfulness score. ### The answer is grounded but useless There is a degenerate corner where the metric can be optimized without improving anything. An answer that makes no factual claims — a refusal, a restatement of the question, a pure hedge — has nothing to be unfaithful about and scores 1.0 or undefined. Tighten the grounding instruction hard enough and faithfulness climbs while the product gets worse, because the model now abstains on questions it used to answer well. This is why abstention rate must be reported next to faithfulness: a metric that improves while the abstention rate doubles has not improved. A related corner is omission. An answer can be complete-looking, every claim supported, and materially misleading because it never mentions the exclusion, the exception or the eligibility condition sitting in the next paragraph. Faithfulness scores what was said; it has no opinion about what was left out. ## What to measure instead, alongside The three axes are genuinely separate and none substitutes for another: - **Faithfulness** — answer versus retrieved context. Catches generator hallucination. - **Answer relevance** — answer versus question. Catches drift and non-responsiveness. - **Answer correctness** — answer versus a gold reference answer. Catches everything the first two are blind to, including bad sources. Only the third requires labels, which is why teams skip it and then discover the blind spot in production. You do not need labels at scale: a modest curated set of questions with human-verified answers, refreshed as the corpus changes, is enough to catch a systematic corpus problem. Correctness low while faithfulness is high is the signature of a source problem, and it is unmistakable once you plot the two together. ## Fixes that live outside the generator Because the cause is upstream, so are the remedies: - **Freshness and versioning.** Carry effective dates and version metadata on chunks, filter or down-rank superseded content at query time, and delete rather than archive-in-place where you can. - **Authority ranking.** Not all sources are equal. Official policy outranks a support-ticket transcript; make that explicit in ranking or filtering rather than hoping the embedding sorts it out. - **Deduplication.** Near-identical chunks from multiple versions are the primary generator of silent conflicts. - **Conflict surfacing.** When retrieved passages disagree, prefer an answer that names the conflict over one that silently picks a side — and measure how often that happens. - **Ingestion hygiene as an ongoing job.** Corpus rot is continuous; a one-time cleanup regresses within months. ## How to talk about this in an interview The strong answer states the reference point in one sentence — faithfulness is measured against the retrieved context, not the world — then names the concrete cases, then says what you would add to catch them. It also concedes the metric's virtue rather than dismissing it: separating generator behaviour from corpus quality is exactly what makes faithfulness a good *component* metric. The mistake is treating a component metric as a product metric.
- What signal would reveal a corpus-freshness problem that faithfulness cannot?Answer correctness against a small human-verified reference set. Plot it beside faithfulness: high faithfulness with low correctness is the signature of accurate answers built on bad sources. Add operational signals too — the age distribution of retrieved chunks, and a flag for contexts containing mutually contradictory passages. Neither requires large-scale labelling; a few hundred curated questions refreshed as the corpus changes is usually enough to catch a systematic problem.
- Two retrieved passages give different values for the same rule. What should a well-behaved generator do?Surface the conflict rather than silently choosing. Ideally it names both values with their source and date, or applies an explicit precedence rule the prompt gives it — prefer the document with the later effective date, prefer official policy over transcripts. Silently picking passes faithfulness either way, so the behaviour has to be specified and measured separately, and the durable fix is to stop retrieving both versions.
- How can tightening a grounding instruction make faithfulness rise while the product gets worse?Because claim-free answers cannot be unfaithful. A stricter instruction pushes the model toward hedges, restatements and refusals, which score 1.0 or undefined and drag the average up while the user gets nothing. The guard is to report abstention rate and answer relevance beside faithfulness, and to treat any faithfulness gain that arrives with a jump in refusals as unproven until relevance holds steady.
saying these in an interview costs you the question
- Treating faithfulness as a measure of factual truth
- Assuming a green grounding score means the corpus is healthy
- Ignoring that refusals and hedges score as perfectly faithful
- Letting the model silently pick between contradictory passages
- Skipping reference-based correctness because labelling is inconvenient