How does claim-level entailment scoring measure a RAG answer's faithfulness?
answer
- Decompose, verify, aggregate
- One verifiable assertion per unit
- Three verdicts beat two
- The denominator is gameable
- Nothing false said, yet still misleading
basics
~20 sSplit the answer into atomic factual claims, check each claim for entailment against the retrieved context, and report the supported fraction. Claim-level scoring localizes exactly which statement is ungrounded, which a single whole-answer verdict cannot do.
solid answer
~50 sThe pipeline has three stages. First, **decomposition**: a model rewrites the answer as a list of atomic, self-contained claims — one verifiable assertion each, with pronouns resolved. Second, **verification**: each claim is checked against the retrieved context and labelled supported, contradicted, or not-addressed, either by a natural-language-inference model or by a judge model with a strict rubric. Third, **aggregation**: faithfulness is the fraction of claims that were supported. An insurance-claim answer, for example, might decompose into six claims — the policy covers water damage, the deductible is $500, the reporting window is 30 days, and so on — and five may be entailed by the retrieved policy text while the deductible is invented. That yields 0.83 and, more usefully, names the offending claim. The cost is one extraction call plus one verification per claim, so it is far more expensive than a single holistic score — and it is only as good as the decomposition.
code
python · 16 lines# Verdicts produced by a per-claim entailment check against the retrieved context
verdicts = [
("the policy covers sudden water damage", "supported"),
("damage from gradual seepage is excluded", "supported"),
("the deductible is $500", "not_addressed"),
("claims must be reported within 30 days", "supported"),
("a licensed adjuster must inspect the property", "supported"),
("the policy limit is $50,000 per incident", "contradicted"),
]
supported = sum(1 for _, v in verdicts if v == "supported")
faithfulness = supported / len(verdicts)
unsupported = [c for c, v in verdicts if v != "supported"]
print(round(faithfulness, 2)) # 0.67
print(unsupported) # names the exact offending claimsgo deeper
Know the three stages by name — decompose the answer into atomic claims, check each against the retrieved context, report the supported fraction — and that the point is localizing which sentence was ungrounded.
Explain the mechanics: what makes a claim atomic and self-contained, why supported/contradicted/not-addressed beats a binary verdict, and how the fraction is computed and gamed.
Show judgment about when the per-claim cost is worth paying versus a cheap holistic score, and name what the method still misses — omission, misleading composition, and a context that was simply wrong.
Own the aggregation policy: whether releases gate on a mean, on zero contradicted claims, or on severity-weighted counts, and how that choice matches the domain's tolerance for a confidently wrong number.
## The idea Asking a judge "is this answer grounded?" produces a verdict you cannot act on. Claim-level entailment scoring turns faithfulness into many small, checkable questions instead of one big vague one. It borrows directly from fact-verification research: decompose the text into atomic claims, verify each against a source, aggregate. ## Stage 1 — decomposition A model rewrites the answer as a list of atomic claims. "Atomic" means one assertion per claim; "self-contained" means pronouns and references are resolved, because each claim will later be shown to a verifier without the rest of the answer. "Your policy covers it, but you'd pay the first $500" becomes two claims: *the policy covers water damage* and *the policyholder pays a $500 deductible*. Decomposition is where most of the error budget goes. Over-splitting produces fragments that are not independently verifiable and dilutes the denominator — twelve trivial claims from a three-sentence answer means one hallucination costs you 1/12 rather than 1/3. Under-splitting hides an invented detail inside a mostly-true compound sentence. Non-factual material — hedges, restatements of the question, offers to help further — should be dropped rather than scored, or every polite closing sentence becomes an unsupported claim. ## Stage 2 — verification Each claim is checked against the retrieved context, with three useful verdicts rather than two: - **Supported / entailed** — the context implies the claim. - **Contradicted** — the context says something incompatible. This is the severe case; the model overrode its own evidence. - **Not addressed** — the context is silent. The claim may be true in the world, but it came from pretraining, not from retrieval. Collapsing the last two into one "unsupported" bucket loses real information: contradiction usually points at a prompt or model problem, while a pile of not-addressed claims usually points at thin retrieval. Two implementations dominate. A dedicated NLI model is cheap, fast and deterministic, but works best on short sentence pairs and struggles when support is spread across several passages. A judge model with an explicit rubric handles multi-hop and paraphrased support far better, costs a call per claim, and needs its own validation before you trust it. Whatever the verifier, it must be given only the retrieved context — not the wider corpus, not the internet — otherwise it certifies claims the generator never had grounds for. ## Stage 3 — aggregation The standard score is supported claims over total claims. Simple, comparable across runs, and lossy in specific ways worth naming: - **Severity is flattened.** An invented deductible amount and an over-general adjective both cost one claim, though only one will generate a complaint. - **The denominator is gameable.** Verbose answers that pad with easily-supported restatements score higher than terse ones with the same single hallucination. - **Refusals score perfectly.** An answer with zero claims is either 1.0 or undefined; a system that abstains more looks more faithful. Common mitigations: report the *count* of unsupported claims alongside the fraction, weight contradictions more heavily than not-addressed, and gate on "zero contradicted claims" rather than on an average, since one confidently wrong number in a regulated domain is not offset by nine correct ones. ## What claim-level scoring still misses Even a perfect per-claim pass leaves gaps. **Omission** is invisible: an answer stating the policy covers water damage, every claim entailed, can be dangerously incomplete because it never mentions the exclusion in the next paragraph — nothing false was said. Relatedly, individually-supported claims can be **arranged** into a misleading whole through selection, ordering or emphasis; entailment is checked claim by claim, so the composition is never examined. **Implicit claims** — presuppositions carried by phrasing rather than asserted outright — often escape the decomposer. And faithfulness says nothing about whether the context itself was right, so a fully-supported answer built on a superseded document scores 1.0. ## When whole-answer scoring is the better trade Holistic scoring — one judge call, one verdict, optionally a short justification — costs a fraction as much and is adequate for coarse monitoring of high traffic, for smoke checks in CI, and as a cheap first tier that escalates only suspicious answers to full decomposition. Claim-level scoring earns its cost where you need to *localize* a failure: regression triage, comparing two prompts, or any domain where you must be able to point at the exact sentence that was not supported. Many teams run holistic scoring online and claim-level scoring offline on a sample, which keeps the diagnostic power without paying for it on every request.
- Why distinguish "contradicted" from "not addressed" instead of using a single unsupported bucket?They have different causes and different fixes. A contradicted claim means the generator overrode evidence it was given — usually a prompt, model or context-ordering problem. A not-addressed claim means the model answered from pretraining because the context was silent — usually a retrieval or coverage problem. Collapsing them tells you the answer is ungrounded without telling you which half of the pipeline to work on, and it hides severity: contradictions are the ones that generate complaints.
- How can a verbose answer game a claim-fraction score?The score is supported over total claims, so padding the answer with restatements of the question, definitions lifted from the context, and other trivially-entailed sentences inflates the denominator with guaranteed passes. One invented figure then costs 1/12 instead of 1/3. The defences are to report the absolute count of unsupported claims beside the fraction, gate on zero contradictions, and normalise or drop non-informative sentences during decomposition.
- Where does the biggest error in this pipeline usually come from?Decomposition, not verification. If claims are over-split, the denominator inflates and real failures are diluted; if under-split, an invented detail rides along inside a mostly-true compound claim and gets marked supported. Non-factual filler scored as claims produces false alarms. Before trusting the numbers, hand-inspect the extracted claims on a small sample and check that each is atomic, self-contained and genuinely verifiable.
saying these in an interview costs you the question
- Scoring the whole answer at once and calling it claim-level
- Letting the verifier see the corpus rather than only the retrieved context
- Treating an invented number and a vague adjective as equally severe
- Assuming a perfect claim score means the answer is complete
- Ignoring that a refusal produces a flawless faithfulness score