How do factual and faithfulness hallucinations differ, and why does the distinction change your fix?
answer
- two different reference points
- one against the world, one against the source
- true but unfaithful is possible
- grounding fixes one class, not the other
- attribution plus claim-level check
basics
~20 sA factual hallucination is wrong about the world; a faithfulness hallucination contradicts or exceeds the source text the model was given, even if the claim happens to be true. The first is fixed by supplying sources, the second by constraining and checking generation against them.
solid answer
~50 sFactual hallucination means the claim is false about the world — the model asserts something reality does not support. Faithfulness hallucination is defined relative to supplied material: the output contradicts, or simply is not entailed by, the documents in the context. The two are independent. A pharmacovigilance summarizer that adds a drug-interaction warning appearing in none of the label sections it was handed is unfaithful even if that interaction is genuinely documented elsewhere; conversely a perfectly faithful summary of a wrong source is faithfully wrong. The distinction drives the fix. Factual errors are attacked by supplying authoritative sources, narrowing the question, and hard-resolving identifiers. Faithfulness errors survive that — the sources were right there — so they need generation constrained to the provided text, an explicit "not covered by these documents" output, and a claim-level check against the source spans before the answer ships.
go deeper
Know the two words and one example each: wrong about the world versus not supported by the document you were handed. Be able to say that supplying source text targets the first kind.
Explain that the two are independent, with the true-but-unsupported case as proof, and map each to its own control — sources and identifier checks against factual error, constrained generation and attribution against faithfulness drift.
Show where faithfulness risk concentrates in a real pipeline — abstractive synthesis, partial coverage, many long sources — and describe the gate you would put before user-visible output, including what happens when a claim comes back unsupported.
Own the trust boundary. Decide which failures are the corpus's problem and which are the generator's, who owns each, and what residual rate each surface is permitted before human review becomes mandatory.
## Two different reference points Both failures look the same on screen: a confident sentence that should not be there. They differ in what you compare the sentence against. **Factual hallucination** is judged against the world. The model says something reality does not support — a birthdate that is wrong, a statute that does not say what it is claimed to say, a person who never held the role attributed to them. Ground truth lives outside the system, so detecting it requires an external authority. **Faithfulness hallucination** is judged against the material in the context. The model was handed documents and produced a claim those documents do not entail. This includes outright contradiction, but far more often it is quiet extrapolation: a hedge dropped, a number rounded into a different number, a general rule stated where the source stated a conditional one, or background knowledge from pretraining fused into the summary as if it had come from the source. ## They are genuinely independent The four combinations all occur, and interviewers probe whether you see that. A claim can be **unfaithful but true**. A pharmacovigilance summarizer is given three sections of a product label and outputs an interaction warning that appears in none of them. Suppose the interaction is real and well documented in the literature — the summary is still unfaithful, and in a regulated pipeline that is a defect regardless of truth, because the artefact is supposed to be a faithful rendering of that label, and a reviewer downstream will attribute the warning to a document that does not contain it. A claim can be **faithful but false**. If the supplied source is stale, wrong, or a low-quality page pulled in by retrieval, a scrupulously grounded answer inherits its error. Grounding moves the trust boundary to the corpus; it does not create truth. And of course claims can be both faithful and true, or unfaithful and false — the case people picture first. ## Why the distinction changes engineering The two classes respond to different controls. Against **factual** error, the levers are informational. Put authoritative material in front of the model instead of relying on parametric recall. Narrow the question so the answer is a lookup rather than an inference. Resolve anything with a canonical registry — citations, product codes, ticket ids — against that registry and fail closed when it does not resolve. Curate the corpus, because a grounded system is now only as correct as what you grounded it on. Against **faithfulness** error, informational levers are already spent — the right document was in the context and the model still drifted. The levers here are constraint and verification. Instruct explicitly that the answer must come only from the supplied passages and that anything not covered must be reported as not covered. Give the model a real abstention output, so "these sections do not mention interactions" is a success state rather than a failed answer. Ask for span-level attribution, so every claim points at the text it came from. Then run a check: a second pass that re-reads the cited passages and labels each atomic claim supported, unsupported or not addressed, with unsupported claims stripped, rewritten or escalated. For high-stakes output, that check should be the gate, not an advisory. ## Where faithfulness errors come from Understanding the mechanism helps you predict them. Generation is always conditioned on both the context and the model's priors, and the two are blended, not switched between. When the context is thin on a point the priors are strong on, the priors leak in — this is why summaries of unusual documents drift toward the typical document of that genre. Instruction-tuned models also carry a strong pull toward being complete and helpful, so a source that only partly answers the question invites the model to finish the thought. And the more the task asks for synthesis across passages, the more room there is for an inference the sources do not jointly support. Practically, that means faithfulness risk rises with abstractive rather than extractive tasks, with longer and more numerous sources, and with questions the corpus only partially covers — precisely the cases where the answer looks most impressive. ## What to say in an interview Define both, stress that they are independent by giving the unfaithful-but-true example, and then map each to its fix: sources and hard identifier checks for factual error, constrained generation plus attribution and a claim-level verification pass for faithfulness. If pressed on measurement, say that faithfulness is the one you can score without external ground truth — because the reference text is in hand — which is why grounded systems are usually monitored on it first.
- Which of the two can you monitor in production without external ground truth, and why?Faithfulness. The reference text is already in the request, so each claim can be checked against the passages that were supplied — no external authority, no labelled dataset of world facts. Factuality needs ground truth that lives outside the system, so it usually requires a curated eval set or human review. That asymmetry is why grounded systems ship faithfulness monitoring first and treat corpus quality as the separate lever for factual correctness.
- A summary is fully faithful to a retrieved document that is three years out of date. Is that a hallucination?Not a faithfulness hallucination — the model did its job. It is a factual failure of the system, located in the corpus rather than the generator, and blaming the model will send you fixing the wrong layer. The mitigations are corpus-side: freshness metadata, recency filters, and having the answer state the date of the source it relied on so a reader can judge staleness themselves.
- Why do abstractive summaries drift more than extractive ones?Abstractive generation rewrites in the model's own words, so every sentence is a fresh conditioning on both the source and the model's priors, and priors leak in where the source is thin. Extractive output copies spans, which bounds the failure to selection rather than invention. The practical middle ground is abstractive prose with mandatory span attribution, so each rewritten claim still points at the text it came from.
Think of a court reporter versus a fact-checker. The reporter's job is faithfulness — the transcript must match what was said, even if the witness lied. The fact-checker's job is factuality — whether what was said is true. Confusing the two jobs produces a transcript that quietly corrects the witness.
saying these in an interview costs you the question
- Treats faithful and true as the same property
- Says grounding makes an answer factually correct
- Assumes an unfaithful claim must also be false
- Blames the model for an out-of-date retrieved source
- Thinks longer, more synthetic summaries are safer