skip to content

Why can't hallucination be fully eliminated from a model trained to predict the next token?

level: middleimportance: should knowfreq 46%

answer

  1. the loss has no truth term
  2. plausible and true are not the same target
  3. generating is at least as hard as judging
  4. facts seen once cannot be generalized
  5. finite parameters, fixed cutoff

basics

~20 s

Next-token pretraining fits a distribution over plausible text, not a table of verified facts. For facts that appear rarely or never in training there is no signal to reproduce them, so a nonzero error rate on such questions is a property of the objective rather than a defect to patch.

solid answer

~50 s

Pretraining is density estimation: the model learns which continuations are likely, and a false but plausible continuation is exactly what that objective cannot penalize as false. Two consequences follow. First, generating a valid statement is at least as hard as deciding whether a statement is valid, so a model's error rate on open-ended generation is bounded below by how well the same evidence supports a yes/no validity judgement — you cannot generate reliably what you could not classify reliably. Second, for arbitrary facts the floor tracks how many of them appear only once in the corpus: if roughly a fifth of some category of facts is seen exactly once, a base model is expected to get about that share wrong, because a single occurrence gives almost nothing to generalize from. Add finite parameters and a training cutoff, and the honest target is a low, measured rate with abstention and grounding on top, not zero.

go deeper

for a junior

Know that the model is trained to predict likely text, not to check truth, and that rare or post-cutoff facts are the ones it gets wrong. Avoid promising that a prompt can eliminate the problem.

for a middle

Explain the mechanism: the loss contains no truth term, facts appearing once give nothing to generalize from, and finite capacity forces lossy reconstruction that fills gaps with typical values.

for a senior

Turn the limit into design. Show that you move facts into context, treat registry-backed identifiers as lookups, keep abstention as a real output, and report a measured residual rate per query class rather than a claim of zero.

for a principal

Set the expectation with stakeholders. Frame residual error as a budget allocated across surfaces, decide what each is allowed, and make the tradeoff between coverage, verification cost and human review explicit rather than implicit.

## The objective does not contain truth Pretraining maximizes the likelihood of the next token given the preceding text. What that fits is a distribution over what text looks like. Truth enters only indirectly, through the fact that true statements are more common in the corpus than any particular false one. Where that statistical correlation is strong — widely repeated, consistently stated facts — the model gets them right with high reliability. Where it is weak, the objective offers no mechanism that prefers the true continuation over an equally fluent false one. There is no term in the loss that says "this string is false"; there is only "this string was less likely in the data". That is the root of the phenomenon. Hallucination is not the model malfunctioning; it is the model doing exactly what it was trained to do on inputs where doing that is not enough. ## Generation is harder than judging A useful way to see why the floor is not zero: reduce generation to classification. Imagine a binary question — is this string a valid, correct statement? Any system that can reliably *produce* correct statements could be used to judge candidate statements, so generation is at least as hard as that classification problem. Learning theory gives every classifier a nonzero error rate whenever the features available do not separate the classes, and for arbitrary facts the features available in text are often exactly that: a well-formed sentence with a wrong number in it looks, statistically, like a well-formed sentence with the right number in it. So the generative error rate is bounded below by the misclassification rate of the corresponding validity problem. Anywhere validity is not statistically determinable from the training distribution — arbitrary dates, obscure identifiers, one-off events — the bound is well above zero, and no amount of scale removes it, because scale improves the estimate of the distribution, not the informativeness of the features. ## The singleton intuition The sharpest version of this concerns facts with no redundancy. Consider a class of facts, say the birthdays of minor public figures, and ask what fraction of them appear exactly once in the training corpus. A fact seen once provides essentially nothing to generalize from: the model cannot distinguish it from noise, cannot cross-check it, and has no second phrasing to anchor it. If about a fifth of that category is singleton, then a base model asked such questions is expected to be wrong on at least roughly that share — the error rate on arbitrary facts tracks the singleton rate. This is why hallucination is so lopsided across topics. The model is close to flawless on facts the internet repeats constantly, and unreliable in exactly the long tail where you most wanted help. It is also why a user's experience swings so violently: the failure clusters on rare queries, not uniformly. ## The other two floors Two further limits are less subtle but just as binding. **Capacity.** Parameters are finite and the training corpus is not. Pretraining is lossy compression; details that are cheap to drop get dropped, and the reconstruction that replaces them is a plausible average. This is why fabricated details are usually *typical* rather than random — the model fills the hole with the modal value of that slot. **Cutoff.** Anything that happened after training simply is not represented, and there is no internal marker distinguishing "after my cutoff" from "I don't recall". Asked about a recent event, a model has no way to notice the absence. ## What this changes about how you build Accepting the floor reframes the engineering. The goal is not a model that never errs; it is a system whose residual error is small, measured, and caught where it matters. That means moving facts out of the weights and into the context, where recall becomes reading. It means treating anything with a canonical registry as a lookup, not a memory. It means making abstention a first-class output, since a model that cannot say "I don't have that" must fill the gap. And it means measuring: publish the residual rate per query class rather than claiming a rate of zero, and put verification in front of the surfaces where an error is expensive. It also sets expectations correctly for a stakeholder conversation. "We reduced fabricated citations from 8% to 0.3% and hard-block the rest at the citation resolver" is a defensible claim. "We fixed hallucination" is not, and an interviewer will hear the difference immediately. ## What separates the strong answer Weak answers say the model "tries to be plausible" and stop. Strong answers name a mechanism: the loss contains no truth term; generation inherits a lower bound from the corresponding validity-classification problem; the floor for arbitrary facts tracks how many of them appear only once; capacity and cutoff add their own floors. Then they pivot to what you do about it, because an interviewer asking this is usually testing whether you will over-promise.

  • If the floor is intrinsic, why are newer models noticeably better on the same questions?
    Because the floor is not the same everywhere. Scale, better data curation and heavy synthetic augmentation move rare facts out of the singleton regime and improve the estimate of the distribution, which lowers the achievable error on the classes that improved. Post-training also teaches better hedging and abstention behaviour. None of that removes the bound for facts genuinely absent or seen once — it shrinks how much of your traffic falls in that regime.
  • Does supplying the source text in context change this argument?
    It changes which problem the model is solving. With the fact on screen, the task shifts from recall against a lossy compressed store to reading, where the evidence is present and the validity question is far better determined. The pretraining floor for that fact effectively no longer binds. A different floor appears — the model can still misread or over-extrapolate, and it inherits any error in the source — but it is a much lower one and, unlike the first, it is checkable against text you hold.
  • How would you communicate this limit to a stakeholder asking for zero hallucinations?
    Reframe from a promise to a budget. State the measured rate per query class, name the surfaces where an error is unacceptable, and describe the controls on those: grounding, hard resolution of identifiers, mandatory verification, human review. Committing to zero sets you up to either refuse to ship or quietly miss the target; committing to a measured rate with hard gates on the expensive paths is both honest and auditable.

saying these in an interview costs you the question

  • Says better prompting can drive the rate to zero
  • Claims enough scale eliminates the limit entirely
  • Describes recall as a lookup in a stored fact table
  • Assumes the training loss penalizes false statements
  • Treats the model as knowing when a fact is missing

context