What is hallucination in an LLM, and why doesn't telling it "don't make things up" fix it?
answer
- fluent and wrong look identical
- no separate store of known facts
- instructions change behaviour, not knowledge
- grounding, abstention, hard checks beat wording
basics
~20 sHallucination is a language model stating false or unsupported claims in the same fluent, confident register as correct ones. A prompt instruction cannot fix it because the model has no internal signal separating what it reliably knows from what it is inventing.
solid answer
~50 sHallucination is when a model produces content that is either false about the world or unsupported by the material it was given, delivered with no stylistic difference from a correct answer — a fabricated case citation is formatted exactly like a real one. The reason a system-prompt instruction only dents it is that the instruction changes behaviour, not knowledge. Telling a model to avoid inventing facts can raise how often it hedges or abstains, but it does not tell the model *which* of its claims are the unreliable ones; that information is not represented anywhere the instruction can reach. What moves the number is changing the inputs and the checks: put the source text in context and require answers to come only from it, make "I don't have that" an acceptable output, resolve identifiers such as citations or IDs against a real index, and verify claims against sources before the answer ships.
go deeper
Be able to define hallucination in one sentence and give a concrete example, such as an invented citation or an invented policy clause. Say plainly that prompt wording helps a little and grounding helps a lot.
Explain why the model cannot flag its own invented claims: next-token training builds no separate store of known facts, so weak recall and strong recall come out sounding identical. Name grounding and abstention as the real controls.
Show how you would engineer around a rate you cannot zero: source text in context, refusal paths that the product actually supports, identifiers resolved against an authoritative index, and a verification step before anything user-visible ships.
Own the framing that residual error is a budget to allocate, not a bug to close. Decide which surfaces get hard verification and human review, what error rate each is allowed, and how that is measured continuously rather than argued about once.
## What the word means A hallucination is output that a language model presents as fact but that is either false about the world or unsupported by the material the model was given. The defining property is not that the model is sometimes wrong — every system is sometimes wrong — but that the wrong output arrives in exactly the same fluent, confident register as a correct one. There is no stylistic tell, no hedge, no error code. An invented court case is formatted like a real one, complete with a plausible reporter volume and page. An invented internal-policy clause reads like the rest of the handbook. That uniformity is what makes hallucination a product problem rather than a curiosity. A system that failed loudly would be easy to wrap in a check. A system that fails in the same voice it succeeds in pushes the entire burden of detection onto the reader, and readers do not audit fluent text. ## Why the model has no "I made that up" flag A base model is trained to predict the next token of text. It learns which continuations are likely given everything before them. Nothing in that objective builds a separate store of "facts I hold" that could be consulted and reported empty. Recalling a fact is not a lookup in a table; it is reconstruction from distributed parameters, and reconstruction degrades smoothly. When the evidence in the weights is strong, the reconstruction is right. When it is thin — a name seen once, a number that never appeared at all — the same machinery still produces a fluent, well-shaped answer, because producing fluent well-shaped answers is what it was trained to do. The model does not experience the second case as different from the first. This is why "just ask it if it is sure" is weak. A confidence statement is itself generated text. It is produced by the same process, conditioned on the same context, and correlates with correctness only loosely unless the model has been specifically trained and calibrated for that behaviour. ## Why the instruction helps a little, and only a little An instruction like "answer only if you are certain; otherwise say you don't know" is not useless. It shifts the output distribution toward hedging and abstention, and on questions the model half-knows it will refuse more often. Two limits bite quickly. First, it supplies no new information. The instruction cannot tell the model which claims are shaky, because the model's own uncertainty is not legible to it in a form the instruction can act on. So abstention becomes roughly untargeted: the model may refuse things it actually knows and still assert things it does not. Second, it competes with everything else the model was trained to be. Post-training rewards answers that are helpful, complete and confident; users and graders alike prefer a direct answer to a shrug. A single line in a system prompt is a weak counterweight to that pressure, which is why models trained on accuracy-only scoring keep guessing even when told not to. ## What actually reduces it Four levers do real work, roughly in order of payoff: **Ground the answer.** Put the relevant source text into the context and instruct the model to answer only from it. This converts an open recall problem, where the model must dredge a fact out of its weights, into a reading problem, where the fact is on screen. It does not eliminate error, but it changes the failure into one you can check mechanically. **Make abstention a legitimate output.** If the product has no path for "the documents I was given do not cover this", the model will fill the gap. Design the refusal, write it into the prompt, and — crucially — make sure your own evaluation gives credit for it rather than scoring it as a miss. **Hard-check the checkable.** Anything with an authoritative index — case citations, SKUs, ticket numbers, API names, drug identifiers — should be resolved against that index before the answer ships, and the answer failed closed if it does not resolve. This is ordinary software, not model behaviour, and it catches the most damaging class of fabrication. **Verify claims against sources.** A second pass that re-reads the cited passages and labels each claim supported, unsupported or not addressed catches drift that grounding alone lets through. ## Saying it in an interview Give the definition in one line, then immediately say why prompting is a weak control: the instruction changes willingness to answer, not knowledge of which answers are safe. Interviewers are listening for whether you reach for prompt wording as the fix — the junior answer — or for grounding, abstention and verification, which is the answer of someone who has shipped this.
- If a user asks the model to rate its own confidence, how much should you trust that number?Only as much as you have measured. A self-reported confidence is generated text produced by the same process that produced the answer, so it can be as fabricated as the claim it describes. It becomes useful only if you check it empirically — bucket answers by stated confidence and measure accuracy per bucket. If the 90%-confident bucket is right 60% of the time, the number is decoration, and you should gate on grounding and external checks instead.
- Does a bigger or newer model make hallucination a solved problem?No. Scale reliably lowers the rate on facts that are well represented in training, and modern models are noticeably better than earlier ones. But rare and long-tail facts stay weakly supported at any scale, anything after the training cutoff is simply absent, and a bigger model that is wrong is wrong more persuasively. Treat scale as a rate reduction, never as a guarantee, and keep the grounding and verification layers regardless of which model you deploy.
- Where would you put a human in the loop if you cannot eliminate the residual rate?Where an error is expensive and irreversible, and where verification is cheap relative to the harm: outbound claims to customers, anything filed with a regulator or a court, changes to records of account. Route those through review while letting low-stakes drafting flow. Make the review efficient by surfacing each claim next to its source span, so the human checks attribution rather than re-reading everything.
saying these in an interview costs you the question
- Claims a strong system prompt eliminates hallucination
- Says the model knows it is guessing and chooses to bluff
- Treats self-reported confidence as a reliable signal
- Believes a larger model removes the problem entirely
- Confuses hallucination with malformed output or a safety refusal