A legal assistant cited a case that doesn't exist. How do you ground and verify its answers?
answer
- the citation is generated text too
- resolve identifiers against a real index
- fail closed when it does not resolve
- claims carry spans, spans get matched
- fresh checker, not the author
basics
~20 sCondition the answer on retrieved source text and forbid claims beyond it, require span-level attribution, resolve every citation against an authoritative case index and fail closed when it does not resolve, then run a verification pass that labels each claim supported, unsupported or not addressed.
solid answer
~50 sThree layers, in order of how much they buy. **Grounding**: put the actual source passages in context, instruct that the answer must come only from them, and give the system a real "these sources do not cover it" output so it is not forced to fill the gap. **Attribution you validate**: require every claim to carry the identifier of the passage it came from — but treat the citation as a claim too, because it is generated token by token and a fabricated reporter volume and page is exactly what a language model produces well. Resolve each identifier against an authoritative index, confirm the quoted span really appears, and drop the answer if it does not. **Verification**: a second pass re-reads the cited paragraphs and marks each atomic claim supported, unsupported or not addressed; unsupported claims are stripped or escalated rather than shipped. On a filing-bound surface, that gate blocks release and a human signs off.
code
json · 24 lines{
"answer_id": "a-1042",
"claims": [
{
"text": "A carrier may limit liability for delayed baggage.",
"source_id": "doc-7",
"quoted_span": "the carrier's liability for delay shall be limited",
"verdict": "supported"
},
{
"text": "The limit does not apply to business travellers.",
"source_id": "doc-7",
"quoted_span": "",
"verdict": "unsupported"
},
{
"text": "Claims must be filed within seven days.",
"source_id": "doc-9",
"quoted_span": "within seven days of receipt",
"verdict": "not_addressed"
}
],
"gate": "blocked"
}go deeper
Know that a citation produced by a model is generated text and can be invented, and that the basic defence is answering from supplied documents and checking that the cited source really exists.
Describe the layers concretely: source passages in context with a not-covered output, per-claim attribution with quoted spans, and identifier resolution against a real index that fails closed.
Add the verification pass as a separate call returning supported, unsupported or not addressed per claim, decide what each verdict does to the response, and sequence the controls by payoff when the budget is limited.
Own where the gate sits and who signs. Decide which surfaces block on verification versus annotate, what the residual rate is allowed to be on each, and how the claim-and-source audit trail supports a defensible answer after an incident.
## Why the citation is the dangerous part A fabricated citation is the worst-shaped failure a language model produces. The format is highly regular, so the model reproduces it perfectly; the content is arbitrary, so it has nothing to reproduce it *from*; and the artefact carries social authority, so a reader treats it as the evidence rather than as another generated claim. The result is a string that looks exactly like proof and is not. Any design that treats "it gave a citation" as reassurance has inverted the risk. Start from that: a citation emitted by the model is a claim requiring verification, at least as much as the sentence it supports. ## Layer one — grounding Grounding means the answer is produced conditioned on source text placed in the context, with the instruction that claims must come from that text. This converts an open recall problem into a reading problem, which is a much easier one, and it gives you something to check against afterwards. Two design details do most of the work. First, the instruction must be specific about scope: answer only from the passages provided, and if they do not address the question, say so. Second, the product must have somewhere for that refusal to go. If "the retrieved authorities do not address this point" has no place in the interface, the system will produce prose instead — you have designed the gap that the model then fills. Make the not-covered response a first-class result with the next action attached, such as which sources were searched and what to broaden. Grounding is necessary and not sufficient. The right passage can sit in the context while the model still merges in a half-remembered holding from pretraining, or generalizes a narrow ruling. That is what the next two layers are for. ## Layer two — attribution that is validated, not trusted Ask for structured output where each claim carries the identifier of the passage it rests on, plus the quoted span. Then check it mechanically: - **Resolve the identifier** against an authoritative index — a real case database, your document store, the corpus you actually retrieved from. If the identifier does not resolve, the answer fails closed. This single check kills the fabricated-case class outright and it is ordinary software, not model behaviour. - **Confirm the span exists** in the resolved document, by exact or near-exact match. A resolvable case with an invented quotation is the second failure mode, and string matching catches it. - **Confirm the cited passage was actually retrieved** for this request. A citation pointing at a real document that never entered the context is evidence the model produced it from memory. A useful discipline: never let the model author an identifier freely. Have retrieval hand it opaque passage ids and require the answer to reference only those, so an invented id is a syntactic error rather than a plausible one. ## Layer three — the verification pass The final layer is a separate call whose only job is checking. Split the answer into atomic claims, hand each one the source paragraphs it cites, and ask for a verdict per claim: **supported** (the passage entails it), **unsupported** (the passage contradicts it or does not entail it), or **not addressed** (the passage is about something else). Making the check a fresh call over the source text — rather than asking the writer to re-read its own work — matters: the checker sees the claim and the passage, not the reasoning that produced the claim, so it is not defending a conclusion it already committed to. What you do with the verdicts is the design decision. On a low-stakes drafting surface, strip unsupported claims and return the rest with a note. On a filing-bound surface, any unsupported claim blocks release and the item goes to a human with the claim and the passage side by side, which makes review a matter of checking attribution rather than re-reading the whole brief. Be honest about the cost. Verification adds a call per answer and latency proportional to how much source text must be re-read, and it is imperfect — a checker can wave through a subtly over-general claim. It is a large reduction, not a proof. ## Sequencing and what to say If asked what you would do first with a limited budget, the order is: hard-resolve citations (cheapest, kills the headline failure), then make abstention real, then ground with attribution, then add the verification pass on the surfaces where an error reaches a court or a customer. Fabricated-citation incidents are usually failures of the first item, which is the least model-dependent of the four. The interview answer that lands names the three layers, insists that the citation is itself generated and must be resolved externally, and closes on where the human sits. The answer that does not land proposes a better prompt.
- The model cites a real case but misstates its holding. Which layer catches that?Not the resolver — the identifier resolves fine. Span matching catches it only if the quoted text was invented. The layer that catches it is the verification pass: the checker reads the cited paragraphs against the claim and returns unsupported, because the passage does not entail the stated holding. This is why resolution alone is insufficient, and why the checker must be given the source text rather than asked to judge plausibility.
- Why run verification as a separate call instead of asking the model to double-check itself in the same response?Because a model continuing its own turn is conditioned on the reasoning that produced the claim and tends to defend it. A fresh call sees only the claim and the source paragraphs, with no commitment to the conclusion, and its task is narrow enough to be scored on its own. It also lets you use a different model or a cheaper one, run claims in parallel, and log the verdicts as an auditable artefact separate from the draft.
- How would you keep the latency cost of verification acceptable on an interactive surface?Scope it. Verify only claims that carry citations or fall in high-risk categories, check claims in parallel rather than in one long pass, and send only the cited paragraphs rather than whole documents. Stream the draft while verification runs, then mark or retract claims that come back unsupported. Reserve the blocking, everything-checked configuration for output that leaves the building — a filing, a customer email — where the latency is irrelevant next to the error.
saying these in an interview costs you the question
- Treats the presence of a citation as proof of the source
- Relies on prompt wording alone to stop fabricated cases
- Lets the model invent document identifiers freely
- Has the model verify itself inside the same response
- Assumes grounding alone guarantees a faithful answer