In RAG, why prompt the model to quote supporting spans before answering?
answer
- evidence first, prose second
- extraction is easier than synthesis
- quotes are strings you can test
- cap quote length or it copies everything
- sourcing check, not entailment check
basics
~20 sQuote-then-answer makes the model select verbatim evidence from the retrieved passages first and synthesise only from those spans. Unsupported claims become visible, the quotes are cheaply checkable against the source, and the cost is extra output tokens and latency.
solid answer
~50 sThe output is split in two: an evidence block listing exact spans copied from the passages, each tagged with its source marker, then the prose answer built from those spans. It helps for two reasons. Extraction is a much easier task than open synthesis, so forcing extraction first makes the model commit to evidence before it starts composing, which reduces drift into background knowledge. And the quotes are strings you can test — a verbatim span either appears in the chunk it cites or it does not, no judge model required. On an employee-benefits corpus the instruction is literally "quote the exact sentence stating the vesting rule, then answer the question from those quotes". The costs are real: output tokens roughly double, latency rises, and a model can still quote correctly and then over-generalise in the prose, so the quote block is a check on sourcing, not on reasoning.
go deeper
Know the shape: the model lists exact sentences from the retrieved passages first, each tagged with its source, and only then writes the answer from those sentences.
Explain why the ordering matters — extraction is easier than synthesis and commits the model to evidence before it composes — and that verbatim spans can be string-checked against the chunk without another model call.
Discuss when the doubled output cost is worth it, how you normalise for PDF and whitespace drift before flagging a mismatch, and why a correct quote with a wrong inference still slips through.
Frame it as a policy decision: which question classes get the expensive traceable path, whether quotes are shown to users or used only internally, and what the reviewer workflow actually does with the evidence block.
## The technique Quote-then-answer asks for a two-part response. Part one is an evidence block: verbatim spans copied out of the injected passages, each carrying the marker of the passage it came from. Part two is the answer, written only from those spans. Some teams format part one as a small structured list so it can be parsed and stripped before display; others keep it visible as the receipts panel. ## Why it works Copying is a strictly easier task than free-form synthesis. When the model must produce the supporting text first, it performs a selection step over the context that is closer to reading comprehension than to generation, and it does so before it has committed to any wording of the answer. That ordering matters: once a model has written a fluent claim, the rest of the response tends to justify it. Extracting first makes the evidence the anchor and the prose the derivative. The second reason is mechanical. A verbatim span is a string. You can test whether it occurs in the chunk it cites without another model call, which converts "is this answer grounded?" from an expensive judgment into a cheap containment check for the sourcing half of the problem. ## Formatting choices Cap the quote length. A model allowed to quote freely will sometimes emit the entire chunk, which is technically verbatim and carries no signal about which sentence mattered — a limit of one or two sentences per quote forces selection. Require the marker on each quote, not on the block. Require that the answer's claims trace to the listed quotes, and forbid new claims appearing only in the prose. If you want a machine-readable trail, ask for the evidence block in a fixed structure so a parser does not have to guess where it ends. ## Interaction with abstention Quote-then-answer pairs naturally with an insufficient-context rule: if the model cannot find a span that states the fact, there is nothing to answer from, and the empty evidence block is the trigger. That is a cleaner condition than asking the model to introspect on whether it knows enough, because the check is about the presence of text rather than a feeling of confidence. ## What it costs Output tokens roughly double for a short answer, which is a real latency and price hit on a high-volume path, and for a streaming UI it delays the first useful token unless you render the evidence block as it arrives. There is also a quality cost in some settings: heavy quoting can make answers read as stitched-together fragments, and for questions that genuinely need synthesis across five passages the evidence block gets long enough that users skip it. ## Where it fails *Correct quote, wrong conclusion.* The model quotes the vesting rule accurately and then states a consequence the rule does not imply. Grounded sourcing is not entailment; the quote block does not catch faulty inference. *Near-verbatim drift.* The model normalises whitespace, fixes a typo, expands an abbreviation or silently repairs a hyphenated line break from a PDF. The span is honest but the exact-match check fails, so normalisation and fuzzy matching are needed before you treat a mismatch as a defect. *Wrong source, right text.* Boilerplate that appears in many documents matches several chunks; the quote verifies against the cited marker by accident. *Garbage in.* If the retrieved passage is itself stale or wrong, a perfectly quoted answer is a perfectly traceable wrong answer. Quoting improves traceability, not corpus correctness — though traceability is what lets you find the bad page and fix it. ## When to use it It earns its cost where a human will check the answer: regulated advice, policy and benefits questions, contract review, anything a reviewer signs off on. It is usually overkill for casual conversational retrieval where nobody clicks through and latency is the product. A middle ground is to run quote-then-answer only for question classes flagged as high-stakes, or to keep the quotes hidden and use them purely as an internal verification input.
- The model returns a quote whose wording differs slightly from the source PDF. Is that a defect?Usually not a hallucination — it is drift. Models normalise whitespace, repair hyphenated line breaks, expand abbreviations and silently fix typos while copying, and PDF text extraction introduces its own noise. Normalise both sides (collapse whitespace, strip soft hyphens, case-fold) and use a similarity threshold before flagging. Reserve the failure verdict for spans that have no close match in the cited chunk at all.
- Does quote-then-answer help if the retrieved passage is itself outdated or wrong?No. It makes the wrongness traceable, not absent. The answer will faithfully reflect a stale policy page, and the citation will resolve. That is still valuable because you can now find and fix the offending document, but the control for corpus correctness lives upstream in ingestion, freshness and provenance, not in the generation prompt.
- What stops the model from quoting the entire chunk to be safe?An explicit length cap and an instruction that each quote must be the minimal span that states the fact, typically one or two sentences. Without that, models do exactly this — it satisfies the letter of the instruction while destroying the signal, since the point of the quote is to identify which sentence mattered. It is worth measuring average quote length as a health metric.
saying these in an interview costs you the question
- Claims quoting eliminates hallucination rather than making it visible
- Never checks the quoted spans against the source text
- Lets the model quote whole chunks, which carries no evidence signal
- Ignores the doubled output tokens and added latency
- Treats a correctly quoted span as proof the conclusion follows