How do you verify RAG citations after generation without a judge model?
answer
- deterministic first, semantic later
- markers must be a subset of injected
- normalise before matching PDF text
- offsets carried through chunking enable highlights
- provenance verified is not entailment verified
basics
~20 sCheck mechanically: every emitted marker must belong to the injected set, every quoted span must be found in the chunk it cites after normalisation, and every load-bearing sentence must carry a marker. Failures trigger a regeneration, a stripped sentence, or an unverified badge.
solid answer
~50 sRun a deterministic pass over the produced answer before it reaches the user. First, set membership — markers not in the injected set are fabricated references and are always a defect. Second, span containment — normalise whitespace, soft hyphens, quotes and case, then look for each quoted span in the chunk it cites, allowing a similarity threshold rather than exact equality because models silently repair PDF extraction artefacts. Third, coverage — flag sentences that make a factual claim yet carry no marker. If you carried character offsets from the source document through chunking, a matched span also resolves to a highlight range in the original PDF, so a reviewer clicks a sentence and lands on the exact text. The whole pass is string work: milliseconds, no extra model call. What it cannot decide is whether a genuinely quoted span entails the claim built on it — that is a separate, more expensive layer.
code
python · 25 linesimport re
from difflib import SequenceMatcher
def normalize(text):
text = text.replace("", "").replace("-\n", "")
text = text.replace("’", "'").replace("“", '"').replace("”", '"')
return re.sub(r"\s+", " ", text).strip().lower()
def verify(quote, marker, injected, threshold=0.92):
if marker not in injected:
return "fabricated-marker"
q, src = normalize(quote), normalize(injected[marker])
if q in src:
return "verified"
best = max(
(SequenceMatcher(None, q, src[i:i + len(q)]).ratio()
for i in range(max(len(src) - len(q) + 1, 1))),
default=0.0,
)
return "verified-fuzzy" if best >= threshold else "quote-not-found"
injected = {"S1": "Employer match vests after\ntwo years of ser-\nvice."}
print(verify("Employer match vests after two years", "S1", injected))
print(verify("Employer match vests immediately", "S1", injected))
print(verify("anything at all", "S9", injected))go deeper
Know that after the answer is produced you can check that every citation marker was one you actually injected, and that quoted text really appears in the passage it cites.
Explain the normalisation needed before matching — whitespace, hyphenated line breaks, smart quotes, case — and why a similarity threshold beats exact equality when the source is extracted PDF text.
Lay out the layered guard: cheap deterministic checks inline on every request, semantic entailment checking reserved for sampled or high-stakes traffic, and a per-surface failure policy of strip, regenerate, flag or escalate.
Decide what verification the product guarantees and what it costs. Own the offset-plumbing discipline that makes click-through highlighting trustworthy, and the threshold policy that keeps warnings credible instead of routinely ignored.
## A guard, not a metric This is a runtime check that runs on every answer in the request path, distinct from offline scoring of a sample. Its job is to catch a broken answer before a user sees it, so it must be cheap enough to run inline and deterministic enough that its verdict is explainable. ## Layer one: marker validity The cheapest test. Parse the markers out of the answer and intersect with the set you injected. Any marker outside that set was invented — the model produced a reference-shaped string from memory or from pattern. This is a hard error with no benign explanation, and it is worth alerting on, because a system that fabricates references is worse than one that cites nothing: the fake citation is what stops the reader from checking. ## Layer two: span containment If the answer format includes quoted spans, test whether each span occurs in the chunk it cites. Naive exact matching fails constantly and for boring reasons, so normalise both sides first: collapse runs of whitespace, remove soft hyphens and rejoin words split across a line break, fold smart quotes and dashes to ASCII, case-fold, and strip page furniture like running headers. Then use a similarity threshold or an approximate substring match rather than equality — models routinely fix a typo, expand an abbreviation or drop a footnote marker while copying, and PDF text extraction introduces its own noise from column ordering and ligatures. Calibrate the threshold against a labelled sample. Too strict and you flag honest quotes constantly, teams start ignoring the warnings, and the guard is dead. Too loose and a paraphrase that changes the meaning slips through. ## Layer three: coverage Split the answer into sentences and ask which ones assert a fact but carry no marker. Uncited factual sentences are where background knowledge leaks in, and they are frequently the ones that turn out to be wrong. Framing and transition sentences legitimately carry no citation, so this check is a flag for review rather than a hard failure, and it is usually reported as a ratio: what fraction of asserting sentences are attributed. ## Offsets and click-through highlighting The same machinery gives you span-level attribution in the UI, provided the offsets survive the pipeline. Record, at extraction time, where each chunk starts and ends in the source document's character stream, and carry that through chunking, cleaning and indexing. When a quoted span matches inside a chunk at a local position, add the chunk's base offset and you have an absolute range in the source. A reviewer clicks a sentence in the answer and the viewer opens the source PDF at the exact highlighted passage. That is a large usability win in review workflows — legal, clinical, benefits — because it collapses verification from "open the document and search" to one click, and it is what converts citations from decoration into a working audit trail. The plumbing is the hard part: any cleaning step that rewrites text without adjusting offsets silently misaligns every highlight downstream, so offset mapping deserves its own tests. ## What to do on failure Options, roughly in order of cost: strip or mark the offending sentence and serve the rest; regenerate once with the failure fed back as an explicit correction instruction; serve the answer flagged as unverified so the interface can dim it or require an extra confirmation; or in high-stakes settings, refuse and escalate. Pick per surface — silently dropping content is fine for a summary panel and unacceptable in a clinical review workflow, where the reviewer needs to see that something failed. ## What it cannot catch Entailment. A span can be quoted perfectly, resolve to a real document and support nothing the sentence claims. Detecting that requires comparing claim to evidence semantically, which is the expensive layer and belongs in a separate stage with its own budget. The right framing is a ladder: deterministic checks first because they are nearly free and catch the crude failures, semantic checking above them on the traffic or the question classes that justify it. It also cannot catch a well-cited answer drawn from a wrong document. Provenance is not correctness — if the corpus contains a superseded policy page, verification will confirm faithfully that the answer came from it. ## Operational notes Budget it honestly: string normalisation and fuzzy matching over a handful of chunks is sub-100ms, but a careless implementation doing pairwise fuzzy comparison over every sentence and every chunk is not. Log per-check outcomes so you can see trends — a rising fabricated-marker rate after a model change is a strong signal, and it is invisible if you only log the final answer.
- What routinely breaks an exact-match check between a quoted span and PDF-extracted text?Hyphenated line breaks, soft hyphens, ligatures, smart quotes and dashes, multiple spaces from column layout, running headers and page numbers interleaved into the stream, and the model quietly normalising or correcting as it copies. Normalise aggressively on both sides and compare with a similarity threshold. Treat a mismatch as a signal to review, not as proof of fabrication, until the threshold has been calibrated on labelled examples.
- What if the model paraphrases instead of quoting — is the check useless?Containment stops working, so you either require verbatim spans in the output format or fall back to a semantic comparison. In practice this is the strongest argument for a quote-then-answer format: it exists partly to make the cheap check possible. Marker-validity and coverage checks still apply to a paraphrasing answer, so you keep the layers that catch fabricated references and unattributed claims.
- Should this run inline in the request path or asynchronously?Marker validity and span containment are string operations costing milliseconds, so run them inline where they can still change what the user sees. Anything needing a model call belongs asynchronously on sampled traffic, or inline only for question classes whose stakes justify the latency. A common shape is deterministic checks on every request and semantic checking on a sample plus all flagged responses.
- How do character offsets survive chunking and cleaning?By treating offsets as data that every transform must maintain. Record the chunk's start and end in the source stream at extraction, and make any step that rewrites text — dehyphenation, header stripping, normalisation — either preserve a mapping back to original positions or run before offsets are taken. Test it explicitly with fixtures, because a silent off-by-N misaligns every highlight in a way nobody notices until a reviewer complains.
saying these in an interview costs you the question
- Assumes a resolvable marker proves the claim is supported
- Exact-string-matches quotes against PDF text without normalisation
- Blocks every answer whenever a fuzzy match falls short
- Thinks post-hoc checking substitutes for fixing retrieval quality
- Rewrites chunk text during cleaning without maintaining source offsets