skip to content

How would you detect unsupported claims across production RAG traffic on a budget?

level: principalimportance: should knowfreq 40%

answer

  1. Cheap first, expensive rarely
  2. Numbers and names are the easy wins
  3. A sample that misses the risky tail teaches nothing
  4. A detector without a policy is decoration
  5. Watch what aggressive gating does to refusals

basics

~20 s

Run a cascade: cheap deterministic checks on every answer, a small entailment model on the suspicious ones, and an expensive judge or human review on a sample. Match spend to blast radius rather than scoring everything the same way.

solid answer

~50 s

Treat detection as tiered rather than uniform. **Tier 1, every request, near-zero cost.** Deterministic checks: do the numbers, dates, currency amounts and named entities in the answer appear in the retrieved context? Do quoted spans actually occur there? Most confident fabrications in RAG are numeric or nominal, and this catches a surprising share for the price of string matching. **Tier 2, flagged requests.** A small entailment model over the answer's claims — cheap, fast, deterministic, adequate for single-passage support. **Tier 3, sampled offline.** Full claim decomposition with a judge model, plus human adjudication on a slice, to estimate the true rate and to check the cheaper tiers are not systematically missing a failure class. Then decide what the signal *does*: block, warn, force a citation, or route to a human — and vary that by stakes. Also stratify sampling, because a uniform sample will barely touch the rare high-risk queries that matter most.

go deeper

for a junior

Know that checking every answer with a large model is too expensive, and that cheap checks — do the numbers and names in the answer appear in the retrieved passages — catch a real share of hallucinations.

for a middle

Describe the tiers and what each is good and bad at: deterministic string and numeric checks, small entailment models, judge models with a claim rubric, human labels for calibration.

for a senior

Show you would stratify sampling toward risky query classes, measure false negatives on cleared answers, and attach a graded action — advisory, regenerate, abstain, escalate — to each stakes level.

for a principal

Own the economics and the policy: what fraction of inference spend detection is allowed to consume, how that buys down the rate of undetected wrong answers, and how the org decides that tradeoff for a domain where a confident wrong number is expensive.

## Framing: detection is a portfolio, not a metric Scoring every production answer with a full claim-decomposition pipeline roughly doubles or triples the cost of the system and adds latency you cannot spend online. Scoring nothing means you learn about hallucinations from customers. The engineering answer is a portfolio of detectors with different cost, coverage and precision, arranged so that expensive signals only run where they pay. ## The cost ladder **Deterministic checks (cents per million answers).** Extract numbers, dates, monetary amounts, identifiers, and named entities from the answer and confirm each appears in — or is derivable from — the retrieved context. Verify that anything presented as a quotation is a literal span of a retrieved passage. These checks are cheap, deterministic, explainable and run inline. Their precision on numeric fabrication is high, which matters because the fabrications that cause real damage in enterprise RAG are overwhelmingly numeric or nominal: an invented amount, an invented deadline, an invented product name. Their recall on soft claims is poor, and they produce false alarms on legitimate arithmetic or unit conversion, so treat a hit as a flag rather than a verdict. **Small entailment models (fractions of a cent).** A compact natural-language-inference model over answer sentences against context passages. Fast, deterministic, cheap enough to run on a large fraction of traffic. Weak on multi-hop support spread across passages and on heavy paraphrase, so calibrate its threshold to favour recall and let the next tier arbitrate. **Judge models with a claim-level rubric (cents per answer).** The most accurate automated tier and the most expensive by one to two orders of magnitude, plus non-trivial latency. Reserve it for flagged answers, for the highest-stakes query classes, and for an offline sample. Any judge you rely on needs its own validity work — agreement against human labels on a slice — before its numbers carry weight. **Human adjudication (dollars per answer).** Not a monitoring tier; a calibration tier. Its job is to produce the ground truth that tells you what the automated tiers are missing. ## Sampling that actually finds things Uniform random sampling is the default and usually the wrong one. Hallucination risk is concentrated: queries with thin retrieval, queries whose context passages contradict each other, long answers, low-similarity top results, novel or out-of-distribution phrasing, and high-stakes query classes. Stratify by these and oversample the risky strata, then reweight when you report an overall rate so the headline number stays unbiased. A 1% uniform sample of mostly-easy traffic tells you almost nothing about the 0.2% of queries that generate incidents. Also sample the *passing* population deliberately. A detector whose false-negative rate you never measure will quietly degrade — new content types, a model upgrade, a prompt change — and only a labelled slice of answers it cleared will reveal it. ## Turning a score into an action Detection is worthless without a policy, and the policy should scale with blast radius: - **Advisory.** Log and dashboard only. Right for low-stakes informational queries where a false alarm costs more than a rare bad answer. - **Degrade.** Strip the unsupported sentence, or fall back to quoting the source passage instead of paraphrasing it. - **Regenerate.** One retry with the flagged claim named and a tighter grounding instruction. Bounded — a single retry, never a loop, or tail latency and cost become unpredictable. - **Abstain.** Return "I could not verify this in the available documents." Honest, and it converts a wrong answer into a mild disappointment. - **Escalate.** Route to a human before the answer is shown. Only affordable for narrow, high-stakes classes. Each step costs latency, cost or usefulness, so tie the choice to the query class rather than applying one policy to all traffic. And keep the abstention rate on a dashboard: aggressive detection policies push a system toward refusing, which looks excellent on grounding metrics and terrible to users. ## Operating it over time Three things drift and each needs a watch. **The corpus** changes, so a detector tuned on last quarter's content classes may misfire on new ones. **The generator** changes with every model or prompt update, shifting the mix of failure modes. **The judge** changes when its underlying model is upgraded, which silently moves your thresholds — pin versions where the provider allows it, and re-run a frozen labelled slice after any judge change so you can tell a real quality shift from a measurement shift. Finally, budget honestly. Detection is a fixed percentage tax on the system, and it is worth stating that percentage explicitly — say, five to ten percent of inference spend — so the tiering is designed to a number rather than discovered when the bill arrives. That framing also makes the tradeoff legible to non-engineers: you are buying a known reduction in undetected wrong answers for a known cost, and the right point on that curve is a business decision about the domain's tolerance for a confidently wrong answer, not a purely technical one.

  • Why do deterministic numeric and entity checks earn their place despite poor recall?
    Because cost and blast radius both favour them. They run inline at effectively no cost, they are deterministic and explainable, and the fabrications that cause real damage in enterprise RAG are disproportionately numeric or nominal — an invented amount, date or identifier. Low recall on soft claims is acceptable when the tier's job is to flag cheaply for a better checker, not to deliver a verdict. Their false alarms on legitimate arithmetic are the main tuning cost.
  • What goes wrong if you sample production traffic uniformly for faithfulness review?
    Risk is concentrated, so a uniform sample spends nearly all its budget on easy queries. Thin retrieval, contradictory passages, long answers and unusual phrasing carry most of the hallucination rate, and a small uniform sample will contain almost none of them. Stratify on those risk signals, oversample the risky strata, and reweight when reporting the overall rate so the headline stays unbiased while the review budget lands where failures live.
  • How do you keep an automated detection pipeline honest as models and corpus change?
    Freeze a labelled slice and re-run it after every judge, generator or corpus change, so you can separate a real quality movement from a measurement movement. Pin judge model versions where possible. Sample answers the detector cleared, not just the ones it flagged, to keep false-negative rate visible. And track abstention rate, since a tightening detection policy can improve every grounding metric while the product quietly stops answering.

saying these in an interview costs you the question

  • Running the most expensive judge on every single request
  • Sampling uniformly and missing the risky tail entirely
  • Producing a score with no policy attached to it
  • Never measuring what the cheap tiers falsely clear
  • Treating an upgraded judge model as a like-for-like replacement

context