skip to content

How do you write a RAG insufficient-context refusal rule that does not over-refuse?

level: seniorimportance: must knowfreq 60%

answer

  1. say what to emit, not just what to avoid
  2. condition must be textually testable
  3. strictness buys safety with recall
  4. false refusals need their own eval set
  5. most false refusals are retrieval defects

basics

~20 s

State a concrete abstention output tied to a testable condition — no passage states the fact — rather than a vague "don't make things up", allow partial answers with an explicit gap, and measure both error types: invented answers when context is missing and refusals when the answer was actually present.

solid answer

~50 s

A workable rule says what to emit, not only what to avoid. Define insufficiency operationally: no injected passage states the requested fact, the passages cover a different entity or period, or they conflict. Specify the abstention text or a structured flag so the application can detect it, and require the model to name what the passages *do* cover so the user can re-query. Permit partial answers — answer the covered part, mark the gap — because all-or-nothing rules throw away good information. Then evaluate two-sided: an unanswerable set measuring how often it answers anyway, and an answerable set with known gold passages measuring false refusals. A strict rule is easy to write and quietly destroys recall, and most of that damage is a retrieval failure being pushed onto the user rather than genuine absence of the fact.

go deeper

for a junior

Be able to say the prompt must tell the model to answer only from the provided passages and to state plainly when they do not cover the question, instead of answering from memory.

for a middle

Explain why the abstention condition must be textually testable — no passage states the fact, or the passages cover a different entity or period — and why the refusal needs detectable wording or a structured flag.

for a senior

Show the two-sided evaluation: an unanswerable set for invented answers and an answerable set with gold passages for false refusals, plus the diagnostic of checking whether the gold passage reached the context before touching the prompt.

for a principal

Own the asymmetry. Decide which error the product can afford in this domain, set target rates before tuning, build the escalation path refusals depend on, and make corpus gaps revealed by abstention feed the content roadmap.

## Abstention is a prompt-side contract Retrieval decides what evidence exists; the generation prompt decides what the model does when the evidence is not there. Left unspecified, a capable model will fill the gap from its parametric knowledge, and the result is an answer that reads exactly like a grounded one — same tone, same confidence, frequently no citation. That is the highest-severity failure in a grounded assistant, because the whole value proposition was "this comes from your documents". ## Say what to do, not only what to avoid "Do not hallucinate" is not an instruction; it names an outcome, not a behaviour. A usable rule specifies a condition, an output, and a shape. For a clinical-guideline assistant: if the provided guideline sections do not address the drug asked about, reply with the fixed sentence "Not covered by the provided guidance", state which sections were provided, and stop. The three parts matter — the condition is testable, the output is detectable, and telling the user what *was* retrieved converts a dead end into a next query. ## Define insufficiency operationally Models are poor at introspecting on whether they "know enough". They are much better at a textual test. Useful formulations: no passage states the requested fact; the passages concern a different entity, drug, plan year or jurisdiction than the question; the passages state the fact but disagree with each other. The last one deserves its own branch — the right behaviour for conflict is usually to surface both positions with their citations rather than silently pick one or refuse outright. ## Make the abstention machine-detectable Free-form apologies are unparseable. Either fix the wording exactly, or better, have the model emit a structured field alongside the answer indicating whether the question was answerable from the context. Detection is what lets you route: offer a broader search, escalate to a human, log the query as a corpus gap. It is also how you measure the abstention rate at all, which is one of the most informative operational metrics a RAG system has — a sudden spike almost always means retrieval broke, not that users started asking harder questions. ## Allow partial answers All-or-nothing rules are the main driver of user-perceived uselessness. If four of five sub-questions are covered, the useful behaviour is to answer those four with citations and explicitly flag the fifth as not covered. Write this into the rule; otherwise a strict instruction pushes the model to refuse the whole response over one uncovered detail. ## The recall cost, and why it is usually retrieval's fault Tightening the abstention rule always trades one error for another. Push hard and the model starts abstaining when the answer *is* in the corpus but is phrased differently from the question, split across two chunks, or stated implicitly in a table. From the user's seat that is indistinguishable from "your product doesn't know", and it is worse than a hedged answer because there is nothing to correct. Crucially, most of those cases are retrieval or chunking defects surfacing at the generation stage. The diagnostic is simple: for each false refusal, check whether the gold passage was actually in the injected set. If it was, the prompt or model is too strict. If it was not, stop tuning the instruction and go fix retrieval. ## Measure both sides You need two evaluation sets. An unanswerable set — questions whose answers are genuinely absent from the corpus, including near-miss distractors that retrieve plausible but non-answering passages — measures how often the system answers when it should not. An answerable set with known gold passages measures false refusals. Report both. Teams that track only hallucination rate will happily ship a system that refuses half the time and call it safe; teams that track only helpfulness ship the opposite. A single "accuracy" number hides the trade entirely. ## The asymmetry is a product decision Where the two errors sit is not an engineering choice. In a clinical setting an unsupported answer can cause harm and a refusal costs a clinician thirty seconds, so the rule leans hard toward abstention and the escalation path must be excellent. In an internal search assistant, over-refusal drives users straight back to grepping the wiki and the product dies of irrelevance. Name the asymmetry explicitly, set the target rates before tuning, and revisit them with real usage data rather than intuition. ## Common pitfalls Burying the rule mid-prompt where it competes with a dozen other instructions. Writing an abstention condition the model cannot test. Treating every refusal as correct behaviour and never sampling them. Letting refusals be silent — logged as a successful response — so the corpus gaps they reveal are never collected and fixed.

  • How do you tell an over-strict prompt from a retrieval failure when the system refuses a question it should answer?
    Check whether the gold passage was actually in the injected context for that request. If it was present and the model still abstained, the abstention rule or the model is too strict and you tune the prompt. If it was absent, the model behaved correctly and the defect is upstream in indexing, chunking, query rewriting or ranking. Without that per-request evidence log the two are indistinguishable and teams routinely tune the wrong stage.
  • What should the model do when two retrieved passages state contradictory answers?
    Surface the conflict rather than resolve it silently. State that the sources disagree, give both positions with their citations, and where the metadata supports it note which is more recent or more authoritative. Silently picking one produces a confidently wrong answer with a resolvable citation, which is the hardest failure for a reviewer to catch. A refusal is also wrong here — the useful information is precisely that the corpus is inconsistent.
  • Why prefer a structured answerable flag over a fixed refusal sentence?
    Because prose drifts. Even with fixed wording specified, models paraphrase, add apologies, or wrap the sentence in extra context, and a string match on your side then misses real abstentions. A separate field is unambiguous, survives translation and tone changes, lets you route to broader search or a human, and makes the abstention rate a metric you can trust rather than a regex you keep patching.
  • Is a rising abstention rate in production always bad?
    No — it is a signal, not a verdict. A spike usually means retrieval regressed, an index went stale, or a new user population is asking about content the corpus never had. Each cause has a different fix, and the third one is genuinely correct behaviour that should feed your content roadmap. Treat abstention rate as a monitored alarm with a triage path, not as a quality score to be minimised.

saying these in an interview costs you the question

  • Says "just tell it not to hallucinate" and stops there
  • Measures hallucination rate but never false-refusal rate
  • Treats every refusal as correct behaviour without sampling them
  • Uses free-form refusal prose the application cannot detect
  • Forces all-or-nothing answers when part of the question is covered
  • Tunes the prompt when the gold passage never reached the context

context