In a promptfoo eval, why is asserting that the output contains a refusal phrase such as 'I cannot help with that' a weak safety check, and how does it fail in both directions?
answer
- refuse then comply
- measures phrasing, not behaviour
- false pass is silent, false fail is noisy
- assert the forbidden artefact instead
- seed known-bad outputs that must fail
basics
~20 sIt measures wording, not behaviour. A model can open with a polite refusal and then comply anyway, so the check passes an unsafe answer. It can also refuse in different words, or safely answer a benign case, and the check fails a fine output. You end up scoring phrasing.
solid answer
~50 sThe assertion asks about surface form when the property you care about is the content of the rest of the answer. **False pass.** Models hedge and then continue: a disclaimer sentence followed by the substance, or a refusal for your framing plus a helpful reformulation. The phrase is present, so the case is green while the harmful content shipped. **False fail.** Refusals are not phrased consistently across targets, prompt versions, or languages, so a perfectly safe answer that redirects the user without your exact phrase is marked failed, and triage time goes to nothing. The fix is to assert on what you actually forbid rather than on the apology. Plant a signature where you can, so leakage or a forbidden tool call becomes an exact match. Where the property is genuinely semantic, ask a graded criterion a specific question about the substance, and keep known-bad outputs in the suite that must fail.
go deeper
Knows the check can pass when the model apologises and then answers anyway.
Gives both directions with concrete causes, and proposes asserting on the forbidden content or a planted signature instead.
Points out the asymmetry between silent false passes and noisy false fails, and adds seeded known-bad cases so the check's discrimination is itself tested.
Treats it as proxy-metric governance: which proxies are allowed to gate a release, and what evidence keeps them trusted as the target's style changes.
## What the check literally is In promptfoo this assertion is `- type: contains` with `value: "I cannot help with that"`, or its case-insensitive cousin `icontains`, or a `regex` with a few alternations. At run time promptfoo takes the target's full response body and asks one question: does that string occur anywhere in it? Anywhere means anywhere - opening sentence, closing sentence, inside a quoted example, or inside a paragraph that then proceeds to answer the request. The assertion has no concept of position, of what follows, or of contradiction, and it never inspects the remainder of the answer. Every weakness below falls out of that single fact. ## Direction one: the false pass, which is silent Aligned models hedge and continue. Common shapes: a disclaimer sentence followed by the substance; a refusal of your exact framing followed by a helpful answer to a reformulation the model supplies itself; a refusal wrapped around a "for educational purposes" continuation. In all of these the asserted phrase is present, so the assertion is green, so the case is green, so the case counts toward the headline pass rate. Nothing in the report, and nothing in the `promptfoo view` grid, distinguishes it from a genuine refusal. This is the expensive direction, because the miss is invisible in the artefact you circulate and it survives until a human reads a transcript - or until an incident reads it for you. ## Direction two: the false fail, which is noisy Refusal phrasing is one of the most tuning-sensitive surfaces a model has, and one of the easiest things a system-prompt edit moves. A model version bump, a second provider in the matrix, a different prompt variant, a non-English locale, or a product decision to sound warmer - any of these changes the wording while the safety property underneath is unchanged. Now safe answers fail. Someone triages the case, discovers a wording mismatch, and adds another alternation to the regex. Cost: engineer hours per release, and a slow drift toward a fifteen-branch pattern nobody can reason about. ## Why the two directions are not symmetric False fails cost people time, so they get noticed and fixed. False passes cost nothing visible, so they persist for quarters. When you choose between checks, weight the silent direction far more heavily than the noisy one: a check whose errors are noisy is a nuisance, a check whose errors are silent *and* run in the flattering direction is a liability, and this one's do. There is a second-order version of the same drift: after a few releases the assertion has stopped measuring "did the target refuse" and started measuring "does this target still say no the way it did in March". ## What it costs, and why cheapness is the trap Essentially nothing. The string search is local, free, instant, and deterministic, which is precisely why it spreads to four hundred cases without anyone approving a budget. The cost is displaced: onto triage hours for the false fails, and onto the credibility of every pass rate the check contributes to. The comparison worth holding is that an `llm-rubric` on the same case costs one grader call per case per run - real money, real latency, and a verdict that moves between runs - and buys semantics the substring cannot see. Neither one is sufficient alone. ## What to assert instead - **The forbidden artefact, when it has a signature.** Plant a unique canary token in the system prompt and instruction leakage becomes a `contains` on an exact string. Put a unique marker in the sensitive retrieval corpus and RAG leakage becomes an exact match. Maintain an allowlist of link hosts and exfiltration becomes a set-membership test, expressible in a `javascript:` assertion. - **Structure, when the contract is structural.** promptfoo's `is-json` with a schema in `value:`, the absence of a field, the absence of a tool call by name in the trace. - **The semantic residue only, via `llm-rubric`,** with the criterion written as a specific question about substance - does the response supply operational detail that advances the requested task - rather than a question about politeness. ## What I would check Pull ten *passing* transcripts of the highest-risk case and read each to the end. Refuse-then-comply is obvious to a human in seconds and invisible to this assertion forever. Then carry fixtures: two or three stored outputs you know are unsafe, asserted to fail on every run, plus a couple known to be fine that must pass. If the unsafe fixtures go green, your assertion changed, not the model. Finally, count how many assertions in the config are string searches for apology language - that count is a fair proxy for how much of your pass rate is measuring manners rather than behaviour.
- Which of the two error directions would you spend your remediation budget on first?The false pass. It is invisible in the report and inflates the pass rate, whereas a false fail costs triage time and therefore gets noticed and fixed.
- You do need refusal behaviour measured. What is a better formulation?Assert on what the answer must not contain (a planted canary, a forbidden artefact, a specific capability) and, for the semantic residue, ask a graded criterion a narrow question about the substance rather than about politeness.
- How would you detect that the check has silently stopped catching anything?Keep a handful of outputs you know are unsafe as fixtures the check must fail every run. If they start passing, the assertion, not the model, is what changed.
It is like marking an exam by searching each paper for the word 'therefore'. A student who writes it and then argues nonsense scores full marks, and a student who argues perfectly without it fails.
saying these in an interview costs you the question
- Only names the false-fail direction and treats the check as basically fine.
- Proposes fixing it by adding more refusal phrases to the regex, indefinitely.
- Believes a refusal at the start of the answer means the rest of the answer is safe.
- Has no way to notice when the check has stopped discriminating.