A red-team report says a model scored 0% attack success on a jailbreak benchmark. Why is that number alone not evidence the model is good, and what second measurement belongs beside it?
answer
- refuse-everything scores zero
- one-tailed metric
- attack rate + benign-refusal rate
- benign set must surface-match
- same build, same decoding
basics
~20 sA model that refuses everything scores zero on any attack suite, so zero can mean hardened or mean useless. The attack rate only reads next to a benign-refusal rate: run a set of harmless prompts, many of them phrased to look sensitive, and report how many were refused. Publish both numbers together.
solid answer
~50 sAn attack benchmark only asks whether harmful content came out. Nothing in it rewards answering a legitimate question, so the degenerate policy "refuse every request" is the global optimum of that scale. A 0% attack-success rate is therefore consistent with a well-tuned assistant *and* with a model that has been trained or wrapped into near-total refusal. The fix is to report the pair. Beside the attack rate you run a **benign set** — harmless prompts, deliberately weighted toward ones that surface-match the harmful ones (medication dosages, security concepts, violence in fiction, self-harm support resources) — and report the share of those the model declined. That second rate is the user-facing cost of the first. One number moving is meaningless; a build is only better if it lowered the attack rate *without* raising the benign-refusal rate. Both rates must come from the same build, same decoding settings and same system prompt, or the pair does not compare anything.
go deeper
Should say that refusing everything wins an attack benchmark, so the score needs a benign-refusal number next to it, and that both come from the same build.
Adds why the benign set must overlap the harmful distribution in surface features, and that the two rates should never be averaged into one index.
Adds the plumbing failure modes behind a suspiciously clean pair, insists on identical decoding and system prompt across both runs, and reads raw transcripts before reporting.
Frames the pair as the deliverable the red team owes a release review — two rates with their denominators and labelling rules stated — and is explicit about which decisions the pair cannot make.
**What the number actually is.** An attack-success rate is a fraction. The numerator is the count of attempts that a judge ruled produced the targeted harmful behaviour; the denominator is the count of attempts made. A published suite such as HarmBench, JailbreakBench or AdvBench fixes both halves for you: it ships the list of behaviours, the attack prompts or the recipe that generates them, and the rule that decides whether a response counted as compliance. Everything the suite measures lives inside that fraction, and the fraction contains no term at all for a request the model ought to have answered and did not. **Why that makes it one-sided.** The instrument can only ever punish answering. Declining costs nothing on this scale, so the degenerate policy "refuse every request" is its global optimum — a stub that replies "I cannot help with that" to every input scores a flawless 0% on any attack suite ever written. A metric that a null model maximises is a ceiling check, not a quality measure. Set the two rates beside each other and the asymmetry is plain: | | attack-success rate | benign-refusal rate | |---|---|---| | counts | attempts that produced the targeted behaviour | harmless prompts the model declined | | denominator | attack attempts made | benign prompts sent | | refuse-everything model | 0% — perfect | 100% — worst | | comply-with-anything model | near 100% — worst | 0% — perfect | Neither column is a safety score on its own. The pair is the measurement. **What the second measure has to look like.** A benign set of ordinary errands is too easy: almost no model declines "summarise this email", so such a set reports a flattering near-zero by construction and tells you nothing. The informative benign prompts are the ones that share surface features with the harmful list — same vocabulary, same topic, same imperative framing — but have a legitimate answer: medication dosing questions, security-education questions, violence inside fiction, self-harm support requests, dual-use technical questions. Those are exactly where an intervention that keys on topic or keywords rather than on intent reveals itself. The benign set is a designed artefact; how you build it determines what the rate can detect. **What the pair costs.** The API bill is the smallest line. AdvBench carries roughly 520 harmful behaviours and HarmBench a few hundred; at one generation per behaviour that is several hundred calls, and at five generations per behaviour a couple of thousand. With prompts and completions of a few hundred tokens each, a mid-priced hosted chat endpoint prices that sweep in single-digit to low-tens of dollars. A benign set of 200-400 prompts adds a comparable handful. Wall-clock is usually set by rate limits rather than compute: a 2,000-call sweep at 60 requests per minute is well under an hour, and running the attack and benign sets is a background afternoon. The expensive resource is a person. Benign-side grading is a judgment about whether the user's request was served, so a human reads 300 responses at twenty to forty seconds each — two to four hours per build, paid again on every build you compare. That is the real reason teams ship the attack rate alone, and skipping it is the defect this whole measurement exists to name. **Where the number misleads.** Four readings go wrong, in rough order of how often they are seen. First, and the point of the question: 0% is read as "hardened" when it is equally consistent with a model or wrapper tuned into near-total refusal. Second, 0% is read as "the true rate is zero". Zero hits over 40 behaviours leaves a 95% upper bound of roughly 7%; the suite has not excluded a real rate of several percent, it has only failed to see one. A zero from a 40-item run and a zero from a 400-item run are not the same claim. Third, 0% on both sides almost always means the harmful side never ran: a broken target wrapper, a truncated prompt file, a timeout path returning empty strings that score as non-hits, or a max-output-length that cuts completions before anything scoreable arrives. Fourth, public benchmark prompts leak into training corpora, so a low rate on a widely-published list can be memorised refusal on those exact strings rather than a defence that generalises. **What you check before believing it.** Confirm the denominator — how many behaviours, how many attempts each, how many errored — and put an interval around a zero rather than reporting a bare 0%. Read a sample of raw transcripts from both sides, not just the judge's labels; empty, errored and truncated responses are the usual culprits. Confirm the benign set genuinely overlaps the harmful distribution in surface features. And confirm both rates came from one build under one system prompt and one decoding configuration, because otherwise the pair does not compare anything. **What the pair does not settle.** It gives a builder the user-facing price of a safety change. It does not choose a filter threshold and it does not decide what refusal rate a product can bear — those belong to whoever owns the product surface.
- Which harmless prompts actually belong in the benign set?Ones that look like the attack list — same topic, same vocabulary, same imperative form — but have a legitimate answer, such as clinical, security-education, fiction-violence or self-harm-support requests. Neutral small talk is filler and reports a flattering zero.
- Both rates came back at zero. What do you check first?The plumbing. Confirm the attack prompts actually reached the model and produced non-empty responses, that the prompt set loaded fully, and that errors and empty strings are not being scored as non-hits. Read a sample of raw transcripts on both sides.
- Can you publish just the benign-refusal rate for a model you did not attack?Yes, but label it as what it is — a usability measurement with no safety claim attached. On its own it is as one-sided as the attack rate is, just pointed the other way.
A smoke alarm that you unplug never sounds a false alarm. Its false-alarm count is perfect, and it is worthless — the number only means something beside how often it fires when there is real smoke.
saying these in an interview costs you the question
- Treating a 0% attack-success rate as proof of safety with no benign-side measurement at all.
- Building the benign set out of easy small talk, so it reports a near-zero refusal rate by construction.
- Averaging the attack rate and the refusal rate into a single 'safety score'.
- Comparing an attack rate from one build against a refusal rate measured on another build or another system prompt.
- Never reading raw transcripts, so an empty or errored response is silently scored as a non-hit.