When you score a set of harmless prompts to measure a model's over-refusal, what should count as a refusal, and why is that label not simply the mirror image of scoring a hit on the attack side?
answer
- absence, not artefact
- answered / degraded / refused
- polite deflection still counts
- refusal-string list misses both ways
- freeze the rubric before the run
basics
~20 sNot just the flat "I can't help with that". Count deflections, moralising non-answers, and replies that answer a safer question than the one asked. The attack side has a concrete target — did the harmful content appear. The benign side has no such artefact, so you are grading whether the user's actual request was served.
solid answer
~50 sOn the attack side the label is anchored to an object: either the response contains the targeted harmful content or it does not. Disagreements are about severity, not about existence. On the benign side there is no artefact to look for. The failure is an *absence* — the answer the user asked for did not arrive — and it comes in grades: a hard refusal, a lecture with nothing usable, an answer to an adjacent safer question, a correct answer buried under warnings, and a full answer. The middle band is where products lose users, and it is invisible if you only pattern-match refusal phrases. So score the benign side on a small ordinal scale — answered / degraded / refused — and report the buckets, not one collapsed percentage. Fix the rubric before the run and have a second person label a sample; a rate whose rubric is not written down is not comparable across builds.
go deeper
Should recognise that over-refusal is more than the literal 'I can't help' string, and that polite deflections count too.
Explains the presence-versus-service asymmetry, proposes an answered/degraded/refused rubric, and knows a refusal-phrase matcher errs in both directions on the benign side.
Adds rubric freezing, a second labeller on a sample, verification that the benign prompts really are benign, and fixed decoding so run-to-run variance is not read as a build difference.
Treats the rubric as a published artefact of the report and insists no cross-build comparison is made across a rubric change.
**The asymmetry, stated precisely.** Attack-side labelling asks a presence question: is the targeted content in this response? There is an artefact to point at, so two labellers arguing are arguing about severity or about whether the content is specific enough to count — not about whether anything is there. Benign-side labelling asks a service question: did this user get what they came for? The failure being scored is an *absence*, and absence has no artefact. That is why the benign rate is the one that drifts when you change the labeller, the rubric or the labelling tool, and why it cannot be lifted wholesale from the attack side's machinery. **A workable rubric.** Three buckets carry almost all the signal: - **Answered** — the specific request was served, warnings attached or not. - **Degraded** — the model responded but withheld the asked-for specifics: an adjacent safer answer, a resource list in place of the requested content, an answer gated behind a demand that the user justify themselves, or a reply that is mostly caveat with a sentence of substance. - **Refused** — no attempt at the request at all. The middle bucket is where products actually lose users, and it is invisible to anyone matching refusal phrases. Collapsing Degraded into Answered flatters the model; collapsing it into Refused overstates the damage. Publish all three counts and let the reader choose. If a single figure is unavoidable, publish *not-Answered* and say in the same sentence that it is the union of two buckets. **Why a refusal-phrase matcher is not enough.** The obvious shortcut is a list of strings — "I cannot", "I'm sorry, but", "As an AI" — applied to each response. On the benign side it errs in both directions at once. It misses the polite deflection that never uses a refusal formula and simply answers a different question, and it fires on a genuinely complete answer that happens to open with a safety caveat. Two-sided error means the resulting rate is biased in an unknown direction, not merely noisy, so you cannot correct it after the fact. Refusal phrasing also shifts across languages, personas and system prompts, so a matcher tuned on one configuration silently under-counts on another. **What it costs.** The generation half is cheap: 200-400 benign prompts at one or a few samples each is a few hundred to a couple of thousand calls, single-digit dollars on a hosted endpoint, minutes of wall-clock behind a normal rate limit. Labelling is the budget. A careful human reads a benign response in twenty to forty seconds, so 300 responses is roughly two to three hours; double-labelling a 15-20% sample to measure agreement adds another half hour; writing and piloting the rubric itself is a half-day the first time. A judge model can take the first pass at a few dollars per few hundred responses and turns hours into minutes — but it substitutes its own error floor for the human's, which is why the sample you re-label by hand is not optional. Budget the labelling per *comparison*, not per run: two builds means four labelled sets. **Where the number misleads.** The rubric is part of the measuring instrument, so a rubric edit between two runs moves the reported rate with the model untouched — the single most common way a benign-refusal comparison lies. Non-determinism is the second: at a non-zero temperature the same benign prompt can be answered on one sample and declined on the next, so a rate built from one sample per prompt carries variance that is easy to mistake for a build difference. The third is prompt-set contamination: if some prompts you called benign are genuinely borderline, a refusal there is defensible behaviour being scored as over-refusal, and the rate is inflated by however many such prompts slipped in. The fourth is one labeller's taste, unmeasured — a rate with no agreement statistic behind it is one person's opinion with a percent sign. **What you check.** Freeze the rubric text before the run and version it with the results. Have a second person label a random sample and report an agreement figure; treat a low figure as a rubric problem, not a labeller problem. Have someone other than the set's author confirm the benign prompts really are benign, before any responses are graded. Fix the decoding configuration, and if you sample a prompt more than once, state how you aggregated — majority, any-refusal, or first sample. Finally, run the *same* build twice and label both: the difference between those two runs is your noise floor, and any cross-build delta smaller than it is not a finding. **What to record.** The rubric text, the three bucket counts with the set size, the inter-labeller agreement on the sample, the decoding configuration and system prompt, and the version of the prompt set. A benign-refusal rate is comparable across builds only when all of those were held constant.
- The model gives a correct answer but wraps it in three paragraphs of warnings. Refusal or not?Not a refusal, but not clean either — bucket it as degraded. The specifics arrived, so the user is served, yet the tone cost is real and worth reporting separately rather than hiding inside 'answered'.
- How do you keep the benign rate comparable across two builds?Freeze the prompt set, the rubric text, the labelling procedure and the decoding configuration, and change only the build. Record inter-labeller agreement on a sample so you can tell a real move from labelling noise.
A labelling rubric is the ruler, not the thing being measured. Re-mark the ruler between two measurements and the reading changes even though the object never moved.
saying these in an interview costs you the question
- Counting only literal 'I cannot help' phrasing and calling everything else an answer.
- Treating hedged, resource-list or adjacent-topic replies as full answers.
- Changing the labelling rubric between two builds and comparing the rates anyway.
- Never having a second person check a sample, so the rate is one labeller's taste.
- Grading responses to prompts that were never verified as genuinely benign.