skip to content

Over-Blocking

Without a benign hard-negative set a guard that refuses everything scores perfectly. Interviewers probe this because the cost of a cautious guard lands on the product, not on the security report.

on this pageshow

explore

questions

4

A guardrail test corpus for an input classifier contains only attack prompts. Why does a classifier that blocks every single input score perfectly on that corpus, and what must be added before the score means anything?

level: juniorimportance: must knowfreq 70%

answer

  1. one label, one observable error
  2. block-all scores 100%
  3. benign hard negatives
  4. two rates, opposite directions
  5. attach the blocked benign examples

basics

~20 s

Every item in the corpus is an attack, so blocking everything means catching everything: the measured block rate hits 100% and nothing penalises refusals. The corpus can only observe one error type. Add benign prompts the product really sees, and report how many of those were wrongly blocked alongside the attack numbers.

solid answer

~50 s

The corpus has a single label: attack. A test suite can only observe errors it has examples of, so an all-attack corpus can observe misses and nothing else. "Block everything" is therefore the score-maximising policy, and the number you publish cannot tell that policy apart from a well-tuned classifier. The fix is a labelled **benign hard-negative set** run through the same harness, so every report carries two numbers that move in opposite directions: share of attacks blocked, and share of benign requests blocked. Tightening the classifier moves both up; only quoting both makes the tradeoff visible. Where the simple answer breaks: benign items that are trivially benign — greetings, weather, small talk — produce a near-zero over-block rate that is just as uninformative as no benign set at all. The benign items have to resemble the traffic the product actually receives, including the awkward-looking-but-legitimate requests.

go deeper

for a junior

Should say the corpus is all attacks, so blocking everything scores perfectly, and that benign prompts are missing.

for a middle

Should name the benign hard-negative set, explain that the two rates move together when you tighten, and note that trivial benign filler is nearly as useless as none.

for a senior

Should treat it as an instrument defect with a reporting consequence, describe freezing the benign set across reruns, and insist on attaching real blocked benign examples.

for a principal

Should talk about who owns the over-block number organisationally, and about refusing to hand over a one-sided report that can only ever argue for a tighter filter.

### What the harness actually computes A guardrail harness runs one loop: take an item, send its text to the guard, reduce the guard's response to a decision, compare that decision to the item's label, increment a counter. The reduction step differs per guard — OpenAI's moderation endpoint returns a `flagged` boolean plus a `category_scores` map; Llama Guard is a generative model whose first output token is `safe` or `unsafe`, followed by violated category codes; Azure AI Content Safety returns a severity level per category and *you* choose the cut-off that means "block". Whatever the guard, the harness ends up with blocked/not-blocked per item and divides. The division is where the defect lives. A classifier has four outcome cells: correctly blocked, missed, correctly passed, wrongly blocked. The labels in your corpus decide which cells can ever be populated. A corpus carrying one label — `attack` — populates one row: correct blocks and misses. The wrongly-blocked cell is empty **by construction**, not by luck, so no arithmetic over that corpus can produce a number that penalises refusing. A policy of "block everything" sets the numerator equal to the denominator: a perfect score, permanently, for a guard that is a closed door. ### Why the tooling pushes you here Attack-only is the default because the attack half arrives free. garak's `--probes` selection instantiates probe classes whose entire purpose is to emit prompts that should fail, and its per-probe pass rates and report grades are computed over that attack-only population; nothing in garak ships a benign corpus. promptfoo's `redteam` generator likewise emits adversarial cases, and benign cases exist only as ordinary eval test cases that you author and assert on yourself. The benign half needs sourcing, scrubbing and labelling by someone who knows the product, so it is the half that slips. ### What it costs The endpoint bill is rarely the constraint. Five hundred benign items cost exactly the same per call as five hundred attack items — a self-hosted 8B-class guard classifies them in minutes on one GPU, and a hosted content-safety service bills per thousand text records, so the marginal spend is small beside an attack sweep firing thousands of generations. The real cost is human. At one to two minutes per item to label and adjudicate, five hundred items is roughly a reviewer-day plus a half-day on the disputed ones, and that reviewer — someone who knows the product's own policy — is usually the scarcest person on the engagement. Treating benign coverage as a money problem is a mis-diagnosis; it is a scheduling problem, and it has to be booked at kickoff, not scavenged from leftover capacity at the end. ### How the number misleads "Attacks blocked: 99.4%" gets read as "the guard is 99.4% right". It is not a statement about rightness; it is a statement about one class. The identical figure is produced by a well-tuned classifier and by a guard that refuses every request, and nothing else in the report separates them. The second reading is worse because it looks like progress. Catch rate rises from 94% to 99% between two runs and is written up as an improvement. Without a benign half, that rise is indistinguishable from someone having lowered the block threshold — the two hypotheses make identical predictions on an all-attack corpus, so the data cannot choose between them. The third is structural, and it is the one to say out loud to a stakeholder: on an all-attack corpus, no change that increases refusals can ever lower the score. The metric is unfalsifiable in exactly the direction the report will be used to argue. Every number in it points at "tighten", and the cost of tightening is invisible to the instrument that produced them. ### What to check before you believe it Run a control. Point the harness at a stub guard that returns "block" unconditionally, and score it. If it comes back at or near 100%, you have just demonstrated that your instrument cannot detect over-blocking — and you have a one-line artefact for the stakeholder who wants the whole budget spent on attacks. Then check the plumbing. Does the harness score items labelled benign at all, or does it treat a non-hit as nothing to record? Are there benign items in the corpus, and who decided they were benign — your team, or someone who knows the product's policy? Did any items error or time out and get dropped silently, which biases the over-block rate downward? Was the guard configuration constant for the whole run, so the two rates describe the same guard? If you inherit a suite with no benign half and cannot build one, quote its catch rate as a one-sided measurement and say so in the same sentence. An unqualified catch rate from an all-attack corpus is a number that cannot be wrong, and therefore cannot inform anything.

  • You inherit a guardrail suite with no benign items and cannot build one before the deadline. What do you put in the report?
    Quote the attack-block rate explicitly as a one-sided measurement, state that no over-block cost was measured, and recommend against tightening the guard on the strength of this run alone.
  • Does adding benign items change how you interpret a rise in the attack-block rate between two runs?
    Yes. Without the benign half a rise is indistinguishable from the guard simply becoming more willing to refuse; with it you can see whether the over-block rate rose too.

Grading a smoke alarm only on how many fires it detects rewards an alarm that shrieks continuously: it scores perfectly every time, and the test never hears from the family who stopped cooking.

saying these in an interview costs you the question

  • Treating a high attack-block rate as evidence the guard is good, with no benign set in the corpus
  • Proposing trivially benign filler (greetings, weather) as the benign half
  • Saying the product team will notice over-blocking on their own, so the red-team report does not need it
  • Assuming the harness scores benign items automatically without checking that it records their decisions

context

open as a page

You are assembling the benign half of a test corpus for a content-moderation guard. Where do the benign items come from, and what makes a benign request a hard negative rather than an easy one?

level: middleimportance: must knowfreq 55%

basics

~20 s

Pull benign items from the product's real traffic and from topics sitting right next to the blocked ones: safety questions, clinical or legal wording, security research, fiction, non-English phrasing. A hard negative looks like an attack on the surface but is a request the product must answer. Label them, review the labels, freeze the set.

open as a page

You are writing up over-blocking results for a moderation guard in front of a chat product. What counts as an over-block event, and why is a single aggregate false-positive percentage not enough for the product team to act on?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Count an over-block whenever a benign request fails to get the answer it deserved: a hard refusal, a hedged non-answer, or a silently rewritten reply. One overall percentage hides which kinds of user are hit, so break it out by benign category and hand over the actual blocked requests as examples.

open as a page

A guardrail assessment has a fixed number of paid calls to a hosted moderation endpoint. How do you decide what share of that query budget goes to benign hard negatives instead of attack probes, and how do you defend the split to a stakeholder who wants everything spent on attacks?

level: principalimportance: should knowfreq 30%

basics

~20 s

Spend enough on benign traffic to detect an over-block rate the product would actually care about, then give the rest to attacks. Size the benign half from the smallest rate worth acting on, not from what is left over. Defend it by noting a catch-rate-only report cannot tell a good guard from a closed door.

open as a page