You are testing a deployed input-classifier guardrail with a test set made entirely of attack prompts. What does that set measure, and what must you add before the result means anything?
answer
- attack half plus benign half
- block-everything scores 100%
- catches and wrong blocks together
- near-miss benign discriminates
- label from the policy, not the tool
basics
~20 sAn attack-only set measures one thing: how many attack prompts the classifier blocks. It cannot show how often it blocks ordinary users, because there is nothing harmless in the set to block wrongly. Add a benign half, labelled from the same policy, so you can report wrong blocks alongside catches.
solid answer
~50 sAn attack-only corpus gives you a catch rate and nothing else. A classifier that blocks every input scores perfectly on it, so the number cannot distinguish a good guardrail from a broken one. The missing half is benign traffic: prompts a real user would send that the policy says must pass. So the corpus needs two labelled halves, and you report two numbers together — the share of attack items blocked and the share of benign items wrongly blocked. Neither is interpretable alone. The benign half is the harder half to build well. Harmless small talk is trivially passed by anything; the benign items that actually discriminate are the ones that sit near the policy line — security questions asked for legitimate reasons, medical or legal questions, fiction, quoted abuse in a moderation complaint. Those are the ones a jumpy guardrail eats, and they are the ones your users will send.
go deeper
Says the set only measures catches, and that you need harmless prompts too so you can see wrongful blocks.
Reports both rates together, and knows a block-everything classifier scores perfectly on the attack half alone.
Structures the benign half into easy and near-miss strata, sources each with stated provenance, and reviews wrongly blocked items for clusters rather than reading only the aggregate.
Ties the wrong-block rate to a business cost the organisation has agreed, and makes the two-sided report the standard artefact any guardrail change must produce.
## What an input guardrail actually returns An **input-classifier guardrail** is a second model or hosted service that inspects the user's prompt before your application model sees it, and returns a verdict your code collapses into one bit. - Llama Guard emits `safe` or `unsafe` as its first output token, followed by the code of the policy category it believes was violated. - OpenAI's moderation endpoint returns a `flagged` boolean plus a `category_scores` map of per-category floats. - Azure AI Content Safety returns a `severity` value per harm category — 0, 2, 4, 6 — that your own code compares against a threshold you configured. Whatever the response shape, the application ends up holding allow-or-refuse, and that bit can be wrong in two independent ways: an attack sails through, or a legitimate request is refused. A test corpus exists to estimate how often each of those happens. ## Why an attack-only corpus can only see one of the two Feed the guard nothing but items the policy says must be blocked and the only quantity you can compute is the share it blocked — the **catch rate**. Now consider the degenerate guard: a function that returns `unsafe` for every input, or the real guard with its severity threshold pinned at 0 so everything trips it. It scores 100% on your suite. It is also useless, and the suite cannot say so. An attack-only set measures the guard's *willingness to block*, not its ability to *discriminate*, and willingness to block is exactly the property that costs you production traffic. ## The two halves and the two numbers Build the corpus as an **attack half** (items the written policy says must be blocked) and a **benign half** (items the policy says must pass), each item labelled by a human against that policy. Run both halves through the same guard version at the same threshold in the same run, and report two rates side by side: - the share of attack items blocked; - and the share of benign items wrongly blocked. Neither is interpretable alone. A tuning change that lifts catches from 70% to 90% while lifting wrong blocks from 2% to 15% is usually a regression, and any single headline number — "accuracy", "F1", a vendor's blended score — hides that trade completely. ## Building a benign half that discriminates Split it deliberately into two strata. - *Easy benign* is ordinary product traffic, sampled from your own logs with scrubbing and consent; it only confirms the guard is not comprehensively broken, because almost anything passes it. - *Near-miss benign* is where over-blocking actually lives: items that share topic and vocabulary with the attack half but whose intent the policy allows — a defender asking how a phishing kit is structured, a clinician asking about overdose thresholds, a novelist writing a villain, a user quoting the abusive message they are reporting. Write these by hand so you know their provenance; a benign half of small talk will report a wrong-block rate near zero no matter how jumpy the guard is. ## What it costs The **labelling dominates**. A 600-item corpus at roughly a minute per easy item and five or more per contested near-miss item, plus a double-labelled sample to check agreement, is a couple of engineer-days per build and a smaller recurring cost per refresh. Inference is usually the cheap half: a hosted moderation call is a fraction of a cent and a 600-item sweep finishes in minutes for a few dollars at most, while a self-hosted 8B-class guard is a GPU-hour question and a batch job. The bill jumps if you test the guard *in situ* — driving each item through the whole application — because you then pay the application model's tokens for every item as well. Put the numbers in the report: items per half, cost and wall-clock per run, and the guard version and threshold they were measured at. ## Where the numbers mislead Three readings go wrong routinely. - A near-zero wrong-block rate measured on an **easy benign half** reads as "no over-blocking" when it only means "we did not test near the line". - A catch rate and a wrong-block rate quoted from **different runs, thresholds or guard versions** are not a trade-off curve, they are two unrelated facts. - And a benign half whose **labels came from what the guard already let through** drives the wrong-block rate to zero by construction — the guard is grading itself. ## What to check before believing the result 1. Run the two **degenerate baselines** through the harness: an always-block stub must report 100% catches and 100% wrong blocks, and an always-pass stub must report 0% and 0%. If either comes back differently, the harness is broken, not the guard. 2. Then read the wrongly blocked benign items one at a time rather than trusting the aggregate; they nearly always cluster into two or three themes — a keyword, a language, a domain — and each cluster is a concrete fix. 3. Finally, confirm the two rates carry item counts, so nobody quotes a percentage that rests on three items.
- A guardrail blocks every input it sees. What does your attack-only suite report, and what does the two-sided suite report?The attack-only suite reports a perfect catch rate. The two-sided suite reports the same perfect catch rate next to a 100% wrong-block rate on the benign half, which immediately exposes it.
- Where do you get benign items that are actually hard for the guardrail?Write or collect items that share topic and vocabulary with the attack half but have legitimate intent — security education, clinical questions, fiction, and quoted abuse inside a moderation report.
- Roughly how many items do you need before the wrong-block rate is worth quoting?Enough that a handful of items does not swing it: with a few dozen benign items a single mislabel moves the rate by percentage points, so quote a count alongside the rate and treat tiny differences as noise.
A metal detector set to beep at every passenger catches every weapon in the airport. What separates a useful detector from a broken one is how many ordinary passengers it stops, and a test set made only of weapons contains no ordinary passengers.
saying these in an interview costs you the question
- Reporting only a catch rate and calling the guardrail good.
- Filling the benign half with small talk that anything passes.
- Deriving benign labels from what the guardrail already allowed.
- Treating a higher block rate as strictly better with no cost side.
- Assuming the vendor's own reported numbers cover your traffic.