skip to content

Test-Set Design

Stated provenance for the attack and benign sets, injecting layer by layer, and a suite that survives a guard upgrade. Interviewers probe it because a stale test set certifies safety it never checked.

on this pageshow

explore

questions

13

You are testing a deployed input-classifier guardrail with a test set made entirely of attack prompts. What does that set measure, and what must you add before the result means anything?

level: juniorimportance: must knowfreq 72%

answer

  1. attack half plus benign half
  2. block-everything scores 100%
  3. catches and wrong blocks together
  4. near-miss benign discriminates
  5. label from the policy, not the tool

basics

~20 s

An attack-only set measures one thing: how many attack prompts the classifier blocks. It cannot show how often it blocks ordinary users, because there is nothing harmless in the set to block wrongly. Add a benign half, labelled from the same policy, so you can report wrong blocks alongside catches.

solid answer

~50 s

An attack-only corpus gives you a catch rate and nothing else. A classifier that blocks every input scores perfectly on it, so the number cannot distinguish a good guardrail from a broken one. The missing half is benign traffic: prompts a real user would send that the policy says must pass. So the corpus needs two labelled halves, and you report two numbers together — the share of attack items blocked and the share of benign items wrongly blocked. Neither is interpretable alone. The benign half is the harder half to build well. Harmless small talk is trivially passed by anything; the benign items that actually discriminate are the ones that sit near the policy line — security questions asked for legitimate reasons, medical or legal questions, fiction, quoted abuse in a moderation complaint. Those are the ones a jumpy guardrail eats, and they are the ones your users will send.

go deeper

for a junior

Says the set only measures catches, and that you need harmless prompts too so you can see wrongful blocks.

for a middle

Reports both rates together, and knows a block-everything classifier scores perfectly on the attack half alone.

for a senior

Structures the benign half into easy and near-miss strata, sources each with stated provenance, and reviews wrongly blocked items for clusters rather than reading only the aggregate.

for a principal

Ties the wrong-block rate to a business cost the organisation has agreed, and makes the two-sided report the standard artefact any guardrail change must produce.

## What an input guardrail actually returns An **input-classifier guardrail** is a second model or hosted service that inspects the user's prompt before your application model sees it, and returns a verdict your code collapses into one bit. - Llama Guard emits `safe` or `unsafe` as its first output token, followed by the code of the policy category it believes was violated. - OpenAI's moderation endpoint returns a `flagged` boolean plus a `category_scores` map of per-category floats. - Azure AI Content Safety returns a `severity` value per harm category — 0, 2, 4, 6 — that your own code compares against a threshold you configured. Whatever the response shape, the application ends up holding allow-or-refuse, and that bit can be wrong in two independent ways: an attack sails through, or a legitimate request is refused. A test corpus exists to estimate how often each of those happens. ## Why an attack-only corpus can only see one of the two Feed the guard nothing but items the policy says must be blocked and the only quantity you can compute is the share it blocked — the **catch rate**. Now consider the degenerate guard: a function that returns `unsafe` for every input, or the real guard with its severity threshold pinned at 0 so everything trips it. It scores 100% on your suite. It is also useless, and the suite cannot say so. An attack-only set measures the guard's *willingness to block*, not its ability to *discriminate*, and willingness to block is exactly the property that costs you production traffic. ## The two halves and the two numbers Build the corpus as an **attack half** (items the written policy says must be blocked) and a **benign half** (items the policy says must pass), each item labelled by a human against that policy. Run both halves through the same guard version at the same threshold in the same run, and report two rates side by side: - the share of attack items blocked; - and the share of benign items wrongly blocked. Neither is interpretable alone. A tuning change that lifts catches from 70% to 90% while lifting wrong blocks from 2% to 15% is usually a regression, and any single headline number — "accuracy", "F1", a vendor's blended score — hides that trade completely. ## Building a benign half that discriminates Split it deliberately into two strata. - *Easy benign* is ordinary product traffic, sampled from your own logs with scrubbing and consent; it only confirms the guard is not comprehensively broken, because almost anything passes it. - *Near-miss benign* is where over-blocking actually lives: items that share topic and vocabulary with the attack half but whose intent the policy allows — a defender asking how a phishing kit is structured, a clinician asking about overdose thresholds, a novelist writing a villain, a user quoting the abusive message they are reporting. Write these by hand so you know their provenance; a benign half of small talk will report a wrong-block rate near zero no matter how jumpy the guard is. ## What it costs The **labelling dominates**. A 600-item corpus at roughly a minute per easy item and five or more per contested near-miss item, plus a double-labelled sample to check agreement, is a couple of engineer-days per build and a smaller recurring cost per refresh. Inference is usually the cheap half: a hosted moderation call is a fraction of a cent and a 600-item sweep finishes in minutes for a few dollars at most, while a self-hosted 8B-class guard is a GPU-hour question and a batch job. The bill jumps if you test the guard *in situ* — driving each item through the whole application — because you then pay the application model's tokens for every item as well. Put the numbers in the report: items per half, cost and wall-clock per run, and the guard version and threshold they were measured at. ## Where the numbers mislead Three readings go wrong routinely. - A near-zero wrong-block rate measured on an **easy benign half** reads as "no over-blocking" when it only means "we did not test near the line". - A catch rate and a wrong-block rate quoted from **different runs, thresholds or guard versions** are not a trade-off curve, they are two unrelated facts. - And a benign half whose **labels came from what the guard already let through** drives the wrong-block rate to zero by construction — the guard is grading itself. ## What to check before believing the result 1. Run the two **degenerate baselines** through the harness: an always-block stub must report 100% catches and 100% wrong blocks, and an always-pass stub must report 0% and 0%. If either comes back differently, the harness is broken, not the guard. 2. Then read the wrongly blocked benign items one at a time rather than trusting the aggregate; they nearly always cluster into two or three themes — a keyword, a language, a domain — and each cluster is a concrete fix. 3. Finally, confirm the two rates carry item counts, so nobody quotes a percentage that rests on three items.

  • A guardrail blocks every input it sees. What does your attack-only suite report, and what does the two-sided suite report?
    The attack-only suite reports a perfect catch rate. The two-sided suite reports the same perfect catch rate next to a 100% wrong-block rate on the benign half, which immediately exposes it.
  • Where do you get benign items that are actually hard for the guardrail?
    Write or collect items that share topic and vocabulary with the attack half but have legitimate intent — security education, clinical questions, fiction, and quoted abuse inside a moderation report.
  • Roughly how many items do you need before the wrong-block rate is worth quoting?
    Enough that a handful of items does not swing it: with a few dozen benign items a single mislabel moves the rate by percentage points, so quote a count alongside the rate and treat tiny differences as noise.

A metal detector set to beep at every passenger catches every weapon in the airport. What separates a useful detector from a broken one is how many ordinary passengers it stops, and a test set made only of weapons contains no ordinary passengers.

saying these in an interview costs you the question

  • Reporting only a catch rate and calling the guardrail good.
  • Filling the benign half with small talk that anything passes.
  • Deriving benign labels from what the guardrail already allowed.
  • Treating a higher block rate as strictly better with no cost side.
  • Assuming the vendor's own reported numbers cover your traffic.

context

open as a page

A deployment screens each user message with an input classifier and then screens the model's reply with a separate output classifier. You run an attack corpus end to end against that live path and every case is blocked. Why does that green result not tell you which of the two classifiers earned the block, and what test design does?

level: juniorimportance: must knowfreq 55%

basics

~20 s

End to end you only see the stack's final verdict, so the first layer that blocks hides every layer behind it. To attribute, run the corpus once per layer with the others put into log-only mode, and record per case which layer fired. Then a green run says which classifier actually earned it.

open as a page

You built the attack half of your guardrail test corpus by downloading a public jailbreak-prompt dataset, and the guardrail's vendor lists that same dataset among its training sources. Why is the resulting catch rate uninformative, and what do you build instead?

level: middleimportance: must knowfreq 62%

basics

~20 s

The classifier was fitted on those exact strings, so a high catch rate scores memorisation, not detection. It tells you nothing about a phrasing it has never seen. Build a held-out attack half you author or transform yourself, keep it unpublished, and record where every item came from so contamination can be checked later.

open as a page

Which cases earn a permanent slot in a guardrail regression suite — the fixed set re-run whenever a content guard is upgraded or the model behind it changes — and which should stay out?

level: middleimportance: must knowfreq 62%

basics

~20 s

Promote cases that once produced a wrong verdict and were then fixed: a payload confirmed to slip past the guard, and a legitimate request it wrongly refused. Each carries the decision you expect from now on. Keep out untriaged scanner output, twenty paraphrases of one technique, and cases whose correct verdict is genuinely arguable.

open as a page

A guardrail regression suite has failed the same three cases for months. This week it passes all three, with no change to the suite. What do you check before recording it as fixed, and what should each run have been recording to let you answer that?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Treat a green that nobody caused as suspect. Check whether the stack changed underneath the suite: guard version or endpoint, its policy configuration, and the model serving responses. Also check the harness — a verdict check that errors and counts as pass looks identical. Every run should record what it tested against, not just the results.

open as a page

In a stack where an input classifier screens user text before the model and an output classifier screens the reply, your attack corpus is blocked at the input stage, so the output classifier is never exercised. What are your options for testing the output classifier alone, and what does each option distort?

level: middleimportance: should knowfreq 45%

basics

~20 s

Two ways. Put the upstream classifier into a mode where it records a verdict but does not stop the request, so payloads reach the model and the reply is screened normally. Or drive the output classifier directly with fixture texts that stand in for replies. The first keeps the real path; the second skips generation entirely.

open as a page

When you promote a case into a guardrail regression suite, what should be stored with it so a run months later can decide pass or fail on its own — and why is storing the response text captured on promotion day a poor expectation?

level: middleimportance: should knowfreq 48%

basics

~20 s

Store the input, the decision you expect (blocked or allowed), the policy category it should fire on, and a one-line reason the case exists. Pinning the exact response text fails because generation is not stable across model swaps or sampling, so the case goes red on harmless wording changes and gets deleted or ignored.

open as a page

When hand-labelling items for a guardrail test corpus, why must each label come from your written content policy rather than from what the guardrail itself returned, and what does that imply for how you choose benign items?

level: seniorimportance: should knowfreq 44%

basics

~20 s

If the classifier's own verdict becomes the label, it agrees with itself by construction and every error rate collapses to zero. The label must be an independent human judgement against the written policy. That also means benign items should be chosen near the policy line, not far from it, or the pass half proves nothing.

open as a page

Two moderation layers in the same stack each block about 90% of a 500-case attack corpus when measured in isolation. Why can you not conclude that dropping either one costs roughly ten points of coverage, and what measurement answers the question properly?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Because the two 90s may cover the same 450 cases. What matters is per-case overlap, not per-layer totals. Record every layer's verdict for every case in one table, then count the cases only one layer caught. That unique-catch column is the layer's marginal contribution; the aggregate rate hides it.

open as a page

Your guardrail test corpus is half attack items and half benign items, but production traffic is overwhelmingly benign. How do you choose the corpus mix, and how do you report the results so the numbers still say something about production?

level: principalimportance: should knowfreq 33%

basics

~20 s

Keep the balanced mix for measurement — you need enough attack items to estimate the block rate at all — but never quote a corpus-level precision as if it were production's. Report the two rates separately, then project the daily wrongly blocked volume using production's real benign volume and its base rate.

open as a page

A defence you are engaged to test has four screening stages: an input classifier, a rules layer, a system-prompt instruction, and an output classifier. One isolation run per stage multiplies your inference bill and your engagement hours. How do you decide how far to decompose, and what do you hand the owner at the end?

level: principalimportance: should knowfreq 25%

basics

~20 s

Decompose only where a decision hangs on the answer: a stage someone wants to remove, or one that costs latency or money. Group the rest. Budget one end-to-end baseline plus an isolation run per stage you are actually questioning, and hand back unique-catch per stage against its cost, not four separate pass rates.

open as a page

A guardrail regression suite has grown to a few thousand cases, and a full run bills a moderation-guard call and a chat-model call for each one. How do you decide what to retire, and what makes retiring a case dangerous?

level: principalimportance: should knowfreq 34%

basics

~20 s

Cut redundancy before cutting evidence. Collapse near-identical paraphrase clusters to one representative plus an occasional sample, and drop cases whose defect class no longer exists in the product. Never retire on the grounds that a case has never failed: a case that never fails is doing its job, and it is often the only witness to one defect.

open as a page

To avoid paying for one full corpus run per layer, an engineer proposes scoring the attack corpus against each moderation layer offline and OR-ing the verdicts to predict what the deployed stack would do. Where does that reconstruction hold, and where does it break?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

It matches for layers that read the same fixed user text, like two input-side classifiers: OR-ing their independent verdicts predicts the stack. It breaks for anything downstream of generation, because the output classifier reads a reply that only exists if the request ran, varies between runs, and changes if an upstream layer rewrote the input.

open as a page