When hand-labelling items for a guardrail test corpus, why must each label come from your written content policy rather than from what the guardrail itself returned, and what does that imply for how you choose benign items?
answer
- policy is the ground truth, not the verdict
- label blind, run the tool after
- selection by verdict hides the misses
- benign twin for each attack item
- disagreements are findings, both ways
basics
~20 sIf the classifier's own verdict becomes the label, it agrees with itself by construction and every error rate collapses to zero. The label must be an independent human judgement against the written policy. That also means benign items should be chosen near the policy line, not far from it, or the pass half proves nothing.
solid answer
~60 sLabelling from the tool's output is circular: the thing you are measuring becomes the ground truth, so disagreement is impossible and the corpus can never fail the guardrail. The same trap appears in softer forms — seeding the corpus by scraping only what the guardrail already blocked, or letting labellers see the verdict while they label. The fix is procedural. Write the policy first, in decidable terms. Label blind: the labeller sees the item and the policy, never the classifier's response. Run the classifier afterwards and treat every disagreement as a finding to inspect, because it is either a guardrail error or a bad label, and both are worth knowing. For the benign half this changes selection. If the policy is the authority, benign means "the policy says this must pass", not "the guardrail let it through". That pushes you toward items near the boundary — legitimate security questions, clinical detail, fiction, quoted abuse in a report — which are precisely the items where the policy has an opinion and the classifier may not.
go deeper
Sees that using the classifier's answer as the label makes it agree with itself and hides errors.
Labels blind against a written policy, runs the guardrail afterwards, and treats mismatches as things to inspect.
Also catches selection-by-verdict and correlated judge errors, writes decidable policy rules, and builds attack/benign twin pairs so the near-boundary behaviour is measurable.
Owns the policy as a versioned artefact, requires labels to be traceable to a policy clause, and uses recurring ambiguous items to drive policy revision rather than one-off labeller calls.
**Why independence is the whole point.** A test corpus produces information because two things are compared: a label, which is a human judgement about what *should* happen, and a verdict, which is what the system under test *did*. Take the label from the verdict and you no longer have two sources — you have one, compared with itself. Every measured error rate then collapses toward zero, and what you have measured is the guard's self-consistency, which is near perfect and worth nothing. The trap survives review because the pipeline sounds sensible when described out loud: run traffic through the guard, keep what it flagged, call that the attack half. **Three softer versions of the same mistake.** - *Selection by verdict.* Sourcing attack items from what the guard already flagged builds a corpus of things it already catches. Its misses are excluded by construction, so the catch rate marches toward one and carries no information about the failures that matter. - *Verdict-visible labelling.* Showing a labeller the guard's `unsafe`/`safe` output, or Azure AI Content Safety's per-category `severity`, while they label anchors them, and the anchoring runs toward agreement. Hide the verdict; reveal it only at adjudication time. - *Correlated judge.* Pre-labelling with an LLM judge is legitimate triage if a human confirms every final label. It stops being legitimate when the judge is the guard under test or a close relative — same family, same safety-tuning data. Then the judge's errors correlate with the guard's, and correlated errors are indistinguishable from agreement. **What a labellable policy looks like.** "No harmful content" cannot be labelled; two competent people will split on it. A labellable clause names the category, the readable-from-text intent condition, and the disposition — for example, that clinical dosing information asked for care is allowed while the same information framed as a request to cause harm is not. Decidability *from the item text alone* is the bar, because item text is all the guard sees. Where intent genuinely cannot be read from the text, that is a policy gap, and the honest move is to park the item in an `ambiguous` pile and count it, not to force a call. **Consequences for benign selection.** If the policy is the authority, "benign" means the policy says this must pass, not the guard let it through. That pushes selection toward the boundary rather than away from it. The sharpest construction is *twin pairs*: for each attack item, write the closest possible benign counterpart — same topic, same vocabulary, intent or framing changed just enough to flip the disposition. Paired items tell you directly whether the guard keys on topic words or on the thing the policy actually cares about, and a guard that blocks both halves of most pairs is a keyword filter wearing a classifier's API. **What it costs.** Blind double-labelling is the expensive part and the part teams cut first. A workable shape is: single-label the bulk, double-label a 100-item random sample blind, compute agreement, and adjudicate the disagreements with the policy open. Budget an hour or two of adjudication per hundred sampled items, plus the rewrite loop when adjudication exposes an undecidable clause. The recurring cost that gets forgotten is re-labelling: labels belong to a policy *version*, so when the policy changes, the affected slice of the corpus must be relabelled or explicitly carried forward with a note. Version the corpus alongside the policy and store the policy clause id on every item, or that cost becomes a full rebuild later. **Where the numbers mislead.** A verdict-labelled corpus reports error rates near zero — the most flattering possible result from the most broken possible construction. A judge-pre-labelled corpus with an unconfirmed lineage reports inflated agreement for the same reason. Subtler: high inter-labeller agreement is not evidence of correctness. Two labellers who share an intuition and never open the policy will agree with each other and drift from the written rule together; agreement measures consistency, correctness needs traceability to a clause. Subtlest: silently dropping the ambiguous pile makes every rate look better, because the discarded items are exactly the hard ones — report the ambiguous count next to the rates. **What to check.** Sample the corpus and confirm each label cites a policy clause. Confirm the labelling interface never displayed the guard's output. Pull every label-versus-verdict disagreement after the run and triage it into guard error or label error — both are findings, and the label errors point at policy sentences to rewrite. Check the double-labelled agreement figure and the size of the ambiguous pile; if agreement is low or the pile is large, the corpus is not ready and the fix is upstream, in the policy.
- You seeded the attack half only from prompts the guardrail flagged in production. What is the resulting catch rate worth?Almost nothing. Its misses are excluded by construction, so the rate is close to one regardless of real coverage; you need items sourced independently of the verdict.
- Is it ever acceptable to use a model to help label?As triage, yes — to pre-sort and surface candidates — provided a human confirms the final label and the helper model is not the guardrail under test or a close relative of it.
saying these in an interview costs you the question
- Building the attack half from whatever the guardrail already flagged.
- Showing labellers the classifier's decision while they label.
- A policy so vague that two labellers cannot reach the same answer from the text.
- Pre-labelling with a judge model that shares a lineage with the guardrail under test.
- Discarding every disagreement as a label error without inspecting it.