A model-based safety classifier (a guard model such as Llama Guard or ShieldGemma) sits in front of your assistant, and your red-team run of 300 harmful prompts came back with zero blocks reported. What do you check before you write that the guard is ineffective, or that it is effective?
answer
- clean sheet = suspect the harness
- positive and benign controls
- fixed hazard taxonomy bounds measurement
- out of taxonomy = not measured, not passed
- log the guard's raw output
basics
~20 sRule out two things. First, that the harness really routed every prompt through the guard and parsed its verdict — send a known-blocked control and confirm it blocks. Second, that your harms map to categories the guard was trained to emit. A harm outside its fixed taxonomy reads clean because it was never measured, not because it is safe.
solid answer
~60 sA uniform result — all clean or all blocked — is a harness result until proven otherwise. **Step one, prove the instrument works.** Include positive controls: a handful of prompts that unmistakably fall inside a category the guard is documented to cover. If those do not trip it, your harness is not calling the guard, is calling it with the wrong field, or is misreading the verdict. Also send obviously benign controls to confirm you are not simply reading a constant. **Step two, map your prompts to the guard's declared categories.** A guard that is a model reports only the fixed set of hazard categories it was trained to emit. Send it a harm outside that set — a business-specific policy breach, a domain your organisation cares about, a category the vendor never trained — and it returns clean. That is not a bypass and not a false negative in the usual sense; it is a measurement gap, and no threshold or prompt change closes it. Report the two causes differently: a broken harness invalidates the run, a taxonomy gap is a finding about what this layer can ever tell you.
go deeper
Suspects the harness first and asks whether the guard was actually called and its verdict read correctly.
Adds positive and benign controls, and knows that a guard model can only report the hazard categories it was trained to emit, so out-of-taxonomy harm reads clean.
Separates the three causes in the write-up, labels prompts with intended categories up front, logs raw verdicts, and re-runs controls when the guard model changes.
Insists the reported denominator is stated, so a leadership reader cannot mistake the guard's category list for the organisation's harm policy.
An empty result sheet is the most misread artefact in guardrail testing, because three completely different situations produce it and only one of them is about the guard's quality. ## How a verdict is produced, and where it goes missing A model-based guard receives the text under review inside a fixed prompt template and emits a short answer — typically a safe/unsafe token, then the codes of any hazard categories it thinks apply. Your harness has to (a) actually call it, (b) hand it the right field, and (c) parse that answer into the pass/fail you record. Each of those three steps fails silently. A harness that reads a `categories` key the provider does not return records "clean" for every case. An integration that fails open on timeout records "clean" for every case. A credential that expires forty minutes into a two-hour run records "clean" from that minute onward. None of these throws. ## The three explanations for zero blocks **1. The instrument was never wired in.** Positive controls settle this. A handful of prompts that unmistakably sit inside a category the guard is documented to cover should trip it; if they do not, stop and fix the harness before reading anything else. Pair them with obviously benign controls, which prove you are not reading a constant "unsafe" either. Controls are the cheapest insurance in the run: a dozen cases against a set of three hundred is roughly four per cent of the budget. Run them at the start *and* at the end, because the failure modes above are time-dependent. **2. The harms fall outside the guard's taxonomy.** This is the one this leaf exists for. A guard model emits verdicts over a hazard category list fixed at training time — a published, finite enumeration. A harm that does not correspond to any entry in that list produces no signal at all. Your organisation's policy on regulated advice, contractual conduct, brand safety or a domain-specific abuse pattern is very likely not in there. The guard returns safe, and it is telling the truth about the only question it was ever asked. The correct sentence in the report is "not measured by this layer" — never "passed", and not "false negative" either, because a false negative implies a signal that existed and was missed. **3. In taxonomy, but under the line.** Here the guard did have the category and simply scored low, or the deployment's blocking threshold sits above what it returned. This is a real finding about this configuration and it is fixable within the layer. ## Where the number misleads The dangerous reading is the substitution of denominators. The run measured "prompts blocked out of prompts sent, over the categories this model can express." The reader hears "harmful content blocked, over our harm policy." Every category your policy names and the guard cannot represent gets silently promoted from unmeasured to safe on the way from your table to a summary slide. A clean sheet is the strongest possible version of that error, because there is not even a partial number to argue with. The second misleading reading runs the other way: reporting "the guard blocked nothing, therefore it is ineffective" when in fact the harness never invoked it. That write-up is worse than useless — it burns credibility and prompts a remediation of a layer that was never tested. ## What you would check Label every prompt with its intended guard category *before* the run, using the guard's own published category names rather than your policy's names; any prompt you cannot label is a declared measurement gap on the way in, not a discovery during triage. Keep the guard's raw returned output next to the pass/fail your harness computed, so a reviewer can distinguish "the guard said nothing" from "the guard said something we discarded". Verify the guard's call count against the expected count. Check error, timeout and retry logs for a fail-open window. Confirm which text the integration actually hands the guard — the newest user turn, the whole thread, or a truncated window. And re-run the controls whenever the guard model or its serving version changes, because both the category list and its behaviour are properties of a specific model, not of the product name on the invoice.
- What is a good positive control for a guard model?A small set of prompts that fall squarely inside a hazard category the guard is documented to cover, plus benign prompts to prove the verdict is not a constant. They exist to prove the wiring, not to find anything.
- The guard returns a verdict for a category, but your harness recorded 'clean'. What happened?Usually a parsing or threshold-mapping bug in the harness: the raw verdict is there but the field, the label mapping or the comparison is wrong. Log the raw response so this is visible rather than inferred.
- How do you phrase a harm the guard's taxonomy does not cover in the report?As untested at this layer, with a recommendation for where it could be measured — a second classifier, a rules layer, or human review — never as a pass.
A guard model is a smoke detector with a fixed sensor: silence can mean there is no fire, or that the battery is dead, or that what is filling the room is carbon monoxide, which this unit was never built to smell. Only the first is good news, and the three are indistinguishable from the panel.
saying these in an interview costs you the question
- Reporting 'the guard blocked nothing, so it is broken' without a positive control.
- Reporting 'no hits, so the deployment is safe' when the harms were never in the guard's category set.
- Writing 'passed' for a harm category the instrument cannot express.
- Trusting a mid-run silence when rate limits or timeouts could have failed open.