skip to content

Safety Classifiers

A guard that is itself a model reports only the hazard categories it was trained to emit, and costs an inference per call. Interviewers probe this because a clean sheet is often an untested one.

on this pageshow

explore

questions

5

A model-based safety classifier (a guard model such as Llama Guard or ShieldGemma) sits in front of your assistant, and your red-team run of 300 harmful prompts came back with zero blocks reported. What do you check before you write that the guard is ineffective, or that it is effective?

level: middleimportance: must knowfreq 70%

answer

  1. clean sheet = suspect the harness
  2. positive and benign controls
  3. fixed hazard taxonomy bounds measurement
  4. out of taxonomy = not measured, not passed
  5. log the guard's raw output

basics

~20 s

Rule out two things. First, that the harness really routed every prompt through the guard and parsed its verdict — send a known-blocked control and confirm it blocks. Second, that your harms map to categories the guard was trained to emit. A harm outside its fixed taxonomy reads clean because it was never measured, not because it is safe.

solid answer

~60 s

A uniform result — all clean or all blocked — is a harness result until proven otherwise. **Step one, prove the instrument works.** Include positive controls: a handful of prompts that unmistakably fall inside a category the guard is documented to cover. If those do not trip it, your harness is not calling the guard, is calling it with the wrong field, or is misreading the verdict. Also send obviously benign controls to confirm you are not simply reading a constant. **Step two, map your prompts to the guard's declared categories.** A guard that is a model reports only the fixed set of hazard categories it was trained to emit. Send it a harm outside that set — a business-specific policy breach, a domain your organisation cares about, a category the vendor never trained — and it returns clean. That is not a bypass and not a false negative in the usual sense; it is a measurement gap, and no threshold or prompt change closes it. Report the two causes differently: a broken harness invalidates the run, a taxonomy gap is a finding about what this layer can ever tell you.

go deeper

for a junior

Suspects the harness first and asks whether the guard was actually called and its verdict read correctly.

for a middle

Adds positive and benign controls, and knows that a guard model can only report the hazard categories it was trained to emit, so out-of-taxonomy harm reads clean.

for a senior

Separates the three causes in the write-up, labels prompts with intended categories up front, logs raw verdicts, and re-runs controls when the guard model changes.

for a principal

Insists the reported denominator is stated, so a leadership reader cannot mistake the guard's category list for the organisation's harm policy.

An empty result sheet is the most misread artefact in guardrail testing, because three completely different situations produce it and only one of them is about the guard's quality. ## How a verdict is produced, and where it goes missing A model-based guard receives the text under review inside a fixed prompt template and emits a short answer — typically a safe/unsafe token, then the codes of any hazard categories it thinks apply. Your harness has to (a) actually call it, (b) hand it the right field, and (c) parse that answer into the pass/fail you record. Each of those three steps fails silently. A harness that reads a `categories` key the provider does not return records "clean" for every case. An integration that fails open on timeout records "clean" for every case. A credential that expires forty minutes into a two-hour run records "clean" from that minute onward. None of these throws. ## The three explanations for zero blocks **1. The instrument was never wired in.** Positive controls settle this. A handful of prompts that unmistakably sit inside a category the guard is documented to cover should trip it; if they do not, stop and fix the harness before reading anything else. Pair them with obviously benign controls, which prove you are not reading a constant "unsafe" either. Controls are the cheapest insurance in the run: a dozen cases against a set of three hundred is roughly four per cent of the budget. Run them at the start *and* at the end, because the failure modes above are time-dependent. **2. The harms fall outside the guard's taxonomy.** This is the one this leaf exists for. A guard model emits verdicts over a hazard category list fixed at training time — a published, finite enumeration. A harm that does not correspond to any entry in that list produces no signal at all. Your organisation's policy on regulated advice, contractual conduct, brand safety or a domain-specific abuse pattern is very likely not in there. The guard returns safe, and it is telling the truth about the only question it was ever asked. The correct sentence in the report is "not measured by this layer" — never "passed", and not "false negative" either, because a false negative implies a signal that existed and was missed. **3. In taxonomy, but under the line.** Here the guard did have the category and simply scored low, or the deployment's blocking threshold sits above what it returned. This is a real finding about this configuration and it is fixable within the layer. ## Where the number misleads The dangerous reading is the substitution of denominators. The run measured "prompts blocked out of prompts sent, over the categories this model can express." The reader hears "harmful content blocked, over our harm policy." Every category your policy names and the guard cannot represent gets silently promoted from unmeasured to safe on the way from your table to a summary slide. A clean sheet is the strongest possible version of that error, because there is not even a partial number to argue with. The second misleading reading runs the other way: reporting "the guard blocked nothing, therefore it is ineffective" when in fact the harness never invoked it. That write-up is worse than useless — it burns credibility and prompts a remediation of a layer that was never tested. ## What you would check Label every prompt with its intended guard category *before* the run, using the guard's own published category names rather than your policy's names; any prompt you cannot label is a declared measurement gap on the way in, not a discovery during triage. Keep the guard's raw returned output next to the pass/fail your harness computed, so a reviewer can distinguish "the guard said nothing" from "the guard said something we discarded". Verify the guard's call count against the expected count. Check error, timeout and retry logs for a fail-open window. Confirm which text the integration actually hands the guard — the newest user turn, the whole thread, or a truncated window. And re-run the controls whenever the guard model or its serving version changes, because both the category list and its behaviour are properties of a specific model, not of the product name on the invoice.

  • What is a good positive control for a guard model?
    A small set of prompts that fall squarely inside a hazard category the guard is documented to cover, plus benign prompts to prove the verdict is not a constant. They exist to prove the wiring, not to find anything.
  • The guard returns a verdict for a category, but your harness recorded 'clean'. What happened?
    Usually a parsing or threshold-mapping bug in the harness: the raw verdict is there but the field, the label mapping or the comparison is wrong. Log the raw response so this is visible rather than inferred.
  • How do you phrase a harm the guard's taxonomy does not cover in the report?
    As untested at this layer, with a recommendation for where it could be measured — a second classifier, a rules layer, or human review — never as a pass.

A guard model is a smoke detector with a fixed sensor: silence can mean there is no fire, or that the battery is dead, or that what is filling the room is carbon monoxide, which this unit was never built to smell. Only the first is good news, and the three are indistinguishable from the panel.

saying these in an interview costs you the question

  • Reporting 'the guard blocked nothing, so it is broken' without a positive control.
  • Reporting 'no hits, so the deployment is safe' when the harms were never in the guard's category set.
  • Writing 'passed' for a harm category the instrument cannot express.
  • Trusting a mid-run silence when rate limits or timeouts could have failed open.

context

open as a page

Your guardrail test against a model-based safety classifier shows that 40 of 100 harmful prompts reached the assistant unblocked. How do you separate the misses whose harm lies outside the guard's fixed hazard taxonomy from the misses that were in taxonomy but scored too low to block, and why does that split change what you recommend?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Label each prompt with the guard category it should fall under before the run. Misses where the guard scored the right category just under the blocking line are threshold problems. Misses where no relevant category exists at all are taxonomy gaps, and no threshold change ever fixes those — they need a different or additional layer.

open as a page

You are red-teaming a chat endpoint that sits behind a safety classifier — a guard that is itself a model, such as Llama Guard or ShieldGemma, which reads text and returns a verdict. Why does each test case in that run cost more than testing the chat endpoint alone, and how should that shape the size of your test set?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Because the guard is itself a model, every case runs extra inferences: one to screen the input, usually another to screen the reply, on top of the chat call. Each case therefore costs roughly two to three times the tokens, money and latency. Budget for that and run a smaller, deliberately chosen set.

open as a page

You are measuring a guard that is itself a language model (a safety classifier fronting a chat product), not a deny-list or regex filter. Which classes of test case do you add because the guard is a model, and what does each one measure?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Add cases that probe judgement rather than matching: the same harm in many paraphrases, other languages, and unusual formatting; long or multi-turn inputs where only part of the text reaches the guard; and text whose framing addresses the guard's own reading of the content. A deny-list needs none of these because it has no judgement to sway.

open as a page

Your organisation's harm policy names categories that the deployed guard model's fixed hazard taxonomy does not cover, and the guard emits categories your policy never mentions. How do you scope and report a guardrail test engagement so that the pass rate you hand leadership is not read as evidence that the policy is enforced?

level: principalimportance: should knowfreq 32%

basics

~20 s

Map every policy category to a guard category first and mark the unmapped ones unmeasurable at this layer. Report two numbers, never one: prompts blocked out of prompts sent, and how many policy categories this instrument can express at all. Unmapped categories go in the report as untested, with an owner, not as passing.

open as a page