skip to content

Your guardrail test against a model-based safety classifier shows that 40 of 100 harmful prompts reached the assistant unblocked. How do you separate the misses whose harm lies outside the guard's fixed hazard taxonomy from the misses that were in taxonomy but scored too low to block, and why does that split change what you recommend?

level: seniorimportance: must knowfreq 50%

answer

  1. label intended category before the run
  2. near-miss score vs near-zero score vs no category
  3. threshold fixes only in-taxonomy misses
  4. paraphrase before declaring a gap
  5. keep the raw verdict for triage

basics

~20 s

Label each prompt with the guard category it should fall under before the run. Misses where the guard scored the right category just under the blocking line are threshold problems. Misses where no relevant category exists at all are taxonomy gaps, and no threshold change ever fixes those — they need a different or additional layer.

solid answer

~60 s

Do the labelling up front, not during triage. For each prompt in the set, record which of the guard's declared hazard categories it is supposed to trip; if you cannot name one, mark it unmappable before you send it. Then read the misses against those labels: - **Right category present, score below the blocking line** — the guard perceived the harm and the deployment chose not to act. Fixable inside the layer: threshold, guard model choice, or screening the response as well as the input. - **Right category present, score near zero** — the guard has the category but does not recognise this presentation of it. A model problem, often surface-form sensitive; test paraphrases before you conclude. - **No mappable category** — a measurement gap. The instrument cannot express this harm, so tuning it is wasted work; the recommendation is another layer, a bespoke classifier, or a non-model control. The split matters because the three groups have different owners and different costs, and lumping them into one "40% miss rate" invites the wrong fix.

go deeper

for a junior

Notices that a miss rate alone does not say why the prompts got through and asks to see the guard's per-case output.

for a middle

Distinguishes an in-taxonomy low score from a harm the guard has no category for, and knows only the first responds to threshold changes.

for a senior

Designs the prompt set with intended-category labels up front, buckets misses with evidence, tests paraphrases before calling a gap, and matches each bucket to a different fix and owner.

for a principal

Uses the decomposition to decide layer architecture — what this classifier will ever be able to cover versus what needs a separate control — rather than iterating on one guard's configuration.

A headline of "40 of 100 got through" is close to information-free. Everything useful is in the decomposition, and the decomposition has to be designed before the run, not reconstructed during triage. ## The mechanism you are reading against Guards report their confidence in two very different shapes, and the shape decides what "threshold" even means. | guard shape | what it returns | what a threshold is | |---|---|---| | generative guard (Llama Guard, ShieldGemma, Granite Guardian) | a safe/unsafe token plus category codes | there is no built-in dial; a score exists only if you take the probability of the unsafe token, which requires logprob access | | hosted scoring service (moderation endpoints, content-safety APIs) | per-category numeric scores or severity levels | a real, tunable comparison in your integration | That table is the first thing to establish, because a recommendation to "lower the threshold" against a generative guard that returns a bare token is a recommendation nobody can implement. If you want the near-miss evidence from such a guard you must capture token probabilities during the run — retrofitting them means paying for the whole run again. ## Buckets, and the evidence that assigns them Build the prompt set as a table before sending anything: prompt, intended harm, intended guard category, expected verdict. A row with an empty intended-category cell is a declared measurement gap on the way in. Then every miss lands in exactly one bucket, on evidence: - **Below the line.** The correct category is present and the score sits close to the blocking point. The guard perceived the harm; the deployment chose not to act. Fixable inside the layer. - **Unrecognised presentation.** The correct category is present but the score is near zero, while sibling paraphrases of the same harm do trip it. A model-behaviour finding, usually surface-form sensitivity. - **Taxonomy gap.** No category in the guard's published set corresponds to the harm at all. Nothing in the layer can be tuned to catch it. - **Harness artefact.** The guard was not called, timed out, or its verdict was mis-parsed. Those cases are invalidated and re-run, not counted as misses. ## What the split costs Paraphrase clusters are what separate bucket two from bucket three, and they are the expensive part: eight wordings per harm, at three inferences per case with two-sided screening, is twenty-four inferences per harm rather than three. Budget the clusters only for categories where a miss actually appeared — you do not need to paraphrase the things that already blocked. Capturing logprobs, where available, is nearly free and is what lets you distinguish "just under" from "not seen at all" without any extra calls at all. ## Where the number misleads **The composite rate.** "40% miss rate" invites the cheapest-looking single remedy — lower the blocking line. That trade buys false positives across *every* category simultaneously, and it buys exactly nothing for the taxonomy-gap bucket, where no score exists to cross any line. The opposite error is just as expensive: recommending an entire additional classification layer when most misses were near-misses in a category the current guard already covers. **The set as a confound.** If thirty of the forty misses come from one generator template or one harm family, the percentage is a property of your prompt set, not of the guard. Report per-category counts under the headline and cap how many prompts come from a single family. **The single wording.** A category declared uncovered on the strength of one phrasing is the most common mis-filing in this work. Send variants before you call a gap; a category that trips on three of five wordings is a recognition problem, and it has a different owner and a different fix than an absent category. ## What you would check That your intended-category labels are drawn from the guard's own published category list rather than your policy's names, and that the two are never compared as one scale. That controls in each category fired. That every miss has raw guard output attached to it. That no bucket is dominated by one prompt family. That you have stated the denominator for each bucket separately. And that any threshold recommendation names the guard shape it applies to, plus the over-blocking cost you measured on a benign set — a threshold move with no false-positive number beside it is an untested product change.

  • A miss shows the correct category with a score just under the blocking point. Is that a guard defect?
    It is a deployment decision, not a model defect: the signal existed and the configured line let it through. The fix and its false-positive cost belong to whoever owns the threshold.
  • How do you keep the miss percentage from being an artefact of your prompt set?
    Balance the set across the categories in scope, cap how many prompts come from one family or generator, and report per-category counts alongside the total rather than a single headline rate.

saying these in an interview costs you the question

  • Reporting a single miss rate with no decomposition.
  • Recommending a lower blocking threshold for harms the guard has no category for.
  • Declaring a taxonomy gap from one wording without trying paraphrases.
  • Mapping prompts to the organisation's policy names rather than the guard's own categories, then comparing the two as if they were one scale.

context