You are measuring a guard that is itself a language model (a safety classifier fronting a chat product), not a deny-list or regex filter. Which classes of test case do you add because the guard is a model, and what does each one measure?
answer
- paraphrase clusters, not single wordings
- language and formatting = generalisation
- what text does the guard actually see
- guard reads untrusted text too
- benign set for over-blocking
basics
~20 sAdd cases that probe judgement rather than matching: the same harm in many paraphrases, other languages, and unusual formatting; long or multi-turn inputs where only part of the text reaches the guard; and text whose framing addresses the guard's own reading of the content. A deny-list needs none of these because it has no judgement to sway.
solid answer
~1 minAgainst a pattern filter you test strings. Against a model guard you test decisions, so the set changes shape: - **Paraphrase clusters.** Five to ten wordings of one harm. A model guard's verdict varies with surface form, so a single wording proves nothing either way, and the cluster gives you a per-category recognition rate instead of a coin flip. - **Language and register variation.** Guard models are usually strongest in their dominant training language and on plain prose; the same harm in another language or in a stylised register measures generalisation, which a regex simply does not have. - **Window and truncation cases.** Establish what text the guard actually receives — the last user message, the whole thread, a truncated window — then place the harmful content where the window may not reach. This is a wiring property of your deployment as much as of the guard. - **Framing that speaks to the guard.** The guard reads the same untrusted text as the model, so content whose framing is aimed at how the guard reads it measures whether verdicts can be influenced by the input. - **Benign-but-sensitive cases.** Model guards over-block adjacent, legitimate topics; without these you report only half the behaviour.
go deeper
Knows a model guard can be tested with rewordings and other languages, unlike a fixed pattern list.
Groups prompts into paraphrase clusters, includes a benign set for over-blocking, and asks what text the integration actually sends to the guard.
Builds the set around the judgement properties — recognition stability, generalisation, window boundaries, verdict influence, false positives — and reports each separately with the evidence.
Decides which of those properties the organisation will monitor continuously and which belong to a one-off engagement, and what an acceptable over-block rate is for the product.
The case list changes because the thing under test changed category. A rule engine has no interpretation to sway: it either contains the string or it does not, and there is no such thing as an unlucky phrasing. A model guard produces a *judgement* over text it was handed, so the properties worth measuring are properties of a judgement — stability, generalisation, what evidence it was given, and how often it is wrong in the expensive direction. Each class below measures exactly one of them, and each should be reported on its own rather than folded into a single block rate. ## Recognition stability — paraphrase clusters Group five to ten wordings of one intended harm and count how many trip the guard. A group that trips 8/10 and a group that trips 1/10 are entirely different postures; averaged together they produce a number that describes neither. Clusters also protect you from the opposite error, which is declaring a category uncovered on the strength of one unlucky phrasing. The output is a per-category recognition rate, which is a defensible statistic; a single binary per category is not. ## Generalisation — language, register, formatting Same intent, different presentation: another language, a transcript or code-block layout, heavy markup, an unusual register. Guard models are strongest in their dominant training language and on plain prose. You are not building a bypass here — you are measuring whether coverage survives presentation variation that ordinary users produce anyway, and multilingual product surfaces make this a first-class risk rather than an edge case. ## What the guard is actually shown — window and truncation Read the integration, then test what you read. Establish whether the guard receives the newest user message, the whole conversation, or a truncated window, and whether the model's reply is screened at all. If only the latest turn goes to the guard, then harm accumulated across turns is out of this layer's scope *by construction* — that is an architecture finding, not a model defect, and it is filed against the integration with an estimate of what screening the full window would cost, since that cost grows with conversation length on every turn. ## Verdict influence The guard consumes the same untrusted text the assistant does, which means its judgement is in principle addressable by that text. Measuring whether verdicts are stable under content whose framing speaks to the guard's own reading is a legitimate defender-side property of the layer. Report it at the level of the class of case and the verdict change you observed — the finding is the instability, never a reusable string. ## False positives — the benign set Ship a benign set of comparable size, and make it hard: medical, legal, security-education and other sensitive-but-legitimate topics, which is precisely where model guards over-block. A test that counts only misses will cheerfully recommend a more aggressive setting, and over-blocking is a product incident with its own support cost. ## What it costs This is the expensive shape of test set. A regex suite of 1,000 strings runs for free; 1,000 model-guard cases with two-sided screening is roughly 3,000 inferences, and clusters multiply that by the cluster size. Budget deliberately: clusters only for categories in scope, a benign set sized to give a usable false-positive rate rather than matched one-for-one, and language variants for the languages the product actually serves. Triage, not generation, is the binding constraint — every case that trips needs a human read. ## Where the number misleads The seductive misreading is that a high block rate on this set means a safe product. It does not, for three reasons stacked on top of each other: the set covers only categories the guard can express; the block rate says nothing about the benign rate beside it; and a rate computed over clusters weights whichever harms you happened to paraphrase most. Report per-category recognition rates, the over-block rate, and the cluster sizes that produced them, or the headline will be quoted without any of that. ## What you would check That clusters are balanced across categories rather than concentrated where prompts were easy to write. That the benign set is genuinely non-trivial and its rate is reported next to the block rate. That you can state, per case, the exact text the integration handed the guard. That language variants match the deployed locales. And that any verdict-influence finding is re-tested after a guard model or serving-version change, since it is a property of that specific model rather than of the product name.
- Why are paraphrase clusters worth the extra inference cost against a model guard?Because the verdict is surface-form sensitive: a single wording gives a binary you cannot trust, while a cluster yields a recognition rate per category and stops you mislabelling a bad phrasing as a coverage gap.
- The guard only screens the newest user turn. Is a harm built up across turns a guard defect?No — it is an integration finding. The layer was never given the evidence. Report it against the architecture and say what screening the full window would cost.
saying these in an interview costs you the question
- One prompt per harm, then a confident per-category verdict.
- Never establishing which text the integration hands to the guard.
- Measuring misses only, with no benign set, and then recommending a more aggressive setting.
- Reporting a reusable bypass string instead of the class of case and the property it demonstrates.