skip to content

The Layers Under Test

A guard model with a fixed taxonomy, a hosted filter with an unreadable policy and a deterministic rail bound what you can measure. Interviewers probe this because you cannot score an unnamed layer.

on this pageshow

explore

questions

15

A hosted content-moderation service (such as Azure AI Content Safety or OpenAI's moderation endpoint) returns every category below your threshold for a prompt your team considers harmful. As a red teamer, what does that clean verdict actually license you to conclude, and what does it not?

level: juniorimportance: must knowfreq 68%

answer

  1. fixed taxonomy, unpublished policy
  2. clean is not safe
  3. three buckets: flagged / under threshold / no category
  4. control set of known positives
  5. store raw scores, re-threshold offline

basics

~20 s

Only that nothing in the service's fixed category set scored above the threshold you applied. The vendor's policy and training data are not published, so a clean verdict is not evidence of safety: the harm may fall outside the categories on offer entirely. Record it as unmeasured, not as passed.

solid answer

~50 s

A hosted filter hands a tester a **fixed taxonomy** plus a numeric severity or confidence per category. What it does not hand you is the policy that decides which text lands in which category, the vendor's own internal cut-offs, or the training distribution. So a below-threshold result is a statement about that taxonomy, not about harm. That makes the miss **unattributable**: you cannot separate (a) a category exists and the classifier failed on this input, from (b) no category covers this harm at all, from (c) it scored just under the threshold *you* chose. Those three demand different write-ups — a defect, a coverage gap, and a tuning recommendation. Practically: store the raw per-category numbers rather than a boolean, keep a small control set of inputs the service reliably does flag so you can tell a real clean run from a broken integration, and bucket results three ways instead of pass/fail.

go deeper

for a junior

Says a below-threshold result only means no listed category scored high, and that it is not proof the content is safe.

for a middle

Separates a classifier miss from an out-of-taxonomy harm from a threshold choice, and keeps raw per-category values rather than booleans.

for a senior

Runs known-positive controls to validate the run itself, buckets misses before reporting them, and words each bucket as a different recommendation.

for a principal

Frames the hosted filter as one measured layer with a stated coverage boundary, and makes the residual out-of-taxonomy risk an explicit line in the risk register rather than a silent gap.

## What the instrument actually is A **hosted content filter** is somebody else's classifier sitting behind an HTTP endpoint. You post text — some services also accept images — and you get back a fixed list of named categories, each carrying a number. - OpenAI's moderation endpoint returns a `category_scores` object of floats, a parallel `categories` object of booleans, and a top-level `flagged`. - Azure AI Content Safety's text analysis returns a coarse `severity` band per harm category. - Perspective API returns a `summaryScore` per attribute it was trained on. Three vendors, three response shapes, one property they all share: **the category list is fixed by the vendor, and the policy that decides which text belongs in which category is not published.** You can observe the mapping's outputs; you cannot read the mapping. ## What "below threshold" is therefore a statement about Not about harm. It is a statement about one fixed taxonomy, evaluated by one unpublished decision rule, compared against a cut-off that in most deployments *you* or the integrating team chose rather than the vendor. Where a service does emit its own boolean, that boolean is computed at a default the vendor picked for a general population, not for your product's risk tolerance. So the honest sentence is: nothing in this service's category set scored above the line we applied, at this time, on this endpoint. ## Three buckets, not two Every result belongs in one of three places, and pass/fail collapses two of them into one lie. - **(1) Flagged** above your cut-off. - **(2) Scored, but under the cut-off** — a *threshold* outcome, owned by whoever deploys the filter, changeable this afternoon, and worth reporting with the actual value because "0.48 against a 0.5 line" and "0.02 against a 0.5 line" are not the same finding. - **(3) No category in the taxonomy plausibly covers the harm** you were probing at all — a *coverage* outcome, which no threshold change reaches, ever. Bucket three is not a defect in the classifier; it is the boundary of the product. ## Where the number misleads, and what that costs The single most common failure is reporting an **aggregate**: "the filter missed 40 of our 100 payloads." That reads as a 40% vendor defect rate and gets escalated as one. If twenty of those forty were harms the taxonomy never claimed to cover, the escalation collapses the moment the vendor answers, and the twenty legitimate misses go down with it — you have spent your credibility to deliver a number that was never a rate of anything. The **denominator** is the other half of the same trap: a miss rate over *your* corpus is not the service's false-negative rate. A false-negative rate needs a population, and the population that matters is the customer's real traffic mix, which your stratified adversarial corpus deliberately does not resemble. Quote it as "misses against our corpus, at our cut-off, during this window", never as a property of the service. ## What it costs to find out Every probe is a billed request against a rate-limited key, and the calls are the expensive, slow part while the analysis is free. That asymmetry dictates the discipline below. ## What you check before believing a clean run - Send a **control set** — a handful of inputs you have previously seen this service flag — inside the same session, with the same key, region and encoding. If the controls come back clean too, you have not measured a filter; you have measured a broken integration, a wrong region, a wrong deployment name, or a payload your client mangled before it left. - Repeat a few payloads to see whether values move at all between identical calls. - And persist the **full response body** keyed to the request and timestamped, because re-thresholding, re-stratifying or answering a reviewer's question then costs nothing, while a run that stored only booleans has to be bought a second time. ## What you must not claim Do not assert what the vendor's policy permits — you have never seen it. And do not try to reconstruct the decision surface by systematically harvesting scores: that is a **model-extraction** exercise, with its own contractual and legal footing, and it is not what a moderation test is for.

  • You have a clean run across your whole corpus. What do you check before you write it up?
    Whether known-positive controls sent in the same session still flag. If they do not, the integration, key, region or encoding is wrong and the run is void.
  • Why store the raw per-category values rather than the pass/fail your run computed?
    Because thresholds are yours, not the vendor's. Raw values let you re-threshold and re-analyse offline for free; booleans force another billed, rate-limited run.
  • Your customer asks for the filter's false-negative rate. What is the honest answer?
    You can only report a miss rate against your own corpus and your own threshold, over the observation window. There is no population-level rate to quote without knowing the vendor's policy or the real traffic mix.

A hosted filter is a metal detector at a door: walking through without a beep tells you that you are not carrying metal, not that you are unarmed. The categories are what the machine is tuned to notice, and nobody will show you the tuning.

saying these in an interview costs you the question

  • Reporting a below-threshold result as "the filter is safe" or "the prompt is benign".
  • Treating every miss as a vendor defect without asking whether any category covered that harm.
  • Storing only booleans, so any later threshold question requires paying for the whole run again.
  • Running with no known-positive control, then reporting a clean sweep that was actually a broken integration.
  • Asserting what the vendor's unpublished policy says based on scores collected from the endpoint.

context

open as a page

A probe you send at an application fronted by a rail framework such as NeMo Guardrails comes back as a short generic refusal. Why can that refusal text not tell you which rail blocked the turn, and what do you inspect instead?

level: juniorimportance: must knowfreq 58%

basics

~20 s

The refusal is a canned message the framework emits for any blocked turn, so an input pattern rule, an intent match, an LLM self-check or an output rail all produce the same text. To attribute the block you need the run's own execution trace: which rails ran, in what order, and each verdict.

open as a page

A model-based safety classifier (a guard model such as Llama Guard or ShieldGemma) sits in front of your assistant, and your red-team run of 300 harmful prompts came back with zero blocks reported. What do you check before you write that the guard is ineffective, or that it is effective?

level: middleimportance: must knowfreq 70%

basics

~20 s

Rule out two things. First, that the harness really routed every prompt through the guard and parsed its verdict — send a known-blocked control and confirm it blocks. Second, that your harms map to categories the guard was trained to emit. A harm outside its fixed taxonomy reads clean because it was never measured, not because it is safe.

open as a page

You are sizing a red-team probe run against a hosted content-moderation service that bills per request and rate-limits your API key. How does that shape the corpus and the run design, and which outcomes must never be scored as a clean result?

level: middleimportance: must knowfreq 58%

basics

~20 s

Every probe costs money and a slot in the rate limit, so build a small stratified corpus instead of a huge random one, deduplicate identical payloads and cache responses. Never score a rate-limit rejection, timeout or error as clean: those are missing verdicts, and counting them as not-blocked inflates your apparent bypass rate.

open as a page

Your guardrail test against a model-based safety classifier shows that 40 of 100 harmful prompts reached the assistant unblocked. How do you separate the misses whose harm lies outside the guard's fixed hazard taxonomy from the misses that were in taxonomy but scored too low to block, and why does that split change what you recommend?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Label each prompt with the guard category it should fall under before the run. Misses where the guard scored the right category just under the blocking line are threshold problems. Misses where no relevant category exists at all are taxonomy gaps, and no threshold change ever fixes those — they need a different or additional layer.

open as a page

Your engagement covers a chat product whose backend calls a hosted content-moderation service. Do you send your probe corpus straight to the vendor's moderation endpoint, or through the product? What does each choice actually measure?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Direct calls measure the vendor's classifier alone. Going through the product measures the deployment: which categories the team acts on, the cut-offs they chose, what text the service actually receives after the app transforms it, and what happens when the call fails. Only the second yields findings the customer can fix, so do both and report them separately.

open as a page

Three payloads get through an application protected by a rail stack: one matched no pattern rule, one was approved by the rail that asks a judge model, and one matched no canonical example so the turn reached the application model on the unmatched path. Why file three separate defects rather than one, and what does the fix look like for each?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Each has a different owner and a different fix. The unmatched pattern is a rule gap: widen or add a rule and it closes deterministically. The judge approval is a score problem: no edit guarantees that payload stays blocked, only the rate moves. The unmatched routing is worse still: that turn was never governed at all.

open as a page

You are red-teaming a chat endpoint that sits behind a safety classifier — a guard that is itself a model, such as Llama Guard or ShieldGemma, which reads text and returns a verdict. Why does each test case in that run cost more than testing the chat endpoint alone, and how should that shape the size of your test set?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Because the guard is itself a model, every case runs extra inferences: one to screen the input, usually another to screen the reply, on top of the chat call. Each case therefore costs roughly two to three times the tokens, money and latency. Budget for that and run a smaller, deliberately chosen set.

open as a page

A rail framework routes each user turn by embedding-similarity against a fixed set of canonical example utterances before deciding what to do with it. Why can a paraphrase that means the same thing flip the outcome, and what does that do to a red-team result you intend to rerun?

level: middleimportance: should knowfreq 44%

basics

~20 s

Matching is nearest-neighbour in embedding space against a fixed example set, with a similarity cutoff. A paraphrase can land nearer a different example, or under the cutoff and match nothing at all, so no governed path fires. The decision follows wording distance, not meaning, so one probe run is a sample, not a verdict.

open as a page

In a rail stack, one layer matches a fixed pattern rule while another asks a judge model whether the turn should be allowed. Why is a single pass from the judge-backed layer weaker evidence in a red-team report than a single pass from the pattern rule, and how do you strengthen it?

level: middleimportance: should knowfreq 46%

basics

~20 s

The pattern rule is deterministic: one trial fully characterises it for that input, and a miss is a rule gap you can point at. The judge layer samples a model, so its verdict moves with sampling, prompt wording and model version. Strengthen it by replaying the identical payload many times and reporting a rate.

open as a page

You are measuring a guard that is itself a language model (a safety classifier fronting a chat product), not a deny-list or regex filter. Which classes of test case do you add because the guard is a model, and what does each one measure?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Add cases that probe judgement rather than matching: the same harm in many paraphrases, other languages, and unusual formatting; long or multi-turn inputs where only part of the text reaches the guard; and text whose framing addresses the guard's own reading of the content. A deny-list needs none of these because it has no judgement to sway.

open as a page

Six weeks after you reported a payload that a hosted content-moderation service failed to flag, the customer says they cannot reproduce it, and the service exposes no version you can pin. How do you resolve the dispute, and how should the finding have been written to survive this?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Assume the service changed behind the endpoint; you cannot pin a version. Re-run the stored payload alongside a control set of cases that flagged during the original run. If the controls still flag and the payload now does too, the behaviour moved. Report that with timestamps instead of defending the old result.

open as a page

Your organisation's harm policy names categories that the deployed guard model's fixed hazard taxonomy does not cover, and the guard emits categories your policy never mentions. How do you scope and report a guardrail test engagement so that the pass rate you hand leadership is not read as evidence that the policy is enforced?

level: principalimportance: should knowfreq 32%

basics

~20 s

Map every policy category to a guard category first and mark the unmapped ones unmeasurable at this layer. Report two numbers, never one: prompts blocked out of prompts sent, and how many policy categories this instrument can express at all. Unmapped categories go in the report as untested, with an owner, not as passing.

open as a page

You lead a fixed-hours red-team engagement on a product that relies on a hosted content-moderation service. How much of the engagement do you spend probing a layer the customer cannot change, and what should the deliverable about it say?

level: principalimportance: should knowfreq 33%

basics

~20 s

Spend little on proving a vendor classifier imperfect; the customer cannot fix it. Spend the hours on the decisions they own: which categories they act on, the cut-offs, what text is screened, and what happens on error or timeout. The deliverable should recommend configuration, fallbacks and a residual-risk statement, not a vendor swap.

open as a page

You have to report guardrail coverage for an application protected by a rail stack that combines deterministic rules, similarity-matched routing and a judge-model self-check. The three layers share no denominator. How do you report coverage without inventing one number that misleads?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Report three numbers with their own denominators and refuse the single one. Rules exercised out of rules configured is countable. Routing is a continuous space with no enumerable denominator, so report a fall-through rate over paraphrase families. The judge layer is a pass rate at a fixed repeat count against a named judge model.

open as a page