skip to content

A hosted content-moderation service (such as Azure AI Content Safety or OpenAI's moderation endpoint) returns every category below your threshold for a prompt your team considers harmful. As a red teamer, what does that clean verdict actually license you to conclude, and what does it not?

level: juniorimportance: must knowfreq 68%

answer

  1. fixed taxonomy, unpublished policy
  2. clean is not safe
  3. three buckets: flagged / under threshold / no category
  4. control set of known positives
  5. store raw scores, re-threshold offline

basics

~20 s

Only that nothing in the service's fixed category set scored above the threshold you applied. The vendor's policy and training data are not published, so a clean verdict is not evidence of safety: the harm may fall outside the categories on offer entirely. Record it as unmeasured, not as passed.

solid answer

~50 s

A hosted filter hands a tester a **fixed taxonomy** plus a numeric severity or confidence per category. What it does not hand you is the policy that decides which text lands in which category, the vendor's own internal cut-offs, or the training distribution. So a below-threshold result is a statement about that taxonomy, not about harm. That makes the miss **unattributable**: you cannot separate (a) a category exists and the classifier failed on this input, from (b) no category covers this harm at all, from (c) it scored just under the threshold *you* chose. Those three demand different write-ups — a defect, a coverage gap, and a tuning recommendation. Practically: store the raw per-category numbers rather than a boolean, keep a small control set of inputs the service reliably does flag so you can tell a real clean run from a broken integration, and bucket results three ways instead of pass/fail.

go deeper

for a junior

Says a below-threshold result only means no listed category scored high, and that it is not proof the content is safe.

for a middle

Separates a classifier miss from an out-of-taxonomy harm from a threshold choice, and keeps raw per-category values rather than booleans.

for a senior

Runs known-positive controls to validate the run itself, buckets misses before reporting them, and words each bucket as a different recommendation.

for a principal

Frames the hosted filter as one measured layer with a stated coverage boundary, and makes the residual out-of-taxonomy risk an explicit line in the risk register rather than a silent gap.

## What the instrument actually is A **hosted content filter** is somebody else's classifier sitting behind an HTTP endpoint. You post text — some services also accept images — and you get back a fixed list of named categories, each carrying a number. - OpenAI's moderation endpoint returns a `category_scores` object of floats, a parallel `categories` object of booleans, and a top-level `flagged`. - Azure AI Content Safety's text analysis returns a coarse `severity` band per harm category. - Perspective API returns a `summaryScore` per attribute it was trained on. Three vendors, three response shapes, one property they all share: **the category list is fixed by the vendor, and the policy that decides which text belongs in which category is not published.** You can observe the mapping's outputs; you cannot read the mapping. ## What "below threshold" is therefore a statement about Not about harm. It is a statement about one fixed taxonomy, evaluated by one unpublished decision rule, compared against a cut-off that in most deployments *you* or the integrating team chose rather than the vendor. Where a service does emit its own boolean, that boolean is computed at a default the vendor picked for a general population, not for your product's risk tolerance. So the honest sentence is: nothing in this service's category set scored above the line we applied, at this time, on this endpoint. ## Three buckets, not two Every result belongs in one of three places, and pass/fail collapses two of them into one lie. - **(1) Flagged** above your cut-off. - **(2) Scored, but under the cut-off** — a *threshold* outcome, owned by whoever deploys the filter, changeable this afternoon, and worth reporting with the actual value because "0.48 against a 0.5 line" and "0.02 against a 0.5 line" are not the same finding. - **(3) No category in the taxonomy plausibly covers the harm** you were probing at all — a *coverage* outcome, which no threshold change reaches, ever. Bucket three is not a defect in the classifier; it is the boundary of the product. ## Where the number misleads, and what that costs The single most common failure is reporting an **aggregate**: "the filter missed 40 of our 100 payloads." That reads as a 40% vendor defect rate and gets escalated as one. If twenty of those forty were harms the taxonomy never claimed to cover, the escalation collapses the moment the vendor answers, and the twenty legitimate misses go down with it — you have spent your credibility to deliver a number that was never a rate of anything. The **denominator** is the other half of the same trap: a miss rate over *your* corpus is not the service's false-negative rate. A false-negative rate needs a population, and the population that matters is the customer's real traffic mix, which your stratified adversarial corpus deliberately does not resemble. Quote it as "misses against our corpus, at our cut-off, during this window", never as a property of the service. ## What it costs to find out Every probe is a billed request against a rate-limited key, and the calls are the expensive, slow part while the analysis is free. That asymmetry dictates the discipline below. ## What you check before believing a clean run - Send a **control set** — a handful of inputs you have previously seen this service flag — inside the same session, with the same key, region and encoding. If the controls come back clean too, you have not measured a filter; you have measured a broken integration, a wrong region, a wrong deployment name, or a payload your client mangled before it left. - Repeat a few payloads to see whether values move at all between identical calls. - And persist the **full response body** keyed to the request and timestamped, because re-thresholding, re-stratifying or answering a reviewer's question then costs nothing, while a run that stored only booleans has to be bought a second time. ## What you must not claim Do not assert what the vendor's policy permits — you have never seen it. And do not try to reconstruct the decision surface by systematically harvesting scores: that is a **model-extraction** exercise, with its own contractual and legal footing, and it is not what a moderation test is for.

  • You have a clean run across your whole corpus. What do you check before you write it up?
    Whether known-positive controls sent in the same session still flag. If they do not, the integration, key, region or encoding is wrong and the run is void.
  • Why store the raw per-category values rather than the pass/fail your run computed?
    Because thresholds are yours, not the vendor's. Raw values let you re-threshold and re-analyse offline for free; booleans force another billed, rate-limited run.
  • Your customer asks for the filter's false-negative rate. What is the honest answer?
    You can only report a miss rate against your own corpus and your own threshold, over the observation window. There is no population-level rate to quote without knowing the vendor's policy or the real traffic mix.

A hosted filter is a metal detector at a door: walking through without a beep tells you that you are not carrying metal, not that you are unarmed. The categories are what the machine is tuned to notice, and nobody will show you the tuning.

saying these in an interview costs you the question

  • Reporting a below-threshold result as "the filter is safe" or "the prompt is benign".
  • Treating every miss as a vendor defect without asking whether any category covered that harm.
  • Storing only booleans, so any later threshold question requires paying for the whole run again.
  • Running with no known-positive control, then reporting a clean sweep that was actually a broken integration.
  • Asserting what the vendor's unpublished policy says based on scores collected from the endpoint.

context