skip to content

Your red-team report says a content-moderation guard let 3% of your attack prompts through. Which configuration facts have to sit next to that number before anyone else can reproduce or compare it?

level: juniorimportance: must knowfreq 68%

answer

  1. threshold travels with the number
  2. config snapshot, not a description
  3. which categories block
  4. who judged a bypass
  5. run date for a hosted guard

basics

~20 s

The block threshold you ran at, which categories were set to block, the guard's identity and configuration snapshot, and the date. A bypass percentage is only true at one operating point, so without those settings written down nobody can rerun your test or compare a later number to yours.

solid answer

~60 s

A bypass percentage is a reading taken at one setting of a dial, so the setting travels with the reading or the reading is worthless. What belongs in the header of the result: - **The operating point** — the numeric block threshold, or the named strictness level, in force during the run. - **Which categories were enabled to block**, since a guard that only blocks two of its categories is a different guard. - **Guard identity and configuration snapshot** — which moderation component, self-hosted or hosted, and the exact config you sent, captured verbatim rather than described. - **The corpus** — how many attack prompts, and which set, so the denominator is fixed. - **What decided a bypass** — the object that judged whether the response counted as a hit (a human, a rubric, an automated judge), because a different judge yields a different 3%. - **The run date**, since a hosted guard can be retuned by its operator without telling you. Without these, the next person's 3% and yours are not the same measurement.

go deeper

for a junior

Should name the threshold or strictness setting, the guard, and the corpus size as things that must be recorded with the number.

for a middle

Adds the enabled categories, a verbatim config snapshot rather than a prose description, and what decided a hit.

for a senior

Points out that a hosted guard cannot be snapshotted, so the run date and any returned version identifier become the reproducibility anchors, and says so in the report.

for a principal

Treats the result header as a contract: makes reproducible metadata a required field of every engagement deliverable so findings stay comparable across quarters and vendors.

## What a bypass percentage actually is A moderation guard is a classifier sitting in front of (or behind) a model. For each input it emits one or more continuous **scores** — usually a confidence per harm category — and a **threshold** turns those scores into a binary verdict: allow or block. Some products expose the threshold as a number, others as a named strictness level (`low`/`medium`/`high`, or a severity cut-off) that maps to a number internally; either way it is a dial. Your "3% bypassed" is the share of attack prompts that landed on the allow side of that dial *and* whose resulting model response a **judge** — a human reviewer, a written rubric, or an automated scorer — called a successful attack. Nothing about the guard's weights changed to produce 3%; the dial position produced it, jointly with which categories were wired to block at all. A guard exposing eight categories with three enforcing is operationally a different guard from the same product with all eight on: same vendor, same version, different measurement. ## The header the number travels in ```yaml result: bypass_rate: 0.03 guard: <which moderation component, hosted or self-hosted> guard_config_snapshot: sha256:... # the exact request config, verbatim operating_point: <threshold value, or the named strictness level> blocking_categories: [ ... ] # which ones were actually enforcing corpus: <name>, n=<count> # fixes the denominator decided_by: <human review | rubric | automated judge, named> run_started: <date> guard_version_seen: <any identifier the response returned> ``` Two of these are the ones people skip. The **snapshot** must be the bytes you sent, not a description: "we ran it on medium" survives contact with nobody, because `medium` can be redefined, and prose silently omits defaults you never set. The **judge** must be named because it decides the numerator: the guard decides what got through, the judge decides what counted, and two runs at an identical threshold diverge wildly when one batch was triaged by a human and the next scored by a model. ## What recording it costs, and what skipping it costs Recording costs minutes and a few kilobytes — it is bookkeeping. Skipping it costs a rerun. A 500-prompt corpus against a metered hosted guard is roughly 500 guard calls, 500 target-model completions to have something to judge, and 500 judge calls if the judge is a model, plus the engineer-days of re-triage that no API bill shows. That asymmetry is the whole argument: the header is a rounding error against the cost of the measurement it protects. ## Where the number misleads The dangerous reading is that 3% is a property of *the guard*. It is a property of guard + threshold + enforcing categories + corpus + judge, and each term has a way of moving silently. - **Vendor comparison.** Guard A at its strictest and guard B at its shipped default are not comparable, and a table with two percentages in it looks exactly as if they were. - **Trend lines.** A quarter-on-quarter chart showing 7% → 3% is an improvement only if the dial did not move between runs. If it did, you have plotted a config edit and labelled it hardening. - **Hosted drift.** For a hosted guard you cannot snapshot the model behind the endpoint at all. The operator can retune it with no notice and no version bump, so an unexplained change months later gets blamed on your corpus. The run date and any returned version identifier are the only anchors you have, and the report should say plainly that reproducibility depends on the operator. - **Denominator confusion.** 3% of an all-attack corpus is not 3% of production traffic; that restatement belongs in the base-rate discussion, and the header is what makes it possible at all. ## What to check before you file Take a 20–50 prompt slice, configure the guard **from the header alone** — not from the notebook still open on your screen, not from memory — and rerun it. If you do not land on the same figure within sampling noise, a field is missing, and in practice the missing field is the judge or an enforcing category you forgot was toggled. Then diff the recorded config against what the console shows today; on a shared environment somebody has usually changed something mid-engagement, and you would rather find that before the number is quoted than after.

  • You cannot snapshot the model behind a hosted moderation endpoint. What do you record instead?
    The exact request configuration you sent, any version or model identifier the response returns, and the run window. Then state in the report that reproducibility depends on the operator not retuning it.
  • Why record the judge that decided a bypass alongside the threshold?
    Because the threshold decides what the guard blocked, and the judge decides what counted as a successful attack. Change either and the same corpus gives a different percentage.
  • Two engineers run the same corpus against the same guard and get 3% and 11%. What do you look at first?
    Compare their recorded operating points and enabled categories, then their judges. One of those three almost always explains the gap before anything about the prompts does.

Reporting a bypass rate with no operating point is like reporting how many cars a speed camera flagged without saying what limit it was set to: the count is real, but it measures the setting as much as the traffic.

saying these in an interview costs you the question

  • Quoting a bypass percentage with no threshold or strictness setting recorded anywhere.
  • Describing the guard configuration in prose instead of capturing it verbatim.
  • Assuming a hosted moderation endpoint behaves identically months later.
  • Not recording what decided that a response counted as a bypass.
  • Treating the number as a property of the guard rather than of guard-plus-setting.

context