A teammate reports "our input-moderation guard has a 3% bypass rate" from a red-team harness run. Before you repeat that number in a report, what must the number state about what was counted?
answer
- numerator: what counted as a bypass
- who decided - regex, judge, human
- denominator: attempts vs distinct attacks
- retries move the rate, not the guard
- quote counts X of Y, not just a percent
basics
~20 sIt must say what counted as a bypass in the numerator: the guard failed to flag, or the model then actually complied - two different numbers. It must say whether the denominator is attempts, distinct attacks, or attempts per attack. And it must say how many retries each attack got, because retries move the rate without the guard changing.
solid answer
~50 sThree things, and none of them is optional. 1. **Numerator definition.** "Bypass" is ambiguous: it may mean the guard's classifier did not raise a flag, or that the flag was raised but the response still contained the harmful content, or that the model behind the guard actually complied. Also: who decided - a keyword match, a separate judge model, a human? Different deciders give different counts on the same transcripts. 2. **Denominator.** Attempts, distinct attack templates, or attacks-within-a-fixed-attempt-budget. The same run yields three different percentages. 3. **Retry policy.** If each attack was tried 50 times and only one attempt in fifty landed, a per-attempt rate of 2% and a per-attack rate of 100% describe the same data. A rate whose denominator the author cannot state is not evidence; it is a number. Report counts (X of Y) next to the percentage so the reader can re-derive it.
go deeper
Should say a percentage is meaningless without knowing what was counted on top and on the bottom, and should ask for the raw counts.
Should distinguish "the classifier did not flag it" from "the model actually complied", and note that retries change a per-attempt rate.
Should push on decider noise, dropped errored calls, and whether a one-in-fifty hit is a real bypass or sampling luck.
Should insist on a written measurement definition that every team quoting the guard's numbers uses, so two teams' rates are the same statistic.
### What the number is made of A bypass rate is one fraction: events somebody decided to call a bypass, over a population somebody decided to call the denominator. Neither half is handed to you by the tool. Both are choices, usually made silently while wiring the harness, and the string "3%" carries neither of them. So the first move on hearing a reported rate is not scepticism about the guard - it is a request for the two definitions. ### The numerator: four different events wear the word "bypass" A *guard* here is the classifier or service that screens a request before the model sees it, or screens the reply before the user does: a hosted one such as OpenAI Moderation or Azure AI Content Safety, or a self-hosted one such as Llama Guard or ShieldGemma. Each returns a different shape, and each shape supports a different notion of "got past". | Reading | What actually happened | Who can observe it | |---|---|---| | Classifier miss | Llama Guard returned `safe`; the omni-moderation response had `flagged: false` | any harness that calls the guard | | Sub-threshold detection | Azure AI Content Safety returned severity 2 for a category your app blocks at 4 - detected, allowed | a harness that logs the score, not just the boolean | | End-to-end compliance | the guard passed the input **and** the model then produced the disallowed content | a harness that also calls the target model | | Judged harm | a scorer - a garak detector, a PyRIT scorer, a promptfoo grader - labelled the final response harmful | a harness with a judge stage | The reading a stakeholder assumes when they hear "bypass" is the third or fourth. The reading most harness output actually gives you is the first, because a harness wired only to the guard endpoint has nothing else to look at. Naming the *decider* is therefore part of defining the numerator: a keyword regex, a hosted classifier, a judge model, or a human triage pass will disagree with each other on the same transcripts, and that disagreement lands straight in the headline percentage. ### The denominator: three populations out of one run - **Attempts** - every request sent. This is the default in tool output because attempts are the unit the loop iterates over: garak's `--generations` flag sets how many completions are drawn per prompt, so its attempt count is prompts times generations; promptfoo's `--repeat` re-runs each test case; a PyRIT orchestrator sends each seed prompt as many times as configured. Retry policy and duplicated templates move this number with the guard untouched. - **Distinct attacks** - one row per attack template, marked bypassed if any attempt landed. This is usually the defender's question, because an attacker needs exactly one technique that works and may retry it freely. - **Effort-weighted attacks** - bypassed within a fixed budget, say within 10 attempts. It restores the cost dimension the other two lack. ### What the run costs Do the multiplication before promising anyone a number. Two hundred templates at twenty attempts each is 4,000 guard calls; if the harness also calls the target model and then a judge model, that is roughly 12,000 requests for one figure. The judge stage usually dominates the bill, because it sends a whole transcript to a capable model rather than a short prompt to a cheap classifier. At the request rates a shared API key survives, and with hosted guards throttling, this is hours of wall-clock rather than minutes. Add the human cost: someone has to hand-confirm a sample of the hits, and the unconfirmed ones are the ones that embarrass you in the read-out. ### Where the number misleads Four failures account for most bad bypass rates. **Retry inflation:** an attack tried fifty times that lands once yields 2% per attempt and 100% per attack from identical data - both true, neither self-explanatory. **Silent denominators:** errored and timed-out calls counted as non-bypasses deflate the rate, while dropping them from the denominator entirely inflates it, and the harness summary rarely says which happened. **Scope creep:** a rate measured on an attack-only corpus is not the share of production traffic that evades moderation, because production traffic is overwhelmingly benign. **Decider noise:** if the judge mislabels even a few percent of transcripts, a 3% headline sits inside its own error bar. ### What to check before repeating it Ask for the per-template breakdown rather than the total; the attempts-per-template budget; the decider and its version; and how errored calls were handled. Then apply the sentence test - write the claim out in full words: ``` of 200 distinct attack templates, 6 got past the input guard at least once within a 20-attempt budget; 8 of 4,000 attempts landed overall ``` If that sentence cannot be written from what you were handed, the rate is not reportable yet, however confidently it was said.
- The harness only wraps the input classifier and never sees the model's reply. What can its bypass rate legitimately claim?Only that the classifier failed to flag the input. It cannot claim the model would have complied - the input may be one the model refuses anyway.
- Why report counts alongside the percentage?Counts let a reader re-derive any other denominator and see the sample size. "6 of 200 templates" and "6 of 4,000 attempts" are both 6, and the percentages are wildly different.
- A run shows 3% per attempt and 100% per distinct attack. Is that a contradiction?No. Every attack works occasionally and none works reliably. Both numbers are true; they answer different questions, and you must report both.
Saying a guard has a 3% bypass rate without naming the denominator is like saying a goalkeeper conceded three percent: three percent of shots faced, of matches played, or of the shots one striker kept re-taking until one finally went in? Same save record, three very different numbers.
saying these in an interview costs you the question
- Treating the percentage as self-explanatory and putting it in a report unchanged.
- Assuming bypass always means the model produced harmful output, when the harness only observed the classifier.
- Ignoring retries: quoting a per-attempt rate from a run where some attacks were retried far more than others.
- Dropping errored or timed-out calls from the denominator without saying so.