A guardrail test harness fired 4,000 prompt attempts, drawn from 200 distinct attack templates, at an input-moderation classifier; 400 attempts reached the model. Why is "the guard has a 10% bypass rate" an incomplete claim, and which denominators would you report instead?
answer
- 400/4000 measures retries, not the guard
- per distinct attack: X of 200
- attacker needs one working technique
- concentration across templates
- quote the attempt budget with the rate
basics
~20 s10% is 400 over attempts, so it partly measures how the corpus was sampled and retried rather than the guard. Report it beside the per-attack rate: how many of the 200 distinct templates got past at least once. If 180 bypassed, the guard is far weaker than 10% suggests; if 4 bypassed 100 times each, far stronger.
solid answer
~50 sThe per-attempt rate has the wrong denominator for the question people ask of it. **Per attempt (400/4,000 = 10%)** answers "if I fire one prompt from this corpus, how likely is it to land?" That is a property of the corpus's composition and retry policy at least as much as of the guard. Duplicate a successful template ten times and the rate rises with no change to the guard. **Per distinct attack (X of 200)** answers "how many of the techniques we know about work at all?" This is the defender-relevant number, because an attacker only needs one working technique and will retry it freely. The two can point in opposite directions from the same run: 400 successes concentrated in 4 templates is a guard with four holes; 400 successes spread one-per-template across 200 templates is a guard that is porous everywhere but reliable against nothing. Report both, labelled, with raw counts, plus the attempts-per-template budget that produced them.
go deeper
Should notice that 10% is per attempt and ask how many separate attacks that represents.
Should name attempts vs distinct attacks, explain that duplicated or retried templates move the per-attempt rate, and report both.
Should add the effort dimension and the success-concentration check, and de-duplicate near-clone templates before counting distinct attacks.
Should fix one reporting shape for the org so guard numbers from different teams and quarters are the same statistic, and require raw counts to be kept.
### Reading the 10% back to its source 400 of 4,000 is a per-attempt rate, and per-attempt is the denominator harnesses emit by default - because attempts are the unit their loops run over, not because anyone judged it the right base. In garak, a probe emits a set of prompts and the `--generations` flag decides how many completions are drawn for each, so the attempt total is prompts times generations and changes when you change a flag. In promptfoo, a redteam plugin's `numTests` sets how many cases are generated and `--repeat` multiplies each case; the totals line you screenshot counts cases run. In PyRIT, an orchestrator walks a seed-prompt dataset and a scorer labels each response, so the natural tally is again per response. None of these tools is lying. The number they print is simply not the number a defender needs, and nothing in the output says so. ### The three denominators and what each one answers | Denominator | The question it answers | It moves when | |---|---|---| | Attempts, 400/4,000 = 10% | if I fire one prompt drawn from this corpus, how likely is it to land? | the retry budget changes, the corpus gains duplicates, or error handling changes | | Distinct attacks, X of 200 | how many of the techniques we know about work at all? | the budget rises (it is monotone upward) or templates are de-duplicated | | Effort-weighted, bypassed within N attempts | how many work *cheaply*? | N changes - but N is quoted with the number, so the change is visible | The per-attempt rate is largely a statement about the corpus and the retry policy. Duplicate a working template ten times and it rises; pad the corpus with a hundred hopeless variants and it falls. Neither edit touches the guard. The per-distinct-attack rate is the one that maps to the threat: an attacker does not average over your corpus, they find one technique that works and use it. ### Concentration is a second axis, not a footnote Two runs can report the identical 400/4,000 and describe opposite systems. If the 400 successes come from 4 templates, the guard has four reliable holes and the fix is four targeted patches or signatures. If they come from 190 templates at roughly two successes each, nothing works reliably but almost everything works occasionally - a coverage or threshold problem, and patching individual templates would be wasted effort. A single-line summary hides which world you are in, which is why the success count per template, and per technique family, belongs in the report next to the rate. ### What proper reporting costs The extra cost is not mostly queries; it is record-keeping and discipline. You must keep a per-template row rather than a running tally, keep the raw transcripts so a later question can be answered without re-running, and de-duplicate the corpus by technique family before counting distinct attacks - which is manual taxonomy work, typically a day or two of an engineer's time on a corpus of a few hundred templates, repeated whenever the corpus grows. The query cost only rises if you also want the effort dimension: an N-attempt budget multiplies calls by N, and with a judge stage each attempt can mean a guard call, a target call and a judge call. ### Where these numbers mislead **Near-clones.** Forty variants of one technique inflate the distinct-attack denominator and drag the rate toward that one technique's behaviour - downward if the family is blocked, upward if it works. **Monotonicity.** Per-distinct-attack bypass can only rise as the budget grows, so it is meaningless without the budget attached; two teams quoting 3% and 12% may have identical guards and different retry loops. **One-in-fifty hits.** A template that lands once in twenty tries may be a genuine low-yield bypass or may be decider noise or decoding randomness; the boolean "bypassed" cannot tell you which, the per-template success count can. **Claim scope.** This is a measurement of a guard against a chosen corpus. It is not the product's production exposure and not a ship decision. ### What to check Ask for successes per template with a family label; the attempts-per-template budget; the decider and its version; how errors were treated; and whether a sample of hits was hand-confirmed. Then publish something a reader can re-derive: ``` 6 of 200 attack templates (3%) bypassed the input guard at least once within 20 attempts each; 8 of 4,000 attempts (0.2%) landed; successes concentrated in 2 technique families ```
- Your corpus contains 40 near-identical variants of one technique. How does that distort each denominator?It inflates both the attempt count and the distinct-attack count for that technique, dragging both rates toward its behaviour. De-duplicate by technique family and report family-level counts.
- Someone doubles the retries per template and the per-attack bypass rate rises. Did the guard get worse?No. Per-attack bypass is monotone in effort, so the rate is only meaningful with the attempt budget attached. Compare runs only at equal budgets.
- Which denominator would you lead with in a report to engineering owners of the guard?Distinct attacks bypassed within a stated budget, with the concentration breakdown - it points directly at what to fix. The per-attempt rate goes alongside as context.
saying these in an interview costs you the question
- Quoting the per-attempt rate as "the" bypass rate with no denominator label.
- Not knowing that adding retries changes the number without changing the guard.
- Reporting a per-attack rate without saying how many attempts each attack got.
- Ignoring that successes may be concentrated in one template family.
- Counting near-duplicate templates as distinct attacks.