skip to content

Two engineers test the same input-moderation guard with the same 200-template attack corpus: one allowed 50 retries per template, the other 1, and they quote different bypass rates. How would you define an effort-weighted bypass rate that makes their runs comparable, and what does it cost you?

level: seniorimportance: should knowfreq 42%

answer

  1. pin the budget N, then compare
  2. attempts-to-first-bypass, censored at N
  3. cannot rescale 1 try up to 50
  4. store the per-template table, not the summary
  5. choose N from throttling and bans

basics

~20 s

Fix the attempt budget and count templates that bypass within it - say, the share of the 200 that got through in 10 attempts or fewer. Better, record attempts-to-first-bypass per template and report its distribution. Both runs must use the same budget, so you re-run the cheaper one. The cost is queries, money and wall-clock against a metered endpoint.

solid answer

~50 s

Comparability comes from pinning effort, not from arithmetic on the numbers you already have. **Definition.** For each template, record attempts-to-first-bypass, censored at a budget N. Report the fraction of templates bypassed within N (a survival-style curve as N grows is even better) and the median attempts-to-first-bypass among those that ever bypassed. Both runs must use the same N, so the 1-retry run has to be re-run - you cannot rescale a 1-try result up to 50 tries, because success probability per template is unknown and wildly unequal. **Why it is the right statistic.** It carries a cost dimension: a technique needing 40 tries is a different threat from one that lands first time, especially where rate limits, per-user throttling or account bans make repetition expensive for a real attacker. **What it costs.** N times the queries and spend; per-template attempt logs rather than a single tally; and enough repeats that a one-in-fifty hit can be told apart from decider noise or sampling variance.

go deeper

for a junior

Should at least see that different retry budgets make the two numbers incomparable and that one run needs redoing.

for a middle

Should propose a fixed attempts-per-template budget and count templates bypassed within it, at equal N for both runs.

for a senior

Should define attempts-to-first-bypass with censoring, keep the per-template table, control decider version and target sampling settings, and price the query cost.

for a principal

Should tie the choice of N to the deployed system's throttling and ban policy, and standardise the budget so quarterly numbers stay comparable.

### Why the two runs are not the same measurement Both engineers ran the same corpus against the same guard, so the instinct is that one number must be adjustable into the other. It is not. Per-distinct-attack bypass is monotone in effort - more tries can only add templates to the "bypassed at least once" set - while the per-attempt rate is a fraction whose denominator also grows, so mostly-failing retries push it down. Effort therefore moves the two headline rates in *opposite* directions, and neither run measured a property of the guard alone. The fix is to promote effort from a hidden confound to an explicit axis of the measurement. ### The statistic to record Treat each attack template as a repeated trial with an unknown per-attempt success probability, and record time-to-first-success in attempts, right-censored at a budget N: ``` template_id, family, attempts_used, bypassed(bool), first_success_attempt, successes_total ``` That single table regenerates every summary anyone will later ask for: the fraction of templates bypassed within 1, within 10, within N; the per-attempt rate over the whole run; the median attempts-to-first-bypass among templates that ever bypassed; and the concentration of successes across families. Storing the table instead of the summary is the actual engineering advice here, because the summary someone wants next quarter is never the one you computed this quarter. ### Making two runs comparable Four things must be pinned, not three. **Equal N** for both runs - the one-retry run has to be re-run, because a single-attempt result cannot be rescaled to fifty attempts without assuming a uniform per-attempt success probability across techniques, an assumption that fails badly (some templates land nearly always, most never). **The same decider and decider version**, since a change in what counts as a hit moves the numerator invisibly. **The same de-duplicated template set**, with family labels. **The same sampling settings on the target** - higher temperature or top-p raises the chance that any given template eventually lands, lowering attempts-to-first-bypass with no change whatsoever in the guard. ### What it costs The visible cost is N-fold queries. At 200 templates and N = 50 that is 10,000 guard calls; with a target call and a judge call per attempt it is around 30,000 requests, and the judge stage - a full transcript sent to a capable model - dominates both the money and the latency. Against a rate-limited hosted endpoint that is typically hours to a day of wall-clock, which inside a fixed engagement window is the real constraint, not the invoice. Add the invisible costs: per-template attempt logs instead of a tally, storage for the raw transcripts, and reviewer time to hand-confirm a sample of hits. Budget those before promising the number, because the usual failure is discovering at day four of a five-day engagement that the re-run at equal N will not fit. ### Where the number misleads **Choosing N to suit the tool rather than the threat.** If the deployed product throttles, rate-limits or bans an account after a handful of refusals, a bypass rate at N = 200 describes an attacker the system does not permit to exist. Pick N from the deployment's tolerance for repetition, and publish the large-N curve as an appendix rather than as the headline. **Reading a censored zero as a negative.** "Not bypassed within N" is not "cannot be bypassed"; censoring destroys exactly that information, and no amount of re-analysis recovers it. **Trusting a one-in-fifty hit.** It may be a genuine low-yield bypass, decider mislabelling, or decoding randomness - report the per-template success count, not just the boolean, and hand-confirm before it becomes a headline. **Comparing medians across runs with different N**, since the median attempts-to-first-bypass is computed only over templates that succeeded, and that population changes with N. ### What to check Before believing either engineer, confirm: the two runs used the same N, the same decider version, the same target sampling settings, and the same de-duplicated corpus; that errored attempts were logged rather than dropped; and that a sample of the hits was manually confirmed. Then report the fraction bypassed within N with the counts, the curve across smaller thresholds, and the justification for N drawn from the deployed system's own throttling policy.

  • Why can't you extrapolate a 1-attempt run to a 50-attempt result?
    Because per-template success probabilities are unknown and very unequal. A single overall estimate assumes homogeneity that never holds across attack techniques.
  • How do you pick N?
    From the deployed system's tolerance for repetition - rate limits, throttling, account bans. If a real attacker gets a handful of tries before being cut off, N should be that, with a larger-N curve reported separately.
  • A template bypasses once in 50 attempts. Is it a finding?
    Provisionally. Confirm the hit manually and check the decider isn't mislabelling; if real, report it with its per-attempt success count so its low yield is visible.

Attempts-to-first-bypass asks how many keys off the ring you tried before the door opened, not merely whether it ever opened. With a thousand keys almost any door opens eventually, so the number only means something once you state how many keys the attacker actually gets before the alarm goes off.

saying these in an interview costs you the question

  • Rescaling a 1-attempt run arithmetically to estimate a 50-attempt result.
  • Comparing two runs at different retry budgets without saying so.
  • Reporting only a boolean bypassed/not per template and discarding attempt counts.
  • Changing the decider or the target's sampling settings between the two runs and still comparing.
  • Picking a huge N because it produces a scarier number, with no threat-model justification.

context