skip to content

Your guardrail test on an all-attack prompt corpus shows a 12% bypass rate. How do you restate that as an expected number of successful attacks per day against the deployed application?

level: middleimportance: must knowfreq 55%

answer

  1. volume x prevalence x bypass rate
  2. prevalence sourced outside the harness
  3. give a range, not one number
  4. corpus mix != attacker mix
  5. bypass is not incident

basics

~20 s

Multiply three numbers: daily request volume, the estimated share of requests that are attacks, and the 12% bypass rate. Ten million requests a day at one attack per thousand gives ten thousand attacks and about twelve hundred bypasses. Publish the prevalence estimate and its source next to the result.

solid answer

~60 s

The arithmetic is volume x prevalence x conditional bypass rate. Everything interesting is in the second factor, which the harness never measured and which you must source separately and state openly. Three caveats belong in the same paragraph as the number. First, the corpus mix is not the attacker mix: if the tool ships mostly one family of technique and real attackers favour another, the 12% is a rate for the wrong population and the restated volume inherits that error. Second, a bypass is not an incident — reaching the model is not the same as producing harmful output that reaches a user, and the report should say which of those you counted. Third, prevalence is an estimate with a wide band, so give the restated volume as a range driven by a low and a high prevalence figure rather than a single confident count. Done this way the number is usable for prioritisation: it converts a percentage nobody can act on into an events-per-day figure that can be compared against other risks.

go deeper

for a junior

Gets the multiplication right and knows the prevalence factor has to come from somewhere outside the test run.

for a middle

Sources prevalence explicitly, gives the answer as a range, and notes that corpus composition may not match real attacker behaviour.

for a senior

Adds that the counted event must match the harm of interest, refuses to use the guardrail's own block counts as the prevalence numerator, and presents a sensitivity table.

for a principal

Uses the restatement to decide whether the finding is actionable at all, and calls out when the real deliverable is a prevalence measurement rather than a bypass number.

### The model, written down so it can be argued with ``` expected successful attacks/day = requests/day # from the application's own logs × P(request is an attack) # prevalence — estimated OUTSIDE the harness × P(not blocked | attack) # the 12% the corpus run measured ``` Only the third factor came from the guardrail test. The first is a log query. The second is the whole problem: the harness never observed a benign request, so it cannot supply prevalence, and any restatement that omits it is applying an attack-conditional rate to a population that is almost entirely benign. Ten million requests a day at one attack per thousand gives ten thousand attack attempts and roughly 1,200 bypasses — not 1.2 million. ### Sourcing the prevalence factor, and what it costs In rough order of trustworthiness: 1. **Hand-label a random sample of raw inbound traffic** against a written definition of "attack". This is ground truth. A labeller works through perhaps 100–200 requests an hour, so a few thousand labelled requests is one to three engineer-days plus the meeting where you agree the definition — and the definition is half the work, because whether curious policy-probing counts can move the answer by an order of magnitude. 2. **A stratified sample with recorded weights** — heavy sampling of flagged traffic, thin sampling of the unflagged remainder, reweighted to the population. Much cheaper at low prevalence, valid only if the weights survive to the calculation. 3. **A prior engagement's labelled sample, or a comparable product's published order of magnitude**, explicitly marked as an assumption rather than a measurement. What you must never use is the deployed guardrail's own block counts. It cannot count what it missed, so that numerator omits exactly the population under study, and it makes the prevalence estimate a function of the miss rate you are about to multiply it by. ### Where the restated number misleads **Precision theatre.** Prevalence is routinely uncertain by an order of magnitude; the bypass rate, on a fixed corpus and model, usually is not. The product therefore inherits an order of magnitude of uncertainty, and printing "1,214 bypasses per day" claims four significant figures for a quantity whose leading digit is a guess. Publish a small sensitivity table across plausible prevalences instead. If the decision flips between the low and the high row, the real finding is "we need to measure prevalence", and that is a legitimate deliverable rather than a failure. **Wrong population.** The corpus's technique mix is not the attacker's technique mix. A conditional rate averaged over the tool's shipped probes describes attacks weighted the way the tool weights them; the restated volume inherits that error invisibly. **Wrong event.** "Bypass" usually means the prompt reached the model. It is not an incident, not harmful output, and not harmful output delivered to a user. Each of those is a strictly smaller number, and the report should say which one it counted. **Wrong reader.** The whole restatement models an ambient, non-adaptive arrival process. A single targeted adversary sends 100% attacks, so their prevalence factor is 1 and they experience the raw 12% — repeatedly, because they can retry. A campaign can move population prevalence by orders of magnitude in days. Give the ambient band and the targeted-case reading on the same page, or the small number becomes false reassurance precisely when it should not. | factor | source | typical uncertainty | |---|---|---| | requests/day | application logs | small, directly measured | | prevalence | labelled sample or stated assumption | often ±1 order of magnitude | | conditional bypass rate | the corpus run | moderate; sampling error on ~500 prompts | ### What to check before publishing Is the counted event the harm you actually care about? Are the volume figure and the prevalence sample drawn from the same time window — a prevalence measured on a weekday afternoon and a volume averaged over a month do not compose cleanly, because adversarial traffic is burstier than benign. Is the prevalence input visible on the same slide as the conclusion, with a date and a source, rather than in an appendix nobody opens? And does the deliverable still carry the conditional rate itself, which is the figure that does not expire when traffic mix shifts and the only one that lets you compare this quarter's screening layer against last quarter's? Done this way, the restatement earns its keep: it converts a percentage nobody can act on into an events-per-day figure that can be ranked against other risks, without quietly pretending the guess at its centre is a measurement.

  • Where do you get the prevalence factor if nobody has ever labelled the traffic?
    Label a random sample yourself, even a few hundred requests, which bounds the order of magnitude. Failing that, present the restatement as a sensitivity table across assumed prevalences and mark the input as an assumption rather than a measurement.
  • Why is the restated daily figure a poor description of a targeted attacker?
    A targeted attacker's traffic is 100% attacks, so the prevalence factor is 1 and they experience the raw conditional rate. The restatement models ambient background traffic only.

saying these in an interview costs you the question

  • Applying the bypass rate to total request volume without a prevalence factor
  • Quoting a single point estimate for a quantity whose main input is an order-of-magnitude guess
  • Estimating prevalence from the deployed guardrail's own block counts
  • Equating every bypass with a realised harmful output

context