A guardrail evaluation runs 500 prompts, every one of them an adversarial attack, and 12% of them reach the model unblocked. A stakeholder reads that as "12% of our traffic is getting through". Why is that reading wrong?
answer
- corpus is 100% attack
- conditional, not unconditional
- denominator = attacks, not requests
- multiply by prevalence for exposure
- label the population
basics
~20 sThe 12% is conditional on the prompt already being an attack, because the corpus is 100% attacks. The denominator is attack prompts, not requests. Real traffic is almost entirely benign, so the share of all requests that are successful attacks is 12% multiplied by how rare attacks actually are.
solid answer
~50 sThe test corpus fixes the prevalence of attacks at 100%. Any rate you measure on it is a **conditional** quantity: given that a prompt is an attack, how often does the screening layer let it through. Production prevalence is nowhere in that number. To talk about traffic you need a second input the test never supplies — what fraction of live requests are attacks at all. If that is one in a thousand, the fraction of *all* requests that are successful attacks is 0.001 x 0.12, roughly one in eight thousand, not one in eight. Two orders of magnitude separate the headline from the exposure. The conditional number is still the right thing to measure — it is the only one that is stable when traffic mix changes, and it is what lets you compare two screening layers. It is just not an exposure figure, and it should never be quoted as one without saying what the denominator was.
go deeper
Says the corpus was all attacks so the percentage is out of attacks, not out of traffic, and that real traffic is mostly benign.
Names the conditional versus unconditional distinction and shows the multiplication by prevalence that converts one into the other.
Points out that both framings mislead in opposite directions, and insists the report carry both numbers with their populations and the prevalence assumption labelled.
Frames it as a reporting-contract problem: decide up front which population each number in the deliverable describes, because the slide will outlive the conversation that explained it.
### What the harness actually computed A guardrail evaluation corpus is a fixed list of prompts that a harness sends, one at a time, at the screening layer sitting in front of the model — an input classifier such as Llama Guard or Prompt Guard, a hosted service such as OpenAI's moderation endpoint or Azure AI Content Safety, or a rail framework's rule set. For each prompt the harness records one binary outcome: the screen blocked it, or the prompt reached the model. Here 500 prompts were sent, 60 were not blocked, and the tool printed `60 / 500 = 12%`. The denominator is the corpus, and every member of the corpus was adversarial by construction. So the quantity is **conditional**: given that a prompt is an attack of the kind this corpus contains, how often did the screen fail to block it. In notation, `P(not blocked | attack)`. No benign prompt was ever sent, so no fact about the mix of production traffic could possibly have entered the arithmetic. The stakeholder's reading — "12% of our traffic" — silently swaps the denominator from *attack prompts in this corpus* to *all requests the application serves*, two populations that differ by three or four orders of magnitude in size. ### Why the corpus is built this way, and what it costs The all-attack shape is deliberate, not sloppy. Attacks are rare in real traffic; if you sampled to match production at, say, one attack per thousand requests, a corpus large enough to contain 500 attacks would be half a million prompts, and you would pay to run and score every one of them. Oversampling the rare event to 100% buys statistical power at a fraction of the cost. The run itself is cheap. Five hundred prompts against a hosted screen is 500 classifier calls, plus 500 generator calls if you also send survivors to the model, plus a judge call per survivor if a model grades the output — order 1,000–1,500 API calls, single-digit to low-tens of dollars, and minutes of wall-clock at modest concurrency. The real expense is upstream and human: curating and labelling the corpus, deciding what counts as an attack, and keeping it current. That is engineer-days, and it is why teams reuse a corpus long after the attack landscape has moved. ### Where the number misleads Both mistranslations do damage, in opposite directions. - **Read as traffic**, 12% sounds like an active emergency. Nobody can find the corresponding incidents, the report loses credibility, and the next real finding is discounted. - **Restated at prevalence** — `0.001 × 0.12`, roughly one request in eight thousand — it can read as a non-issue. But a motivated attacker sends nothing but attacks: for them prevalence is 1 and they experience the full 12%. The diluted figure describes ambient background traffic and nobody else. Two further ways the 12% itself can be the wrong rate. It is a weighted average over whatever techniques the corpus happens to contain, so if the tool ships mostly one family and real attackers favour another, the conditional rate is measured on the wrong population before prevalence is even discussed. And "reached the model unblocked" is not "produced harmful output that reached a user"; a bypass count is an upper bound on harm, not a count of it. | | measures | needs prevalence? | stable when traffic mix changes? | |---|---|---|---| | Conditional bypass rate, 12% | screen quality against attacks | no | yes | | Unconditional exposure | share of all requests that are successful attacks | yes | no | ### What to check before quoting it Ask what the denominator was — all corpus prompts, distinct techniques, or attempts including retries, since retrying one technique until it works inflates a per-attempt rate. Ask whether the corpus mix resembles what attackers actually send. Ask whether the counted event is "reached the model" or "produced harm". And ask whether any prevalence figure exists at all, or is being assumed silently. The defensible report gives both numbers side by side, each labelled with the population it describes, plus the prevalence assumption and its source. That version survives being pasted into a slide deck by someone who was not in the room — which is the only test that matters, because the slide outlives the conversation.
- If the conditional rate is the misleading one, why not build the test corpus to match production traffic instead?Because attacks would then be a fraction of a percent of the corpus, and you would need enormous samples to measure the bypass rate at all. You oversample deliberately and restate afterwards; you do not throw away statistical power to make one number easier to read.
- Which of the two figures do you use to compare last quarter's screening layer against this quarter's?The conditional rate. It is unaffected by traffic mix shifting between the two measurements, so a change in it reflects a change in the guardrail rather than a change in who happened to be sending traffic.
A drug trial that enrols only people who already have the disease can tell you how often the drug works, but it cannot tell you how many people in the city will get better - that needs the disease's prevalence, which the trial never measured.
saying these in an interview costs you the question
- Treating the all-attack rate as a production traffic percentage without noticing the denominator
- Concluding the corpus was badly built and should have been sampled to match production
- Claiming the restated exposure number makes the finding unimportant
- Cannot state what the denominator of their own headline number was