skip to content

Base Rates

In production almost no traffic is an attack, so a rate measured on an all-attack corpus misleads about volume. Interviewers probe this because that gap makes a scary headline number unusable.

on this pageshow

explore

questions

5

A guardrail evaluation runs 500 prompts, every one of them an adversarial attack, and 12% of them reach the model unblocked. A stakeholder reads that as "12% of our traffic is getting through". Why is that reading wrong?

level: juniorimportance: must knowfreq 68%

answer

  1. corpus is 100% attack
  2. conditional, not unconditional
  3. denominator = attacks, not requests
  4. multiply by prevalence for exposure
  5. label the population

basics

~20 s

The 12% is conditional on the prompt already being an attack, because the corpus is 100% attacks. The denominator is attack prompts, not requests. Real traffic is almost entirely benign, so the share of all requests that are successful attacks is 12% multiplied by how rare attacks actually are.

solid answer

~50 s

The test corpus fixes the prevalence of attacks at 100%. Any rate you measure on it is a **conditional** quantity: given that a prompt is an attack, how often does the screening layer let it through. Production prevalence is nowhere in that number. To talk about traffic you need a second input the test never supplies — what fraction of live requests are attacks at all. If that is one in a thousand, the fraction of *all* requests that are successful attacks is 0.001 x 0.12, roughly one in eight thousand, not one in eight. Two orders of magnitude separate the headline from the exposure. The conditional number is still the right thing to measure — it is the only one that is stable when traffic mix changes, and it is what lets you compare two screening layers. It is just not an exposure figure, and it should never be quoted as one without saying what the denominator was.

go deeper

for a junior

Says the corpus was all attacks so the percentage is out of attacks, not out of traffic, and that real traffic is mostly benign.

for a middle

Names the conditional versus unconditional distinction and shows the multiplication by prevalence that converts one into the other.

for a senior

Points out that both framings mislead in opposite directions, and insists the report carry both numbers with their populations and the prevalence assumption labelled.

for a principal

Frames it as a reporting-contract problem: decide up front which population each number in the deliverable describes, because the slide will outlive the conversation that explained it.

### What the harness actually computed A guardrail evaluation corpus is a fixed list of prompts that a harness sends, one at a time, at the screening layer sitting in front of the model — an input classifier such as Llama Guard or Prompt Guard, a hosted service such as OpenAI's moderation endpoint or Azure AI Content Safety, or a rail framework's rule set. For each prompt the harness records one binary outcome: the screen blocked it, or the prompt reached the model. Here 500 prompts were sent, 60 were not blocked, and the tool printed `60 / 500 = 12%`. The denominator is the corpus, and every member of the corpus was adversarial by construction. So the quantity is **conditional**: given that a prompt is an attack of the kind this corpus contains, how often did the screen fail to block it. In notation, `P(not blocked | attack)`. No benign prompt was ever sent, so no fact about the mix of production traffic could possibly have entered the arithmetic. The stakeholder's reading — "12% of our traffic" — silently swaps the denominator from *attack prompts in this corpus* to *all requests the application serves*, two populations that differ by three or four orders of magnitude in size. ### Why the corpus is built this way, and what it costs The all-attack shape is deliberate, not sloppy. Attacks are rare in real traffic; if you sampled to match production at, say, one attack per thousand requests, a corpus large enough to contain 500 attacks would be half a million prompts, and you would pay to run and score every one of them. Oversampling the rare event to 100% buys statistical power at a fraction of the cost. The run itself is cheap. Five hundred prompts against a hosted screen is 500 classifier calls, plus 500 generator calls if you also send survivors to the model, plus a judge call per survivor if a model grades the output — order 1,000–1,500 API calls, single-digit to low-tens of dollars, and minutes of wall-clock at modest concurrency. The real expense is upstream and human: curating and labelling the corpus, deciding what counts as an attack, and keeping it current. That is engineer-days, and it is why teams reuse a corpus long after the attack landscape has moved. ### Where the number misleads Both mistranslations do damage, in opposite directions. - **Read as traffic**, 12% sounds like an active emergency. Nobody can find the corresponding incidents, the report loses credibility, and the next real finding is discounted. - **Restated at prevalence** — `0.001 × 0.12`, roughly one request in eight thousand — it can read as a non-issue. But a motivated attacker sends nothing but attacks: for them prevalence is 1 and they experience the full 12%. The diluted figure describes ambient background traffic and nobody else. Two further ways the 12% itself can be the wrong rate. It is a weighted average over whatever techniques the corpus happens to contain, so if the tool ships mostly one family and real attackers favour another, the conditional rate is measured on the wrong population before prevalence is even discussed. And "reached the model unblocked" is not "produced harmful output that reached a user"; a bypass count is an upper bound on harm, not a count of it. | | measures | needs prevalence? | stable when traffic mix changes? | |---|---|---|---| | Conditional bypass rate, 12% | screen quality against attacks | no | yes | | Unconditional exposure | share of all requests that are successful attacks | yes | no | ### What to check before quoting it Ask what the denominator was — all corpus prompts, distinct techniques, or attempts including retries, since retrying one technique until it works inflates a per-attempt rate. Ask whether the corpus mix resembles what attackers actually send. Ask whether the counted event is "reached the model" or "produced harm". And ask whether any prevalence figure exists at all, or is being assumed silently. The defensible report gives both numbers side by side, each labelled with the population it describes, plus the prevalence assumption and its source. That version survives being pasted into a slide deck by someone who was not in the room — which is the only test that matters, because the slide outlives the conversation.

  • If the conditional rate is the misleading one, why not build the test corpus to match production traffic instead?
    Because attacks would then be a fraction of a percent of the corpus, and you would need enormous samples to measure the bypass rate at all. You oversample deliberately and restate afterwards; you do not throw away statistical power to make one number easier to read.
  • Which of the two figures do you use to compare last quarter's screening layer against this quarter's?
    The conditional rate. It is unaffected by traffic mix shifting between the two measurements, so a change in it reflects a change in the guardrail rather than a change in who happened to be sending traffic.

A drug trial that enrols only people who already have the disease can tell you how often the drug works, but it cannot tell you how many people in the city will get better - that needs the disease's prevalence, which the trial never measured.

saying these in an interview costs you the question

  • Treating the all-attack rate as a production traffic percentage without noticing the denominator
  • Concluding the corpus was badly built and should have been sampled to match production
  • Claiming the restated exposure number makes the finding unimportant
  • Cannot state what the denominator of their own headline number was

context

open as a page

Your guardrail test on an all-attack prompt corpus shows a 12% bypass rate. How do you restate that as an expected number of successful attacks per day against the deployed application?

level: middleimportance: must knowfreq 55%

basics

~20 s

Multiply three numbers: daily request volume, the estimated share of requests that are attacks, and the 12% bypass rate. Ten million requests a day at one attack per thousand gives ten thousand attacks and about twelve hundred bypasses. Publish the prevalence estimate and its source next to the result.

open as a page

To restate a guardrail bypass rate at production prevalence you need to know what fraction of live requests are attacks. How do you estimate that, and why is deriving it from the deployed guardrail's own block counts circular?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Estimate it by hand-labelling a random sample of raw traffic, not from what the guardrail blocked. The block counts include only attacks it caught, so they omit exactly the misses you are trying to size, biasing prevalence low. Report the estimate as a range with the sample size that produced it.

open as a page

A prompt-screening classifier in front of an LLM app catches 90% of attacks in your test corpus and flags 1% of benign prompts. If roughly 1 production request in 1,000 is an attack, what share of the classifier's flags are genuine attacks, and how should that land in your report?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Per million requests: 1,000 attacks produce 900 flags, and 999,000 benign prompts produce about 9,990 flags. So roughly 10,890 flags contain 900 real attacks, about 8%. Around eleven in twelve flags are benign. Report that share, because a rate measured on an all-attack corpus never reveals it.

open as a page

Restating a guardrail bypass rate at production attack prevalence assumes prevalence is a stable property of your traffic. When is that assumption unsafe, and how do you present the restated number so it does not become false reassurance?

level: principalimportance: should knowfreq 30%

basics

~20 s

Prevalence is partly chosen by attackers, so it is not a fixed property of your traffic. A targeted attacker sends 100% attacks and experiences the raw conditional rate; a campaign can raise prevalence overnight. Present both the ambient restatement and the targeted case, and state the prevalence assumption on the same page.

open as a page