skip to content

Measuring a Guard

Denominators, over-blocking, threshold curves and production base rates decide what a bypass number means. Interviewers probe it because "we got through" is not a finding until it carries a rate.

on this pageshow

explore

questions

18

A guardrail evaluation runs 500 prompts, every one of them an adversarial attack, and 12% of them reach the model unblocked. A stakeholder reads that as "12% of our traffic is getting through". Why is that reading wrong?

level: juniorimportance: must knowfreq 68%

answer

  1. corpus is 100% attack
  2. conditional, not unconditional
  3. denominator = attacks, not requests
  4. multiply by prevalence for exposure
  5. label the population

basics

~20 s

The 12% is conditional on the prompt already being an attack, because the corpus is 100% attacks. The denominator is attack prompts, not requests. Real traffic is almost entirely benign, so the share of all requests that are successful attacks is 12% multiplied by how rare attacks actually are.

solid answer

~50 s

The test corpus fixes the prevalence of attacks at 100%. Any rate you measure on it is a **conditional** quantity: given that a prompt is an attack, how often does the screening layer let it through. Production prevalence is nowhere in that number. To talk about traffic you need a second input the test never supplies — what fraction of live requests are attacks at all. If that is one in a thousand, the fraction of *all* requests that are successful attacks is 0.001 x 0.12, roughly one in eight thousand, not one in eight. Two orders of magnitude separate the headline from the exposure. The conditional number is still the right thing to measure — it is the only one that is stable when traffic mix changes, and it is what lets you compare two screening layers. It is just not an exposure figure, and it should never be quoted as one without saying what the denominator was.

go deeper

for a junior

Says the corpus was all attacks so the percentage is out of attacks, not out of traffic, and that real traffic is mostly benign.

for a middle

Names the conditional versus unconditional distinction and shows the multiplication by prevalence that converts one into the other.

for a senior

Points out that both framings mislead in opposite directions, and insists the report carry both numbers with their populations and the prevalence assumption labelled.

for a principal

Frames it as a reporting-contract problem: decide up front which population each number in the deliverable describes, because the slide will outlive the conversation that explained it.

### What the harness actually computed A guardrail evaluation corpus is a fixed list of prompts that a harness sends, one at a time, at the screening layer sitting in front of the model — an input classifier such as Llama Guard or Prompt Guard, a hosted service such as OpenAI's moderation endpoint or Azure AI Content Safety, or a rail framework's rule set. For each prompt the harness records one binary outcome: the screen blocked it, or the prompt reached the model. Here 500 prompts were sent, 60 were not blocked, and the tool printed `60 / 500 = 12%`. The denominator is the corpus, and every member of the corpus was adversarial by construction. So the quantity is **conditional**: given that a prompt is an attack of the kind this corpus contains, how often did the screen fail to block it. In notation, `P(not blocked | attack)`. No benign prompt was ever sent, so no fact about the mix of production traffic could possibly have entered the arithmetic. The stakeholder's reading — "12% of our traffic" — silently swaps the denominator from *attack prompts in this corpus* to *all requests the application serves*, two populations that differ by three or four orders of magnitude in size. ### Why the corpus is built this way, and what it costs The all-attack shape is deliberate, not sloppy. Attacks are rare in real traffic; if you sampled to match production at, say, one attack per thousand requests, a corpus large enough to contain 500 attacks would be half a million prompts, and you would pay to run and score every one of them. Oversampling the rare event to 100% buys statistical power at a fraction of the cost. The run itself is cheap. Five hundred prompts against a hosted screen is 500 classifier calls, plus 500 generator calls if you also send survivors to the model, plus a judge call per survivor if a model grades the output — order 1,000–1,500 API calls, single-digit to low-tens of dollars, and minutes of wall-clock at modest concurrency. The real expense is upstream and human: curating and labelling the corpus, deciding what counts as an attack, and keeping it current. That is engineer-days, and it is why teams reuse a corpus long after the attack landscape has moved. ### Where the number misleads Both mistranslations do damage, in opposite directions. - **Read as traffic**, 12% sounds like an active emergency. Nobody can find the corresponding incidents, the report loses credibility, and the next real finding is discounted. - **Restated at prevalence** — `0.001 × 0.12`, roughly one request in eight thousand — it can read as a non-issue. But a motivated attacker sends nothing but attacks: for them prevalence is 1 and they experience the full 12%. The diluted figure describes ambient background traffic and nobody else. Two further ways the 12% itself can be the wrong rate. It is a weighted average over whatever techniques the corpus happens to contain, so if the tool ships mostly one family and real attackers favour another, the conditional rate is measured on the wrong population before prevalence is even discussed. And "reached the model unblocked" is not "produced harmful output that reached a user"; a bypass count is an upper bound on harm, not a count of it. | | measures | needs prevalence? | stable when traffic mix changes? | |---|---|---|---| | Conditional bypass rate, 12% | screen quality against attacks | no | yes | | Unconditional exposure | share of all requests that are successful attacks | yes | no | ### What to check before quoting it Ask what the denominator was — all corpus prompts, distinct techniques, or attempts including retries, since retrying one technique until it works inflates a per-attempt rate. Ask whether the corpus mix resembles what attackers actually send. Ask whether the counted event is "reached the model" or "produced harm". And ask whether any prevalence figure exists at all, or is being assumed silently. The defensible report gives both numbers side by side, each labelled with the population it describes, plus the prevalence assumption and its source. That version survives being pasted into a slide deck by someone who was not in the room — which is the only test that matters, because the slide outlives the conversation.

  • If the conditional rate is the misleading one, why not build the test corpus to match production traffic instead?
    Because attacks would then be a fraction of a percent of the corpus, and you would need enormous samples to measure the bypass rate at all. You oversample deliberately and restate afterwards; you do not throw away statistical power to make one number easier to read.
  • Which of the two figures do you use to compare last quarter's screening layer against this quarter's?
    The conditional rate. It is unaffected by traffic mix shifting between the two measurements, so a change in it reflects a change in the guardrail rather than a change in who happened to be sending traffic.

A drug trial that enrols only people who already have the disease can tell you how often the drug works, but it cannot tell you how many people in the city will get better - that needs the disease's prevalence, which the trial never measured.

saying these in an interview costs you the question

  • Treating the all-attack rate as a production traffic percentage without noticing the denominator
  • Concluding the corpus was badly built and should have been sampled to match production
  • Claiming the restated exposure number makes the finding unimportant
  • Cannot state what the denominator of their own headline number was

context

open as a page

A teammate reports "our input-moderation guard has a 3% bypass rate" from a red-team harness run. Before you repeat that number in a report, what must the number state about what was counted?

level: juniorimportance: must knowfreq 62%

basics

~20 s

It must say what counted as a bypass in the numerator: the guard failed to flag, or the model then actually complied - two different numbers. It must say whether the denominator is attempts, distinct attacks, or attempts per attack. And it must say how many retries each attack got, because retries move the rate without the guard changing.

open as a page

A guardrail test corpus for an input classifier contains only attack prompts. Why does a classifier that blocks every single input score perfectly on that corpus, and what must be added before the score means anything?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Every item in the corpus is an attack, so blocking everything means catching everything: the measured block rate hits 100% and nothing penalises refusals. The corpus can only observe one error type. Add benign prompts the product really sees, and report how many of those were wrongly blocked alongside the attack numbers.

open as a page

Your red-team report says a content-moderation guard let 3% of your attack prompts through. Which configuration facts have to sit next to that number before anyone else can reproduce or compare it?

level: juniorimportance: must knowfreq 68%

basics

~20 s

The block threshold you ran at, which categories were set to block, the guard's identity and configuration snapshot, and the date. A bypass percentage is only true at one operating point, so without those settings written down nobody can rerun your test or compare a later number to yours.

open as a page

Your guardrail test on an all-attack prompt corpus shows a 12% bypass rate. How do you restate that as an expected number of successful attacks per day against the deployed application?

level: middleimportance: must knowfreq 55%

basics

~20 s

Multiply three numbers: daily request volume, the estimated share of requests that are attacks, and the 12% bypass rate. Ten million requests a day at one attack per thousand gives ten thousand attacks and about twelve hundred bypasses. Publish the prevalence estimate and its source next to the result.

open as a page

A guardrail test harness fired 4,000 prompt attempts, drawn from 200 distinct attack templates, at an input-moderation classifier; 400 attempts reached the model. Why is "the guard has a 10% bypass rate" an incomplete claim, and which denominators would you report instead?

level: middleimportance: must knowfreq 66%

basics

~20 s

10% is 400 over attempts, so it partly measures how the corpus was sampled and retried rather than the guard. Report it beside the per-attack rate: how many of the 200 distinct templates got past at least once. If 180 bypassed, the guard is far weaker than 10% suggests; if 4 bypassed 100 times each, far stronger.

open as a page

You are assembling the benign half of a test corpus for a content-moderation guard. Where do the benign items come from, and what makes a benign request a hard negative rather than an easy one?

level: middleimportance: must knowfreq 55%

basics

~20 s

Pull benign items from the product's real traffic and from topics sitting right next to the blocked ones: safety questions, clinical or legal wording, security research, fiction, non-English phrasing. A hard negative looks like an attack on the surface but is a request the product must answer. Label them, review the labels, freeze the set.

open as a page

In a red-team report on a moderation guard, why present attack bypass across the guard's whole score range rather than as one percentage at the threshold the team currently runs?

level: middleimportance: must knowfreq 58%

basics

~20 s

One threshold collapses the whole score range into a flattering number and hides where the guard actually breaks. A curve shows how bypass changes as the threshold moves, so the reader sees whether the guard is genuinely strong or just parked at a lucky setting, and what a small retune would cost.

open as a page

To restate a guardrail bypass rate at production prevalence you need to know what fraction of live requests are attacks. How do you estimate that, and why is deriving it from the deployed guardrail's own block counts circular?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Estimate it by hand-labelling a random sample of raw traffic, not from what the guardrail blocked. The block counts include only attacks it caught, so they omit exactly the misses you are trying to size, biasing prevalence low. Report the estimate as a range with the sample size that produced it.

open as a page

A prompt-screening classifier in front of an LLM app catches 90% of attacks in your test corpus and flags 1% of benign prompts. If roughly 1 production request in 1,000 is an attack, what share of the classifier's flags are genuine attacks, and how should that land in your report?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Per million requests: 1,000 attacks produce 900 flags, and 999,000 benign prompts produce about 9,990 flags. So roughly 10,890 flags contain 900 real attacks, about 8%. Around eleven in twelve flags are benign. Report that share, because a rate measured on an all-attack corpus never reveals it.

open as a page

Two engineers test the same input-moderation guard with the same 200-template attack corpus: one allowed 50 retries per template, the other 1, and they quote different bypass rates. How would you define an effort-weighted bypass rate that makes their runs comparable, and what does it cost you?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Fix the attempt budget and count templates that bypass within it - say, the share of the 200 that got through in 10 attempts or fewer. Better, record attempts-to-first-bypass per template and report its distribution. Both runs must use the same budget, so you re-run the cheaper one. The cost is queries, money and wall-clock against a metered endpoint.

open as a page

You are writing up over-blocking results for a moderation guard in front of a chat product. What counts as an over-block event, and why is a single aggregate false-positive percentage not enough for the product team to act on?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Count an over-block whenever a benign request fails to get the answer it deserved: a hard refusal, a hedged non-answer, or a silently rewritten reply. One overall percentage hides which kinds of user are hit, so break it out by benign category and hand over the actual blocked requests as examples.

open as a page

You test a moderation guard that returns a per-prompt score, and most of your attack prompts that were blocked scored only barely above the block threshold. What does that clustering tell you, and what goes in the report?

level: seniorimportance: should knowfreq 34%

basics

~20 s

It means the guard barely separates those attacks from allowed traffic: the blocks are luck of the line, not margin. A small loosening, a guard update, or a slightly reworded attack flips many of them through. Report the margin distribution and how bypass moves for small threshold changes, not just the current number.

open as a page

A hosted moderation service returns only "allowed" or "blocked" for each prompt, with no numeric score. How do you still show a client where that guard breaks across its strictness range instead of handing over one bypass percentage?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Sweep whatever the service does expose instead of a score: each selectable strictness level and each category toggle. Run the same corpus at every setting and report a step curve over settings rather than a smooth score curve. Pin each configuration exactly, and state that the resolution is limited to the settings offered.

open as a page

Restating a guardrail bypass rate at production attack prevalence assumes prevalence is a stable property of your traffic. When is that assumption unsafe, and how do you present the restated number so it does not become false reassurance?

level: principalimportance: should knowfreq 30%

basics

~20 s

Prevalence is partly chosen by attackers, so it is not a fixed property of your traffic. A targeted attacker sends 100% attacks and experiences the raw conditional rate; a campaign can raise prevalence overnight. Present both the ambient restatement and the targeted case, and state the prevalence assumption on the same page.

open as a page

Your team quotes a bypass rate for the same input-moderation guard every release, but the attack corpus grows each quarter as new templates are added. Which denominator do you standardise on so two quarters' numbers mean the same thing, and how do you handle the newly added templates?

level: principalimportance: should knowfreq 30%

basics

~20 s

Split the corpus. Freeze a versioned baseline slice and report the bypass rate over that fixed set of distinct attacks at a fixed attempt budget - that number alone is compared quarter to quarter. Report newly added templates as a separate first-run number, never blended in. Blending makes the trend move when the corpus changes rather than the guard.

open as a page

A guardrail assessment has a fixed number of paid calls to a hosted moderation endpoint. How do you decide what share of that query budget goes to benign hard negatives instead of attack probes, and how do you defend the split to a stakeholder who wants everything spent on attacks?

level: principalimportance: should knowfreq 30%

basics

~20 s

Spend enough on benign traffic to detect an over-block rate the product would actually care about, then give the rest to attacks. Size the benign half from the smallest rate worth acting on, not from what is left over. Defend it by noting a catch-rate-only report cannot tell a good guard from a closed door.

open as a page

The team receiving your guardrail bypass findings also owns the guard's block threshold and can change it the day after you deliver. How do you make the figure you hand over still mean something a month later?

level: principalimportance: should knowfreq 24%

basics

~20 s

Hand over a function, not a constant: bypass across the threshold range, the operating point it was measured at, and the raw per-attempt scores so any future setting can be re-derived. Fix the operating point before results are seen, and make a threshold change trigger a re-measurement rather than a reinterpretation of your number.

open as a page