A prompt-screening classifier in front of an LLM app catches 90% of attacks in your test corpus and flags 1% of benign prompts. If roughly 1 production request in 1,000 is an attack, what share of the classifier's flags are genuine attacks, and how should that land in your report?
answer
- two conditional rates, neither is precision
- benign majority dominates the flags
- 1M requests: 900 vs 9,990
- benign error rate is the bigger lever
- composition moves with prevalence
basics
~20 sPer million requests: 1,000 attacks produce 900 flags, and 999,000 benign prompts produce about 9,990 flags. So roughly 10,890 flags contain 900 real attacks, about 8%. Around eleven in twelve flags are benign. Report that share, because a rate measured on an all-attack corpus never reveals it.
solid answer
~60 sThe two rates you measured are both conditional — one on the prompt being an attack, one on it being benign. Neither answers the question anyone operating the system actually asks: **when this thing fires, how often is it right?** That answer depends on prevalence, and at one in a thousand the benign population is so much larger that its small error rate dominates. Working it per million requests: 900 true flags against about 9,990 false ones, so roughly 8% of flags are genuine. Halving the benign flag rate to 0.5% roughly doubles that share; the catch rate barely moves it. That asymmetry is the useful finding, and it is invisible in any report that only quotes performance on attack prompts. In the write-up, state the prevalence you assumed, give the flag composition alongside the detection rate, and note that everything downstream of the flag — anything a human or a second system does with it — inherits this composition. Sizing that downstream process is somebody else's decision; surfacing the number it depends on is yours.
go deeper
Recognises that most traffic is benign so most flags will be benign, even without doing the arithmetic.
Does the per-million calculation correctly and names prevalence as the reason the two conditional rates do not answer the question.
Identifies the benign error rate as the dominant lever at low prevalence, insists the assumed prevalence is stated, and questions whether the benign sample resembles real traffic.
Turns the mixture into a recommendation about what mode the layer runs in and what evidence would be needed to change it, while leaving downstream process design to its owner.
### Three different rates, and only one of them answers the operator's question A prompt-screening classifier — Llama Guard, Prompt Guard, ShieldGemma, a hosted moderation endpoint — emits a score per prompt, thresholded into a flag. Evaluating it produces two rates, and they point in opposite directions: - **Catch rate (recall)**, here 90%: `P(flagged | attack)`. Measured on the attack corpus. - **Benign flag rate (false positive rate)**, here 1%: `P(flagged | benign)`. Measured on a hard-negative set of benign prompts. Both are conditional on knowing what the prompt *was*. The person operating the system never knows that. They see a flag and must decide what to do with it, so the quantity they need is the reverse conditional — `P(attack | flagged)`, the share of flags that are genuine. That number is not a property of the classifier at all. It is a property of the classifier **and** the population it runs against, and the population is set by prevalence. ### The arithmetic Per 1,000,000 requests at 0.1% prevalence: | | attack (1,000) | benign (999,000) | |---|---|---| | flagged | 900 | 9,990 | | not flagged | 100 | 989,010 | Flags: 10,890, of which 900 are genuine — about **8.3%**. Roughly eleven flags in twelve are benign prompts from ordinary users. Nothing in the all-attack evaluation could have revealed this, because the benign majority that supplies 9,990 of those flags was never in the corpus. ### What it costs, and which lever moves it The operating cost of this layer is 10,890 flags per million requests, whatever the organisation does with a flag — hard block, soft challenge, log-only, or a queue. That volume figure, and its composition, is what a stakeholder needs before choosing the mode; how the downstream review process is then sized and staffed belongs to whoever owns that process, not to the assessment. The arithmetic also exposes an ordering that is counter-intuitive until you see it. From this operating point: - Raising the catch rate 90% → 95% adds ~50 true flags to 900, against ~9,990 false ones. Genuine share moves 8.3% → about 8.7%. Essentially nothing. - Halving the benign flag rate 1% → 0.5% removes ~5,000 of the false flags. Genuine share moves 8.3% → about 15%. Roughly a doubling. At low prevalence the benign error rate is the dominant lever, because it is multiplied by the enormous benign population while the catch rate is multiplied by the tiny attack one. A recommendation to "improve detection" is, at this operating point, a recommendation that changes almost nothing an operator will feel. ### Where the number misleads **A flattering benign set.** If the hard-negative set is made of obviously innocuous prompts — weather questions, greetings — the 1% is fiction. Real benign traffic contains security questions, fiction writing, medical and legal queries, translations, and quoted text: exactly the material that looks adversarial to a screen. Measure the benign rate on language your actual users produce, or the whole table is optimistic. **A single point estimate for a moving population.** The composition scales with prevalence and nothing else has to change. At 1-in-100 the same classifier yields ~9,000 true and ~9,900 false flags, about 48% genuine. Under an attack burst prevalence rises and flags become *more* trustworthy — which is the opposite of most people's intuition and matters when an alerting threshold was tuned at quiet-hours prevalence. **Reading the catch rate as precision.** The classic error is answering "90%". It is the most common wrong answer in interviews and in reports, and it is off by roughly a factor of eleven here. **A threshold that was moved.** Both rates come from one operating point on a score curve. Whoever tunes the threshold moves both, in opposite directions, and can therefore choose the number your report quotes. State the threshold, or report across the range. ### What to check Does the benign sample resemble real user language, and how large was it — 1% of a 200-prompt benign set is two prompts, which pins nothing. Was prevalence estimated from labelled traffic or assumed? Is the threshold stated? And does the flag composition sit next to the detection rate in the deliverable, rather than in a separate section a reader can skip while keeping the reassuring 90%?
- The team wants to raise the catch rate from 90% to 95%. What does that do to the share of flags that are genuine?Almost nothing: true flags go from 900 to 950 while false flags stay near 9,990, moving the genuine share from about 8.3% to about 8.7%. At this prevalence the benign error rate is the lever that matters.
- Prevalence turns out to be 1 in 100, not 1 in 1,000. What changes?Per million: 9,000 true flags and about 9,900 false ones, so roughly 48% of flags are genuine. The classifier did not change; only the population it runs against did.
An airport metal detector with a 1% false alarm rate goes off far more often for belt buckles than for weapons, simply because almost everyone walking through is wearing a belt and almost nobody is carrying a weapon.
saying these in an interview costs you the question
- Answering 90% — reading the catch rate as the share of flags that are correct
- Doing the arithmetic but never stating the prevalence it assumed
- Recommending a higher catch rate as the fix at this prevalence
- Measuring the benign error rate on prompts nobody would ever mistake for an attack