skip to content

A red-team suite reports a 5% attack-success rate. The harm judge that decided which replies counted as successful attacks measured 8% false positives and 12% false negatives on a labelled sample. Why can the 5% not be read as "5% of attempts really succeeded", and what can you honestly state instead?

level: middleimportance: must knowfreq 60%

answer

  1. observed = p(1-FNR) + (1-p)FPR
  2. floor at FPR, ceiling at 1-FNR
  3. 5% under an 8% floor
  4. error rates are distribution-specific
  5. minimum resolvable difference

basics

~20 s

Judge error sets a floor. With an 8% false-positive rate, even a model that never complies would score around 8% flagged hits, so a 5% reading sits inside the judge's noise. Report the judge, its measured error rates, and adjudicate a sample before claiming any true rate.

solid answer

~60 s

The published rate is the **judge's flag rate**, not the true success rate. If the true rate is `p`, what you observe is roughly `p*(1 - FNR) + (1 - p)*FPR`. Put in 8% false positives and 12% false negatives: a model that never complies still reads about 8%, and a model that always complies reads about 88%. The whole scale is compressed into that band, and **5% is below the bottom of it** — the observation is not even consistent with a true rate, which means the judge's error rates were estimated on a different distribution than the one you just ran. What you can state: the raw count of judge-flagged attempts; the judge artifact and how its error rates were measured; a hand-adjudicated sample of the flagged transcripts with an exact interval on the confirmed hits; and, if you want a corrected point estimate, the inverted form `(observed - FPR) / (1 - FPR - FNR)` with an interval that propagates uncertainty in both error rates. Below roughly the false-positive rate, report counts and adjudications, not a rate.

go deeper

for a junior

Should say the reported number is what the judge flagged, and that the judge itself makes mistakes, so it is not the true success rate.

for a middle

Writes down the relationship between the flag rate and the true rate, points out the false-positive floor, and notices 5% falls below it.

for a senior

Adds that the judge's error rates are distribution-dependent and probably were not measured on this run, proposes stratified hand-adjudication, and sets a minimum resolvable difference for comparisons.

for a principal

Treats the judge's measured error as part of the published metric's definition and refuses to let any comparison or trend be reported below the instrument's resolution.

**The arithmetic, stated once.** Let the harm judge — the classifier or judge model that labels each transcript hit or not — have false-positive rate `a` (probability it flags a reply that did not succeed) and false-negative rate `b` (probability it clears a reply that did). Apply it to attempts whose unknown true success rate is `p`. The rate it flags is ``` q = p(1 - b) + (1 - p)a ``` true successes it catches, plus non-successes it wrongly flags. Two consequences follow immediately, and nearly every misreading of a published red-team number is one of them. **Consequence one: the scale has a floor and a ceiling.** As `p` sweeps from 0 to 1, `q` sweeps only from `a` to `1 - b` — here from 0.08 to 0.88. Judge error does not merely add symmetric noise around the truth; it compresses and shifts the entire measurable range. A model that never complies still reads 8%. So a reported 5% is not a small true rate: it is *below the floor of the instrument*, an observation the model above cannot produce for any value of `p`. **Consequence two: you can invert, with conditions.** Solving for `p` gives the standard prevalence correction `p̂ = (q - a) / (1 - a - b)`. It is unbiased only if `a` and `b` genuinely hold on the distribution you just scored, and its variance grows without bound as `a + b` approaches 1 — the denominator is how much signal the judge has left after its errors. Near the floor the correction also happily returns negative estimates, which is the arithmetic telling you the inputs are inconsistent, not a result. **So what does a 5% under an 8% floor actually mean?** One of two things, both findings about the instrument rather than the model. Either the judge's error rates do not transfer to this run — overwhelmingly the common case, because they are typically measured on the benchmark's own reference completions while you are scoring a different model, a different attack family and often unusual output formats — or the labelled sample used to estimate them was too small to distinguish 8% from 3%. A ninety-item labelled set gives an 8% estimate with an interval several points wide in each direction; nobody reports that interval, and the point estimate then gets treated as a constant of nature. **What this costs to do properly.** The correction itself is free. The expensive part is earning `a` and `b` on the right distribution: hand-adjudicating a stratified sample is roughly two to four minutes per transcript for a trained reviewer, so a 200-transcript audit is about a person-day, and you owe it per attack family rather than once for the suite, because a judge's operating point is distribution-dependent. Score a run of heavily obfuscated replies and `b` rises sharply; score a run full of degenerate output and `a` rises. Budget the audit as a recurring line item against the run, not as a one-off validation. **What propagates downstream.** Every figure computed from these labels inherits the floor — per-behaviour breakdowns, model-versus-model comparisons, quarter-over-quarter deltas. Two models at 5% and 7% are indistinguishable under an 8%-false-positive judge, and a two-point quarterly movement is not a movement. Derive a minimum resolvable difference from `a`, `b` and the sample size, publish it, and refuse to narrate changes below it. The commonest bad summary in this space — "the judge is about 90% accurate, so the number is roughly right" — collapses two error rates that act on different populations into one figure and hides which direction the bias runs. **What I would deliver instead of a bare 5%.** The flagged count and its denominator, not just the ratio. The judge's identity and version. Measured `a` and `b`, with the size and provenance of the sample they came from. A stratified hand-adjudication of both flagged and unflagged transcripts, with an exact binomial interval on the confirmed hits. And, if a point estimate is wanted at all, the corrected value with an interval that propagates uncertainty in both error rates, labelled as corrected. Below roughly the false-positive rate, the honest output is confirmed counts and adjudications, not a rate.

  • Where do a harm judge's false-positive and false-negative rates usually come from, and why might they not apply to your run?
    From a labelled sample, usually the benchmark's own reference completions. Your run scores a different model, a different attack family and often odd formats, so the judge's operating point shifts — re-estimate per attack family.
  • Two models score 5% and 7% under the same judge. Can you say one is safer?
    No. Both sit at or under the judge's false-positive floor, so the gap is inside the instrument's error. Adjudicate the flagged transcripts by hand and compare confirmed hit counts instead.
  • Which error rate matters more when the true success rate is very low?
    The false-positive rate. At low prevalence almost every flag is a false one, so precision collapses even for a judge with a respectable overall accuracy.

A judge with an 8% false-positive rate is a thermometer that reads 8 degrees in an ice bath. Every reading it gives you starts from there, so a reported 5 tells you something is wrong with the thermometer, not that the water is unusually cold.

saying these in an interview costs you the question

  • Reads the flag rate as the true rate and never mentions that the judge has an error rate at all.
  • Applies the prevalence correction mechanically without asking whether the error rates were measured on this distribution.
  • Says 'the judge is about 90% accurate so the number is roughly right' — a single accuracy figure hides which direction the errors go.
  • Reports quarter-over-quarter movements smaller than the judge's false-positive rate as real change.

context