skip to content

A red-team run against a customer-facing chat assistant produced an attack-success rate of roughly one attempt in five on a jailbreak suite. The governance workstream wants that entered in the AI risk register as a control marked pass or fail. What is lost in that flattening, and how do you record it so the pass or fail is defensible later?

level: middleimportance: must knowfreq 58%

answer

  1. rate over a chosen suite, not real traffic
  2. denominator, mix, probabilistic control
  3. threshold is a decision, write it down
  4. judge error lives inside the rate
  5. per-family sub-rates in the evidence record

basics

~20 s

You lose the denominator, the mix of attacks that produced the rate, and the fact that the control is probabilistic rather than on or off. Record the rate, the suite it came from, the number of attempts, the target configuration and the date beside the verdict, and state the threshold that turned the rate into pass or fail.

solid answer

~50 s

A register expects binary controls; a red-team result is a rate over a chosen population of attempts. Three things disappear in the conversion. First the **denominator** — one in five of what? Of a suite you selected, which is not a sample of real user traffic. Second the **mix**: a rate averaged over easy and hard attack families hides that the dangerous families may have succeeded far more often. Third the **probabilistic nature** of the control itself: the same suite re-run tomorrow gives a different number, so 'fail' is a threshold decision, not an observation. So write the verdict *and* its basis: rate, attempt count, which suite, the target and guard configuration, the mechanism that decided a response counted as a success, the threshold applied, and who set it. If the threshold is not written down, next quarter's team cannot tell whether the number moved or the standard did.

go deeper

for a junior

Recognises that a percentage is not a yes or no, and that the number of attempts and the suite used have to be recorded next to the verdict.

for a middle

Explains the denominator problem, breaks the rate out by attack family, and insists the pass threshold is written down and owned.

for a senior

Adds measurement error — the judge or detector has its own false-positive and false-negative rates — and defines what a comparable re-test requires before any trend is claimed.

for a principal

Sets the standard: who owns thresholds, how often evidence is refreshed, and how control-effectiveness measurements feed risk decisions without being laundered into likelihoods.

### The verdict is allowed to be binary Governance frameworks exist to compress. Refusing to give a pass or fail does not preserve nuance; it moves the judgement to somebody with less context, usually in the room, usually without the caveats. The job is to make the verdict **traceable**, not to resist it. ### What the rate actually is, mechanically An attack-success rate is successes divided by attempts on a suite that you or a tool author chose. Each term has an owner. The *attempts* are the cross product the harness generated: in garak, the probe classes selected by `--probes` times the prompts each contributes times `--generations` completions per prompt; in promptfoo, the `redteam` config's `plugins` times its `strategies` times `numTests`. The *successes* are whatever a mechanism said counted — a garak detector, a PyRIT scorer, a promptfoo grader, or a model acting as judge. So "one attempt in five" is a statement about a population you constructed, adjudicated by a component with its own error rate. ### What the run costs Worth stating in the entry, because it bounds how often the number can be refreshed. A 2,000-attempt suite against a metered hosted endpoint is 2,000 generation calls; if success is adjudicated by an LLM judge, that is another 2,000 calls, so the judged run is roughly double the naive estimate. At commodity per-call prices that is tens of dollars of inference — cheap — but provider rate limits, not money, set the wall clock, and a suite that is trivially affordable can still take hours to complete and a day of engineer time to triage. The consequence for governance is direct: a measurement that costs a day cannot be re-run on demand at a committee meeting, so the entry has to stand on the record you wrote at the time. ### Where the number misleads *The denominator.* One in five **of what**? Of an adversarial suite, not of user traffic. Multiplying it into a risk score as a likelihood is the single most common misuse: it converts "the control leaks under deliberate attack" into "one in five conversations goes wrong", which is a different and far larger claim. *The mix.* An overall rate averages easy and hard families. One in five that is nearly all low-consequence refusal bypasses is a different control state from one in five concentrated in exfiltration attempts through the retrieval path. Break the rate out per family before flattening it; the sub-rates are what you show when the verdict is challenged. *The judge's own error, which lives inside the rate.* This is the part most teams never quantify. If a judge has a false-positive rate on benign responses and a false-negative rate on genuine successes, the observed rate is not the true one: ``` observed = true x (1 - FNR) + (1 - true) x FPR # true 15%, FNR 10%, FPR 8% # 0.15 x 0.90 + 0.85 x 0.08 = 0.135 + 0.068 = 0.203 ``` An observed 20% is consistent with a true 15% and a mildly over-firing judge. If the pass threshold sat at 18%, the control failed on judge error. Nothing about the model changed. *Run-to-run noise.* Sampling on the target, provider-side change, and non-deterministic judging all move the figure. A verdict at a threshold the measurement's own spread straddles is a coin toss with a document attached. ### What you would check before writing the verdict Hand-label a random sample of scored attempts — 50 to 100 is usually enough to see a badly calibrated judge — and report the judge's measured false-positive and false-negative rates beside the attack-success rate. Re-run the suite unchanged at least once and record the spread, so you know whether a difference between two quarters could be noise. Confirm the suite and its size were pinned, that the target configuration including any guard was captured, and that the success criterion is written in words a different team could apply. ### The minimum evidence record behind one pass/fail cell Suite identity and size, attempt count, target and guard configuration, success criterion and the judge's measured error, the overall rate, the per-family rates, the threshold, who owns the threshold, and the date. Ten fields behind one cell. That is the honest cost of the conversion, and it is what makes next quarter's re-test a comparison instead of a fresh argument. If the threshold is not written down, nobody can later tell whether the number moved or the standard did.

  • Next quarter the rate moves from one in five to one in eight. What must be true before you call that an improvement?
    Same suite, same size, comparable target and guard configuration, same success criterion, and enough attempts that the difference exceeds run-to-run noise. Otherwise you have measured a change in the measurement.
  • Governance asks you to convert the rate into a likelihood value for a risk score. What do you say?
    That the rate is conditional on an adversary running that suite, not on ordinary usage. Offer it as a control-effectiveness input and let exposure and attacker motivation be modelled separately, rather than passing it off as a likelihood.
  • Who should set the pass threshold?
    The risk owner, informed by the red team — not the red team alone. The team supplies the measurement and its uncertainty; accepting the residual risk is an accountable business decision.

Reporting a jailbreak suite's success rate as the chance of a production jailbreak is like quoting a car's crash-test score as your odds of crashing. The number describes how the thing holds up under a deliberate, standardised assault, not how often that assault happens.

saying these in an interview costs you the question

  • Reporting the rate as the probability a real user will succeed.
  • Averaging over attack families and losing that the severe ones succeeded most.
  • No record of the suite, attempt count or success criterion behind the verdict.
  • Treating the pass threshold as a technical fact instead of an owned risk decision.
  • Comparing this quarter's rate with last quarter's after the suite or the target changed.

context