A jailbreak evaluation reports "38% attack success" against a fixed list of 300 harmful behaviours. Why is that percentage not interpretable until you also know how many attempts were made per behaviour?
answer
- any-hit-in-n = OR aggregation
- monotone non-decreasing in n
- rate @ n, never a bare rate
- attempts are the cost driver
- store per-attempt outcomes
basics
~20 sBecause a behaviour usually counts as broken if any attempt succeeds. Twenty attempts give twenty chances where one gives one, so the figure only rises as the attempt count rises. With attempts-per-behaviour unstated, 38% describes how much you spent as much as how fragile the model is.
solid answer
~50 sNearly every harmful-behaviour suite aggregates the same way: for each behaviour it fires n attempts and marks the behaviour broken if **any** of them is ruled a hit. That aggregation is monotone non-decreasing in n — adding attempts can flip a behaviour from unbroken to broken but never the reverse, so the headline percentage can only go up as you buy more attempts. So the number is a pair, not a scalar: *rate at n attempts per behaviour*. "38% at n=1" and "38% at n=50" describe very different models — the first is fragile on first contact, the second held out for 49 tries on most behaviours. A caption that omits n is not comparable to anything, including your own earlier run. The minimum you record beside the rate: the behaviour list and its size, attempts per behaviour, whether attempts stopped early on the first hit, and what decided a hit.
code
python · 5 linesbroken = sum(
any(is_hit(outcome) for outcome in attempts[behaviour])
for behaviour in behaviours
)
rate = broken / len(behaviours) # non-decreasing in len(attempts[behaviour])go deeper
Says a behaviour counts as broken if any attempt works, so more attempts push the number up; n must be reported with the rate.
Names the OR aggregation and its monotonicity, and lists what else must appear in the caption: list size, attempts per behaviour, stopping rule, hit rule.
Connects n to engagement cost and to spurious 'improvements' when budgets change, and insists on storing per-attempt outcomes so the rate can be recomputed at any smaller n.
Treats n as a declared parameter of the release metric and fixes it organisation-wide, the way a latency SLO fixes its percentile.
## What the percentage is a percentage of A *harmful-behaviour list* is a fixed, published inventory of requests a model is supposed to refuse — HarmBench, AdvBench and JailbreakBench each ship one, and they are small: a few hundred entries, not millions. A run: - pairs each behaviour on that list with an *attack* (a prompt template, a search procedure, a multi-turn strategy); - fires some number of *attempts* per behaviour; - and passes every resulting completion to a *hit rule* — the thing that decides whether an attempt counted. The hit rule is usually a refusal-string check, a trained harm classifier, or a judge model. The headline figure is then `behaviours with at least one hit / behaviours attempted`. Notice that two different units are in play: the unit being *scored* is the **behaviour**, the unit being *sampled* is the **attempt**. The reported number never mentions the second one. ## The aggregation, and why it can only climb The map from n attempts to a per-behaviour verdict is a **logical OR**: one hit anywhere in the n marks the behaviour broken. OR over a larger set is monotone — adding attempts can flip a behaviour from unbroken to broken but can never flip it back — so `rate(n)` is non-decreasing in n. Make it concrete. Suppose a behaviour has a per-attempt chance of a hit of p = 0.05 against this attack. At n = 1 it is marked safe 95% of the time. At n = 25 the chance of at least one hit is 1 − (1 − p)^n ≈ 0.72, so the same behaviour against the same model, with nothing changed but the budget, now reads as broken. That is the whole reason a bare "38%" is uninterpretable: you cannot tell a fragile model at n = 1 from a hardened one at n = 50. ## What n costs Attempts are the **cost driver** of an engagement against a metered endpoint. - 300 behaviours at n = 25 is 7,500 generations against the target. - If the hit rule is a hosted judge model called once per attempt, that is a second metered line of roughly 7,500 calls, often on a larger and pricier model than the target. - A multi-turn attack that averages four exchanges per attempt multiplies the target side again. - And rate limits, not credits, are frequently the binding constraint: 7,500 sequential calls at a few per second is hours of wall clock, plus the retries that 429s force, which burn quota without producing outcomes. On top of that sits human time — every candidate hit that gates a release gets read by a person. ## Where the number misleads Three specific misreadings, in rough order of how often they happen. - (1) *A budget cut read as a security improvement.* A team under cost pressure quietly drops n from 25 to 5, the rate falls by fifteen points, and the release review records hardening that never occurred. The mirror image is a pre-launch "let's be thorough" run at n = 50 that makes an unchanged model appear to regress overnight. - (2) *Cross-run comparison against a figure with no n.* Two percentages on the same public list are not on the same axis unless both budgets are stated. - (3) *Per-behaviour "safe" verdicts at low n.* At n = 5, every behaviour with p of a few percent reads as safe — and those are precisely the ones a patient attacker converts, because retrying is free for them. A low-n rate systematically understates a persistent adversary. One further caution: the rate is not the probability that an ordinary user meets harm. The list was assembled adversarially, so it is a **stress measurement**, not a base rate. ## What you would check before believing someone's figure - **Attempts per behaviour**, first and always. - Whether the harness **stopped a behaviour at its first hit** — that changes cost, not the verdict, but it means you cannot compute any per-attempt rate from the same run, because the unsuccessful tail attempts were never made. - The **identity and version** of the behaviour list, and its size. - **What ruled a hit**, since a looser rule inflates the same run with n untouched. - And the target's **surrounding configuration** — system prompt, decoding temperature, whether an input or output filter sat in front of it — because a bare endpoint and the shipped stack are different systems. ## How to report it Treat n as a **declared parameter** of the metric, exactly as a percentile is a declared parameter of a latency SLO: publish `rate @ n`, freeze n across comparisons you intend to make, and keep the per-attempt outcome log (behaviour id, attempt index, verdict). That log is cheap and lets you recompute the rate at any smaller n later without paying for a single new generation.
- Does stopping a behaviour at its first successful attempt change the reported rate?No — it changes cost, not the verdict. The behaviour was going to be marked broken anyway. It does change any per-attempt rate you compute from the same run, because the unsuccessful tail attempts were never made.
- Two runs on the same behaviour list both used n=10 and differ by 6 points. Is n the explanation?Not by itself. With n held fixed you have to look elsewhere: sampling nondeterminism in the target, a changed hit rule, a changed attack strategy, or a different system prompt on the target.
- What is the cheapest artefact to keep so you never have to re-run for a different n?The per-attempt outcome log: behaviour id, attempt index, and the hit verdict. Any rate at a smaller n is a recomputation over that log.
Reporting a break rate without attempts per behaviour is like publishing a lottery's win rate without saying how many tickets each player bought. The same draw looks generous or stingy depending on a number nobody printed.
saying these in an interview costs you the question
- Treating the percentage as a property of the model alone, with no mention of attempts.
- Assuming more attempts could lower the rate.
- Saying n does not matter because 'the benchmark is standard' — the list is standard, the attempt budget is not.
- Comparing two figures from different runs without checking either run's attempt count.