Running one attack against the same fixed harmful-behaviour list at 1, 5 and 25 attempts per behaviour gives 12%, 29% and 41% of behaviours broken. What shape is that curve, and what does it mean when it flattens?
answer
- saturating, concave, non-decreasing
- easy behaviours harvested first
- plateau bounds the attack, not the model
- stop when a doubling buys < delta
- one sweep at max n, truncate to get the curve
basics
~20 sIt rises and flattens. The easy behaviours fall in the first few attempts, so each extra attempt buys less. Flattening means further attempts of this attack will find little more: you have separated the behaviours this attack can break from a residual core that resists it at any budget you can afford.
solid answer
~50 sThe curve is non-decreasing and concave — a saturating curve, not a line. Behaviours are not equally hard: the ones with a high per-attempt chance of a hit are almost all found by attempt three or four, and after that each new attempt is spent on the hard remainder, so the marginal yield drops. The plateau height is the practical ceiling of *this attack against this target*, and the residual is the interesting artefact: which behaviours never fell, and why. The operational reading is a stopping rule. Pick n where the last doubling of the budget moved the rate by less than you care about — if 25 attempts buys two points over 12, buying 50 is a poor use of an engagement's queries. The trap is comparing a saturated number with an unsaturated one. A 41%-at-25 figure and a 29%-at-5 figure describe the same target; only the budget differed.
go deeper
Recognises the curve goes up and levels off, and that more attempts eventually stop helping.
Explains the concavity through heterogeneous per-behaviour difficulty, and uses the flattening point as a stopping rule for the attempt budget.
Gets the whole curve from one logged sweep by truncation, knows when correlated multi-turn attempts break that trick, and labels a still-climbing figure as a lower bound.
Uses the shape as a signal about the list itself — a saturated curve means retire or harden the behaviour list rather than fund more attempts.
## What is being plotted Fix one attack, one target, one hit rule and one behaviour list. For each attempt budget n, compute `behaviours with at least one hit in their first n attempts / behaviours on the list`. Plotting that against n gives the *best-of-n curve*. In the figures given — 12% at n = 1, 29% at n = 5, 41% at n = 25 — the budget grew 25-fold while the rate grew a little over threefold. That gap is the whole story. ## Why it is concave Behaviours are not equally hard. Each has its own **per-attempt chance of a hit**, p, against this attack: - a handful sit near p = 0.5 and fall on the first or second try; - a middle band sits at a few percent and needs several; - and a tail sits near zero and resists any budget you can buy. Best-of-n harvests them in that order. For a single behaviour, the chance of at least one hit in n attempts is 1 − (1 − p)^n; the aggregate curve is that expression mixed over the distribution of p across the list. Each expression is concave in n, and a mixture of concave curves is concave, so the aggregate rises fast and then bends. (The mathematics of that expression, and any interval around the plateau, belong to statistics. What the curve buys a red teamer is a *purchasing decision*.) ## What the plateau means — and what it does not The ceiling bounds *this attack against this target under this hit rule*. It is not a robustness property of the model. Swap the attack strategy, loosen the hit rule, change the system prompt or put an output filter in front, and you get a different curve on the same list. So the reportable object is `attack A, target T, hit rule J, list L: 41% at n = 25, +2 points from n = 12 to n = 25` — the **delta over the last doubling** is as informative as the level. ## The stopping rule The operational use of the shape is deciding when to stop buying attempts. Pick the n past which one more doubling of the budget moves the rate by less than the smallest difference you would act on. If going from 12 to 25 attempts bought two points and your release threshold moves in five-point steps, 50 attempts is a poor use of an engagement's queries — that money buys more behaviours, a second attack, or triage time instead. ## Running it cheaply Do not run one sweep per value of n. Run a **single sweep** at the largest budget you intend to buy, log every attempt outcome, and recompute the rate at 1, 2, 4, 8, 16 by truncating each behaviour's attempt list to its first k. Three separate sweeps at n = 1, 5, 25 cost 31 attempts per behaviour where one logged sweep at 25 costs 25 — on a 300-behaviour list that is 9,300 generations instead of 7,500, plus the matching judge calls, for information you already had. The truncation trick has one precondition: attempts within a behaviour must be **exchangeable**. All of these break it: - a multi-turn strategy that conditions on earlier refusals; - an attack that adapts using feedback from previous attempts; - or a cached deterministic decode. The first k attempts of an adaptive run are not a fair n = k sample, and if you truncate anyway you must say so in the caption. ## Where the number misleads A curve still **climbing steeply** at your stopping point means your headline is a budget artefact, not a measurement of the target; it must be labelled a lower bound, with the observed slope quoted, because a better-funded attacker will publish a higher number on the same list tomorrow. The opposite failure is a curve that **flattens close to 100%**: that is not a triumphant attack result, it is a dead benchmark. The list has stopped discriminating between targets, so every model you test will score the same and the metric can no longer detect a regression. The fix there is harder behaviours or a stricter hit rule, not more attempts. Both readings are invisible from a single point — which is why quoting one number without the shape around it is the core failure this leaf is about. ## What you would check - That the curve was derived from one logged sweep rather than several inconsistent ones; - that attempts were exchangeable if truncation was used; - that the plateau is quoted with the attack and hit rule attached; - that the last-doubling delta is stated; - and that nothing about the target changed mid-sweep — a silently updated endpoint or a changed system prompt turns a saturation curve into a chart of two different systems.
- You want the rate at n = 1, 2, 4, 8, 16. How many sweeps do you pay for?One, at n=16, with every attempt outcome logged; the smaller-n rates are truncations of that log. This only holds when attempts within a behaviour are independent.
- The curve is still climbing steeply at your budget ceiling. What do you write in the report?That the figure is a lower bound at that attempt budget, with the observed slope quoted, and that a better-funded attacker would report a higher number on the same list.
- The curve flattens just under 100%. Is that a good result for the attacker?It is a bad result for the benchmark: the list no longer separates targets. Report it and move to harder behaviours rather than buying more attempts.
saying these in an interview costs you the question
- Expecting a straight line, or expecting the rate to reach 100% given enough attempts.
- Reading the plateau as a property of the target rather than of this attack against this target.
- Re-running the whole sweep once per n instead of truncating a single logged sweep.
- Reporting a still-climbing point as a final robustness figure with no lower-bound caveat.
- Truncating a multi-turn run that conditions on earlier attempts and calling it a fair smaller-n sample.