skip to content

One jailbreak suite marks a behaviour as broken if any of n sampled attempts is ruled harmful and reports the fraction of behaviours broken; another reports the fraction of individual attempts ruled harmful. How do those two rates relate, and can you convert between them?

level: middleimportance: should knowfreq 52%

answer

  1. union over samples is monotone in n
  2. behaviours-broken vs attempts-hit
  3. reachability vs typicality
  4. no conversion without per-item counts
  5. n must be printed beside the rate

basics

~20 s

They are different statistics. The any-of-n behaviour rate is never lower than the per-attempt rate on the same transcripts, and it rises as n rises even though the model is unchanged. You cannot convert one to the other from the headline numbers, because the conversion needs the per-behaviour distribution of hits, which the aggregate discards.

solid answer

~60 s

Write both out. Per-attempt: hits divided by all attempts. Any-of-n per behaviour: behaviours with at least one hit, divided by behaviours. The any-of rate is weakly greater, and it is monotone in n: adding samples can only turn a behaviour from unbroken to broken, never back. So the same model, same attack, same judge yields a bigger headline as the sampling budget grows. Conversion needs information the aggregate threw away. If every behaviour had the same per-attempt hit probability p, the any-of-n rate would be 1 minus (1 minus p) to the n. Real behaviour lists are strongly heterogeneous — a few items fall on nearly every sample, most never fall — so that uniform assumption badly misstates the mapping in both directions. To convert honestly you need the per-behaviour hit counts, not the summary. The same trap appears on the method axis: a suite that marks a behaviour broken if any of several attack methods breaks it reports a union, and the union grows as methods are added.

go deeper

for a junior

Should notice the two rates count different things and not treat them as interchangeable.

for a middle

Should state that the any-of rate is weakly greater and monotone in the budget, and that conversion needs per-behaviour hit counts.

for a senior

Should extend the union framing to attack methods, insist the budget is printed beside the rate, and describe reporting the hit distribution rather than only the aggregate.

for a principal

Should set which of the two the organisation treats as canonical for reachability claims and require the budget in every quoted figure.

Two quantities are both called attack-success rate, and reports get built on top of the confusion. ### Two denominators, one name - **Per attempt.** Numerator: attempts the judge ruled harmful. Denominator: every attempt sent, i.e. `behaviours x attempts`. This measures **typicality** — how often the attack lands. - **Any-of-n per behaviour.** Numerator: behaviours with at least one ruled-harmful attempt among n samples. Denominator: the behaviours on the list. This measures **reachability** — how much of the list an attacker willing to retry can reach under a stated budget. They answer different questions. Abuse-rate and user-exposure reasoning wants the first; a worst-case coverage claim wants the second. ### Why the union framing is monotone ``` any_of_rate(n) = |{ b in B : exists i <= n with hit(b, i) }| / |B| ``` The set in the numerator can only grow as n grows: a further sample can flip a behaviour from unbroken to broken and never back. So the any-of-n rate is weakly greater than the per-attempt rate on the same transcripts and non-decreasing in n. Nothing about the model changed; the budget did. The identical structure appears on the method axis — a suite that marks a behaviour broken if *any* of several attack methods breaks it is reporting a union that grows every time a method is added to the harness. ### Why you cannot convert between them If every behaviour shared one per-attempt hit probability p and samples were independent, then `any_of_rate(n) = 1 - (1 - p)^n`, and the two would be interchangeable. Real behaviour lists violate both premises hard. They are strongly **heterogeneous**: a handful of items fall on nearly every sample while most never fall, so a low per-attempt rate can coexist with a high or a low any-of rate depending entirely on which items the hits sit in. And samples are often **correlated** — an adaptive attack conditions on earlier failures, a multi-turn loop carries state, a shared seed template makes attempts near-duplicates. Converting honestly needs the per-behaviour hit counts, which the aggregate discards. A published headline simply does not contain the information. ### What the budget costs, and what it buys n is a linear multiplier on the whole sweep: `behaviours x n` generations plus a matching number of judging calls. Going from 1 to 25 attempts per behaviour is 25x the spend and 25x the wall clock on a rate-limited endpoint. What that money buys is sharply diminishing — the rate climbs fastest over the first few samples, because the easy items fall immediately, and later samples mostly re-break behaviours already counted. The practical consequence is that budget and headline are traded against each other: a team that wants a bigger number buys more samples, and a team on a tight bill reports n=1 and a flattering zero. ### Where the number misleads An any-of rate quoted without n is an unbounded claim — it can be raised arbitrarily by sampling more, so it is not comparable to any other row in any table. Two specific misreadings follow. A rise between two quarters or two vendors is read as the model getting weaker when the sampling budget grew. And a 0% figure at n=1 is read as robustness when it is close to uninformative; the same list at n=25 may well be substantially non-zero. In the other direction, a per-attempt rate is read as coverage: 4% of attempts landing sounds small, but if all of those hits are concentrated in six behaviours that fall every single time, the reachable surface is six behaviours, permanently. ### What you would check Find the aggregation rule and n in the methodology, not the abstract — this is the single most frequently omitted pair of facts in a results table. Ask for, or reconstruct from raw logs, the per-behaviour hit counts, and look at the shape of that distribution rather than only its mean: a list where a few items fall constantly is indistinguishable in the aggregate from one where every item falls occasionally, and the two demand completely different defences. If you hold the logs, plot the any-of rate as a function of n and publish the curve — it makes the budget dependence visible and it is the honest artefact. Then check that any comparison you are drawing holds both n and the method set fixed across rows.

  • Which of the two rates would you put in a defence-prioritisation report, and why?
    Usually both, labelled: the any-of-n rate with n stated shows what an attacker willing to retry can reach, while the per-attempt rate shows how often it lands, and defence effort differs between a list where a few items fall constantly and one where many fall rarely.
  • Under what assumption could you map a per-attempt rate p to an any-of-n rate?
    Only if every behaviour shared the same p and samples were independent, giving 1 minus (1 minus p) to the n. Real lists are heterogeneous and adaptive attacks correlate samples, so the mapping misstates both directions.

Rolling dice: 'what fraction of attempts came up six' stays near one in six however long you roll, but 'what fraction of dice ever showed a six' climbs towards 100% purely by rolling more. An any-of-n rate quoted without n is that second number with the number of rolls hidden.

saying these in an interview costs you the question

  • Comparing an any-of-n rate with a per-attempt rate as if they were one metric
  • Quoting an any-of rate without stating the budget n or the method set
  • Assuming a single hit probability per behaviour so the two can be converted
  • Reading a rise in the any-of rate as the model getting weaker when the budget grew
  • Averaging per-attempt and per-behaviour rates together

context