skip to content

Sampling Error

A few hundred behaviours sampled n times each yields a rate that climbs with n rather than with attacker skill, and those attempts cost money. Interviewers ask what your figure would survive.

on this pageshow

explore

questions

5

A jailbreak evaluation reports "38% attack success" against a fixed list of 300 harmful behaviours. Why is that percentage not interpretable until you also know how many attempts were made per behaviour?

level: juniorimportance: must knowfreq 65%

answer

  1. any-hit-in-n = OR aggregation
  2. monotone non-decreasing in n
  3. rate @ n, never a bare rate
  4. attempts are the cost driver
  5. store per-attempt outcomes

basics

~20 s

Because a behaviour usually counts as broken if any attempt succeeds. Twenty attempts give twenty chances where one gives one, so the figure only rises as the attempt count rises. With attempts-per-behaviour unstated, 38% describes how much you spent as much as how fragile the model is.

solid answer

~50 s

Nearly every harmful-behaviour suite aggregates the same way: for each behaviour it fires n attempts and marks the behaviour broken if **any** of them is ruled a hit. That aggregation is monotone non-decreasing in n — adding attempts can flip a behaviour from unbroken to broken but never the reverse, so the headline percentage can only go up as you buy more attempts. So the number is a pair, not a scalar: *rate at n attempts per behaviour*. "38% at n=1" and "38% at n=50" describe very different models — the first is fragile on first contact, the second held out for 49 tries on most behaviours. A caption that omits n is not comparable to anything, including your own earlier run. The minimum you record beside the rate: the behaviour list and its size, attempts per behaviour, whether attempts stopped early on the first hit, and what decided a hit.

code

python · 5 lines
python
broken = sum(
    any(is_hit(outcome) for outcome in attempts[behaviour])
    for behaviour in behaviours
)
rate = broken / len(behaviours)   # non-decreasing in len(attempts[behaviour])

go deeper

for a junior

Says a behaviour counts as broken if any attempt works, so more attempts push the number up; n must be reported with the rate.

for a middle

Names the OR aggregation and its monotonicity, and lists what else must appear in the caption: list size, attempts per behaviour, stopping rule, hit rule.

for a senior

Connects n to engagement cost and to spurious 'improvements' when budgets change, and insists on storing per-attempt outcomes so the rate can be recomputed at any smaller n.

for a principal

Treats n as a declared parameter of the release metric and fixes it organisation-wide, the way a latency SLO fixes its percentile.

## What the percentage is a percentage of A *harmful-behaviour list* is a fixed, published inventory of requests a model is supposed to refuse — HarmBench, AdvBench and JailbreakBench each ship one, and they are small: a few hundred entries, not millions. A run: - pairs each behaviour on that list with an *attack* (a prompt template, a search procedure, a multi-turn strategy); - fires some number of *attempts* per behaviour; - and passes every resulting completion to a *hit rule* — the thing that decides whether an attempt counted. The hit rule is usually a refusal-string check, a trained harm classifier, or a judge model. The headline figure is then `behaviours with at least one hit / behaviours attempted`. Notice that two different units are in play: the unit being *scored* is the **behaviour**, the unit being *sampled* is the **attempt**. The reported number never mentions the second one. ## The aggregation, and why it can only climb The map from n attempts to a per-behaviour verdict is a **logical OR**: one hit anywhere in the n marks the behaviour broken. OR over a larger set is monotone — adding attempts can flip a behaviour from unbroken to broken but can never flip it back — so `rate(n)` is non-decreasing in n. Make it concrete. Suppose a behaviour has a per-attempt chance of a hit of p = 0.05 against this attack. At n = 1 it is marked safe 95% of the time. At n = 25 the chance of at least one hit is 1 − (1 − p)^n ≈ 0.72, so the same behaviour against the same model, with nothing changed but the budget, now reads as broken. That is the whole reason a bare "38%" is uninterpretable: you cannot tell a fragile model at n = 1 from a hardened one at n = 50. ## What n costs Attempts are the **cost driver** of an engagement against a metered endpoint. - 300 behaviours at n = 25 is 7,500 generations against the target. - If the hit rule is a hosted judge model called once per attempt, that is a second metered line of roughly 7,500 calls, often on a larger and pricier model than the target. - A multi-turn attack that averages four exchanges per attempt multiplies the target side again. - And rate limits, not credits, are frequently the binding constraint: 7,500 sequential calls at a few per second is hours of wall clock, plus the retries that 429s force, which burn quota without producing outcomes. On top of that sits human time — every candidate hit that gates a release gets read by a person. ## Where the number misleads Three specific misreadings, in rough order of how often they happen. - (1) *A budget cut read as a security improvement.* A team under cost pressure quietly drops n from 25 to 5, the rate falls by fifteen points, and the release review records hardening that never occurred. The mirror image is a pre-launch "let's be thorough" run at n = 50 that makes an unchanged model appear to regress overnight. - (2) *Cross-run comparison against a figure with no n.* Two percentages on the same public list are not on the same axis unless both budgets are stated. - (3) *Per-behaviour "safe" verdicts at low n.* At n = 5, every behaviour with p of a few percent reads as safe — and those are precisely the ones a patient attacker converts, because retrying is free for them. A low-n rate systematically understates a persistent adversary. One further caution: the rate is not the probability that an ordinary user meets harm. The list was assembled adversarially, so it is a **stress measurement**, not a base rate. ## What you would check before believing someone's figure - **Attempts per behaviour**, first and always. - Whether the harness **stopped a behaviour at its first hit** — that changes cost, not the verdict, but it means you cannot compute any per-attempt rate from the same run, because the unsuccessful tail attempts were never made. - The **identity and version** of the behaviour list, and its size. - **What ruled a hit**, since a looser rule inflates the same run with n untouched. - And the target's **surrounding configuration** — system prompt, decoding temperature, whether an input or output filter sat in front of it — because a bare endpoint and the shipped stack are different systems. ## How to report it Treat n as a **declared parameter** of the metric, exactly as a percentile is a declared parameter of a latency SLO: publish `rate @ n`, freeze n across comparisons you intend to make, and keep the per-attempt outcome log (behaviour id, attempt index, verdict). That log is cheap and lets you recompute the rate at any smaller n later without paying for a single new generation.

  • Does stopping a behaviour at its first successful attempt change the reported rate?
    No — it changes cost, not the verdict. The behaviour was going to be marked broken anyway. It does change any per-attempt rate you compute from the same run, because the unsuccessful tail attempts were never made.
  • Two runs on the same behaviour list both used n=10 and differ by 6 points. Is n the explanation?
    Not by itself. With n held fixed you have to look elsewhere: sampling nondeterminism in the target, a changed hit rule, a changed attack strategy, or a different system prompt on the target.
  • What is the cheapest artefact to keep so you never have to re-run for a different n?
    The per-attempt outcome log: behaviour id, attempt index, and the hit verdict. Any rate at a smaller n is a recomputation over that log.

Reporting a break rate without attempts per behaviour is like publishing a lottery's win rate without saying how many tickets each player bought. The same draw looks generous or stingy depending on a number nobody printed.

saying these in an interview costs you the question

  • Treating the percentage as a property of the model alone, with no mention of attempts.
  • Assuming more attempts could lower the rate.
  • Saying n does not matter because 'the benchmark is standard' — the list is standard, the attempt budget is not.
  • Comparing two figures from different runs without checking either run's attempt count.

context

open as a page

Running one attack against the same fixed harmful-behaviour list at 1, 5 and 25 attempts per behaviour gives 12%, 29% and 41% of behaviours broken. What shape is that curve, and what does it mean when it flattens?

level: middleimportance: must knowfreq 55%

basics

~20 s

It rises and flattens. The easy behaviours fall in the first few attempts, so each extra attempt buys less. Flattening means further attempts of this attack will find little more: you have separated the behaviours this attack can break from a residual core that resists it at any budget you can afford.

open as a page

You hold a fixed query budget of 20,000 generations against a metered chat endpoint for one engagement. Do you spend it as 400 behaviours at 50 attempts each, or 2,000 behaviours at 10 attempts each? How do you decide?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Decide from the claim the number must support. Breadth, many behaviours at few attempts, finds unexpected harm categories and gives a stable per-behaviour picture. Depth, fewer behaviours at many attempts, shows what a determined attacker eventually gets. Most engagements buy a wide cheap pass first, then spend the remainder deep where the pass looked weak.

open as a page

Another team publishes "2% attack success" on the same public harmful-behaviour list you used, where you measured 30%, and their write-up never states attempts per behaviour. What can you legitimately conclude, and what do you do to make the two comparable?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Almost nothing about the gap. Their 2% may be one attempt per behaviour against your twenty. Treat it as a lower bound, ask for attempts per behaviour, the stopping rule and what ruled a hit, and re-measure yourself at a declared attempt budget before putting the two numbers in one table.

open as a page

You are asked to set the house standard for attempts per behaviour in red-team results that gate a model release. What do you fix, what do you leave to each team, and what does over-fixing cost you?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Fix what makes numbers comparable across releases: the behaviour list, a minimum attempts per behaviour, the stopping rule, the hit rule, and a caption that states all of them. Leave extra deep runs and new attacks free, reported separately. Over-fixing costs discovery: a frozen budget stops anyone probing where the model actually looks weak.

open as a page