skip to content

What the Number Is

One headline rate hides a denominator, a judge and a resampling choice, and each of those moves it further than any defence does. Interviewers probe whether you can take a published figure apart.

on this pageshow

explore

questions

19

A red-team benchmark reports an attack-success rate (ASR) for a model. What are the numerator and the denominator of that fraction, and what must the results table state before the number can be read at all?

level: juniorimportance: must knowfreq 78%

answer

  1. hits over what
  2. behaviour list vs attempt vs pair
  3. any-of aggregation rule
  4. attack method is part of the metric
  5. raw weights vs deployed endpoint

basics

~20 s

The numerator is the count of attempts the suite judged harmful; the denominator is the unit the suite chose to count over. That unit may be behaviours on a fixed list, individual attempts, or behaviour-and-attack pairs. Before reading the rate you need the behaviour list, the attempts per item, the attack method, and who ruled an attempt a hit.

solid answer

~50 s

ASR is a fraction, and almost all of the meaning sits in how the fraction was formed: - **Numerator** — attempts (or behaviours) that some judging procedure ruled a success. A string match on refusal phrases, a trained classifier and a judge model give different counts on identical transcripts. - **Denominator** — the counting unit. Three common ones: *per behaviour* on a fixed list, *per attempt* across all samples, and *per behaviour-times-attack-method* cell. - **Aggregation** — whether a behaviour counts once it falls to any attempt or any method, or is averaged. So a bare percentage is unreadable. Demand the behaviour list and its size, attempts per behaviour, the attack method(s) held fixed, the decoding settings, the target wrapper (raw weights or a deployed endpoint with its system prompt), and the judging procedure. Change any of those and the same model produces a different rate with no change in its safety.

go deeper

for a junior

Should say ASR is hits divided by attempts or behaviours and ask which one before quoting it.

for a middle

Should name the denominator options, note that the attack method and the judging procedure are part of the metric, and explain why the same model yields different rates.

for a senior

Should insist on the full tuple behind the number and point out that a suite usually targets raw weights rather than the deployed stack.

for a principal

Should argue for reporting ASR as a tuple across an organisation and for treating per-category breakdowns, not the scalar, as the decision input.

An attack-success rate (ASR) is a measurement produced by an instrument, never a property the model carries around with it. Reading one means reconstructing the instrument. ### The fraction written longhand `ASR = hits / opportunities`. Neither half is fixed by the name. The **numerator** is the count of model responses that some *ruling procedure* declared a win for the attacker. The ruling procedure is a concrete piece of software: a substring match against a list of refusal phrases ("I can't help with that"), a trained harm classifier, or a judge model prompted with a rubric. These disagree on identical transcripts, so the numerator is a property of the judge as much as of the model. The **denominator** is whatever the suite decided one *opportunity* is. Three are in common use: | denominator | one unit is | the question it answers | |---|---|---| | per behaviour | one item on a fixed harmful-behaviour list, counted broken if any sampled attempt lands | how much of the list is reachable | | per attempt | one generation against one item | how often an attack lands | | per behaviour x method cell | one item under one named attack method | which method beats which category | The same transcripts yield three different percentages under these three rules, with no change to the model. ### The four knobs behind any published rate 1. **The item list.** HarmBench, JailbreakBench and AdvBench each ship their own curated behaviour list, and they differ in size (roughly a hundred behaviours for JailbreakBench's list, several hundred for HarmBench and AdvBench), in category mix, and in difficulty. TrustLLM aggregates across several dimensions again. A different list is a different population, not a second reading of one quantity. 2. **Attempts per item.** One generation per behaviour, or twenty-five with the behaviour marked broken on the first hit, are different statistics from the same harness. 3. **The attack held fixed.** Every ASR is the rate *of some method*. A single-turn template and an optimisation-based or multi-turn method are different instruments pointed at the same list. 4. **The target configuration.** Raw open weights with a plain chat template answer differently from the same weights behind a product system prompt, a retrieval corpus and an output filter. Benchmarks usually measure the former; decoding settings (temperature, top-p, max tokens) belong here too. ### What a run costs Multiply it out: `behaviours x attempts x methods` target generations, plus roughly one judge call per transcript when the judge is itself a model. A 400-behaviour list at 25 samples under three methods is 30,000 generations and about 30,000 judge calls — on a metered chat API at a few dollars per million tokens and a thousand-odd tokens per exchange, tens to low hundreds of dollars and several hours of wall clock, mostly bounded by rate limits rather than compute. Two lines dominate and surprise people: the judge roughly doubles the bill, and an optimisation-based attack that searches per behaviour replaces the API line item with GPU-hours, which is where a benchmark sweep goes from a coffee-break job to a weekend one. ### Where the number misleads The single most common error is reading the percentage as a model property and carrying it across contexts. Three specific misreadings follow. First, a low rate is read as safety when it may be the mark of a permissive-to-refuse model that also declines benign requests — the attack suite alone never shows that cost. Second, a rate is compared against another suite's rate as though the denominators agreed. Third, re-scoring the archived transcripts with a stricter classifier moves the number, which is proof that part of what you are reading is the judge's error floor, not the model's behaviour. And an ASR of zero means only that nothing on *that* list, under *that* method, at *that* budget, was ruled a hit. ### What you would check before believing it Open the methodology section and the harness config, not the abstract, and fill in the tuple: list identity and size, attempts per item, aggregation rule, attack method and budget, target wrapper and decoding settings, ruling procedure. If the paper publishes per-item results, recompute the headline from them — the recomputation catches aggregation-rule surprises immediately. Then hand-read fifteen or twenty transcripts the judge labelled hits and a similar number it labelled misses; the disagreement rate you find is the floor under every digit of the published figure. Finally, note what was actually tested: if it was raw weights and you ship a product, the number belongs in a model-selection conversation, not a product-risk one.

  • The same transcripts are re-scored with a stricter harm classifier. Does the ASR change?
    Yes. The numerator is whatever the ruling procedure counted, so re-scoring the identical transcripts moves the rate without any change to the model or the attack.
  • Is an ASR of zero on a suite evidence the model cannot be jailbroken?
    No. It says that on that behaviour list, with that attack method, that many attempts and that judge, nothing was ruled a hit. It says nothing about behaviours or methods outside the list.
  • Where in a benchmark run is the most actionable information, if not the headline rate?
    In the per-category and per-method breakdown, plus the failing transcripts, because those point at a specific defence gap rather than a single scalar.

saying these in an interview costs you the question

  • Quoting an ASR percentage without being able to say what the denominator counted
  • Treating ASR as a property of the model rather than of model plus list plus attack plus judge
  • Assuming every suite counts one attempt per behaviour
  • Ignoring that the target may be raw weights rather than the deployed product
  • Believing a lower ASR always means a safer system

context

open as a page

A jailbreak benchmark harness rules an attempt successful when the model's reply does not contain any phrase from a fixed refusal-string list ("I'm sorry", "I cannot", "As an AI"). What errors does this substring rule push into the attack-success rate it produces?

level: juniorimportance: must knowfreq 70%

basics

~20 s

It counts by wording, not content. A reply that apologises then complies is scored a refusal; an empty, off-topic or garbled reply with no listed phrase is scored a success. Unlisted refusal wordings, other languages and paraphrases all leak through, so the rate drifts both up and down.

open as a page

A red-team report says a model scored 0% attack success on a jailbreak benchmark. Why is that number alone not evidence the model is good, and what second measurement belongs beside it?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A model that refuses everything scores zero on any attack suite, so zero can mean hardened or mean useless. The attack rate only reads next to a benign-refusal rate: run a set of harmless prompts, many of them phrased to look sensitive, and report how many were refused. Publish both numbers together.

open as a page

A jailbreak evaluation reports "38% attack success" against a fixed list of 300 harmful behaviours. Why is that percentage not interpretable until you also know how many attempts were made per behaviour?

level: juniorimportance: must knowfreq 65%

basics

~20 s

Because a behaviour usually counts as broken if any attempt succeeds. Twenty attempts give twenty chances where one gives one, so the figure only rises as the attempt count rises. With attempts-per-behaviour unstated, 38% describes how much you spent as much as how fragile the model is.

open as a page

You have one attack-success rate published by HarmBench and one published by JailbreakBench for the same open-weights model. What has to line up before you can put both figures in a single comparison table, and what usually does not?

level: middleimportance: must knowfreq 66%

basics

~20 s

Both rates must count over the same population and be produced the same way: the same behaviour list, the same attempts per behaviour and aggregation rule, the same attack method, the same target configuration, the same procedure for ruling a hit. Independently published suites share none of these, so their headline rates belong on different axes.

open as a page

A red-team suite reports a 5% attack-success rate. The harm judge that decided which replies counted as successful attacks measured 8% false positives and 12% false negatives on a labelled sample. Why can the 5% not be read as "5% of attempts really succeeded", and what can you honestly state instead?

level: middleimportance: must knowfreq 60%

basics

~20 s

Judge error sets a floor. With an 8% false-positive rate, even a model that never complies would score around 8% flagged hits, so a 5% reading sits inside the judge's noise. Report the judge, its measured error rates, and adjudicate a sample before claiming any true rate.

open as a page

Running one attack against the same fixed harmful-behaviour list at 1, 5 and 25 attempts per behaviour gives 12%, 29% and 41% of behaviours broken. What shape is that curve, and what does it mean when it flattens?

level: middleimportance: must knowfreq 55%

basics

~20 s

It rises and flattens. The easy behaviours fall in the first few attempts, so each extra attempt buys less. Flattening means further attempts of this attack will find little more: you have separated the behaviours this attack can break from a residual core that resists it at any budget you can afford.

open as a page

Two teams publish attack-success rates for the same open-weights model on the same behaviour list. One scored replies with the trained harm classifier the benchmark distributes; the other scored them with a hosted chat model prompted with a harm rubric. The rates differ by 14 points. How do you work out whether the model or the scoring choice explains the gap, and what do you require before putting both numbers in one table?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Hold the transcripts fixed and vary the scorer. Get both teams' raw completions and run both judges over both sets. If each judge gives similar numbers on either set, the gap is the scoring choice, not the model. Without archived completions the two rates cannot share a table.

open as a page

One jailbreak suite marks a behaviour as broken if any of n sampled attempts is ruled harmful and reports the fraction of behaviours broken; another reports the fraction of individual attempts ruled harmful. How do those two rates relate, and can you convert between them?

level: middleimportance: should knowfreq 52%

basics

~20 s

They are different statistics. The any-of-n behaviour rate is never lower than the per-attempt rate on the same transcripts, and it rises as n rises even though the model is unchanged. You cannot convert one to the other from the headline numbers, because the conversion needs the per-behaviour distribution of hits, which the aggregate discards.

open as a page

When you score a set of harmless prompts to measure a model's over-refusal, what should count as a refusal, and why is that label not simply the mirror image of scoring a hit on the attack side?

level: middleimportance: should knowfreq 48%

basics

~20 s

Not just the flat "I can't help with that". Count deflections, moralising non-answers, and replies that answer a safer question than the one asked. The attack side has a concrete target — did the harmful content appear. The benign side has no such artefact, so you are grading whether the user's actual request was served.

open as a page

A stakeholder points at a published safety-benchmark attack-success rate for the base model behind your product and asks why your own red-team scan of the deployed assistant produced a very different rate. What does the published figure's denominator actually enumerate, and why do the two numbers not sit on one axis?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The published denominator enumerates that suite's fixed behaviour list, attacked by that suite's method, against the model as the suite configured it - usually raw weights with a plain chat template, no product system prompt, no retrieval, no filters. Your scan enumerates your own attack set against the deployed stack. Different population, target and ruling.

open as a page

When a judge model scores red-team transcripts to decide which attempts landed, the text it reads contains the attacker's own prompt and the target's reply. Why does that make the judge itself an attack surface, and how would you check whether your published attack-success rate has been distorted by it?

level: seniorimportance: should knowfreq 35%

basics

~20 s

The judge reads attacker-controlled text, so the transcript is untrusted input to it. Payloads can address whatever reads them next and push a scoring decision, and encoded or obfuscated replies fall outside a trained classifier's training data. Hand-adjudicate a random sample of both scored-hit and scored-miss transcripts.

open as a page

A safety-tuned build drops your attack-success rate on the same jailbreak suite from 22% to 6%. What do you measure, and how do you design the comparison, to show that this is real hardening rather than the model simply refusing more of everything?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Re-run the frozen benign prompt set against both builds and compare refusal rates. Real hardening leaves the benign rate flat while the attack rate falls; a blanket refusal shift moves both together. Hold the prompt sets, system prompt, decoding settings and labelling rubric identical, and change only the build.

open as a page

You hold a fixed query budget of 20,000 generations against a metered chat endpoint for one engagement. Do you spend it as 400 behaviours at 50 attempts each, or 2,000 behaviours at 10 attempts each? How do you decide?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Decide from the claim the number must support. Breadth, many behaviours at few attempts, finds unexpected harm categories and gives a stable per-behaviour picture. Depth, fewer behaviours at many attempts, shows what a determined attacker eventually gets. Most engagements buy a wide cheap pass first, then spend the remainder deep where the pass looked weak.

open as a page

Another team publishes "2% attack success" on the same public harmful-behaviour list you used, where you measured 30%, and their write-up never states attempts per behaviour. What can you legitimately conclude, and what do you do to make the two comparable?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Almost nothing about the gap. Their 2% may be one attempt per behaviour against your twenty. Treat it as a lower bound, ask for attempts per behaviour, the stopping rule and what ruled a hit, and re-measure yourself at a declared attempt budget before putting the two numbers in one table.

open as a page

Your organisation publishes a quarterly safety attack-success rate for every model it ships, computed by an automated harm judge, and better judges keep appearing. How do you decide when to change the judge, and what do you owe readers of the historical series?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat the judge as part of the metric definition. Pin a versioned judge for the published series, archive every raw completion, and when you upgrade, re-score the whole history and publish both series over an overlap window. Never swap silently; state the judge and its measured error rates.

open as a page

Two candidate builds give you an attack-success rate on a red-team suite and a refusal rate on a benign prompt set, and between the builds the two rates move in opposite directions. As the red team, how do you report that pair so nobody is misled, and what claims do you refuse to make from it?

level: principalimportance: should knowfreq 27%

basics

~20 s

Report both rates per build with their prompt-set sizes, labelling rules and run configuration, and state the direction each moved. Refuse to combine them into one index, to declare a winner, and to imply the two rates are in the same units. Add the qualitative slice: which legitimate requests the safer build now declines.

open as a page

Your organisation wants to track one safety benchmark's attack-success rate as a quarterly figure across model upgrades. What has to be frozen for that series to mean anything, and what silently breaks it?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Freeze everything except the model: the pinned item list, the attack method and its budget, attempts per item and the aggregation rule, decoding settings, the target wrapper, and the exact judging procedure. Silent breakers are a refreshed item list, an updated judge, a changed default sampling setting, and the list leaking into training data.

open as a page

You are asked to set the house standard for attempts per behaviour in red-team results that gate a model release. What do you fix, what do you leave to each team, and what does over-fixing cost you?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Fix what makes numbers comparable across releases: the behaviour list, a minimum attempts per behaviour, the stopping rule, the hit rule, and a caption that states all of them. Leave extra deep runs and new attacks free, reported separately. Over-fixing costs discovery: a frozen budget stops anyone probing where the model actually looks weak.

open as a page