A red-team benchmark reports an attack-success rate (ASR) for a model. What are the numerator and the denominator of that fraction, and what must the results table state before the number can be read at all?
answer
- hits over what
- behaviour list vs attempt vs pair
- any-of aggregation rule
- attack method is part of the metric
- raw weights vs deployed endpoint
basics
~20 sThe numerator is the count of attempts the suite judged harmful; the denominator is the unit the suite chose to count over. That unit may be behaviours on a fixed list, individual attempts, or behaviour-and-attack pairs. Before reading the rate you need the behaviour list, the attempts per item, the attack method, and who ruled an attempt a hit.
solid answer
~50 sASR is a fraction, and almost all of the meaning sits in how the fraction was formed: - **Numerator** — attempts (or behaviours) that some judging procedure ruled a success. A string match on refusal phrases, a trained classifier and a judge model give different counts on identical transcripts. - **Denominator** — the counting unit. Three common ones: *per behaviour* on a fixed list, *per attempt* across all samples, and *per behaviour-times-attack-method* cell. - **Aggregation** — whether a behaviour counts once it falls to any attempt or any method, or is averaged. So a bare percentage is unreadable. Demand the behaviour list and its size, attempts per behaviour, the attack method(s) held fixed, the decoding settings, the target wrapper (raw weights or a deployed endpoint with its system prompt), and the judging procedure. Change any of those and the same model produces a different rate with no change in its safety.
go deeper
Should say ASR is hits divided by attempts or behaviours and ask which one before quoting it.
Should name the denominator options, note that the attack method and the judging procedure are part of the metric, and explain why the same model yields different rates.
Should insist on the full tuple behind the number and point out that a suite usually targets raw weights rather than the deployed stack.
Should argue for reporting ASR as a tuple across an organisation and for treating per-category breakdowns, not the scalar, as the decision input.
An attack-success rate (ASR) is a measurement produced by an instrument, never a property the model carries around with it. Reading one means reconstructing the instrument. ### The fraction written longhand `ASR = hits / opportunities`. Neither half is fixed by the name. The **numerator** is the count of model responses that some *ruling procedure* declared a win for the attacker. The ruling procedure is a concrete piece of software: a substring match against a list of refusal phrases ("I can't help with that"), a trained harm classifier, or a judge model prompted with a rubric. These disagree on identical transcripts, so the numerator is a property of the judge as much as of the model. The **denominator** is whatever the suite decided one *opportunity* is. Three are in common use: | denominator | one unit is | the question it answers | |---|---|---| | per behaviour | one item on a fixed harmful-behaviour list, counted broken if any sampled attempt lands | how much of the list is reachable | | per attempt | one generation against one item | how often an attack lands | | per behaviour x method cell | one item under one named attack method | which method beats which category | The same transcripts yield three different percentages under these three rules, with no change to the model. ### The four knobs behind any published rate 1. **The item list.** HarmBench, JailbreakBench and AdvBench each ship their own curated behaviour list, and they differ in size (roughly a hundred behaviours for JailbreakBench's list, several hundred for HarmBench and AdvBench), in category mix, and in difficulty. TrustLLM aggregates across several dimensions again. A different list is a different population, not a second reading of one quantity. 2. **Attempts per item.** One generation per behaviour, or twenty-five with the behaviour marked broken on the first hit, are different statistics from the same harness. 3. **The attack held fixed.** Every ASR is the rate *of some method*. A single-turn template and an optimisation-based or multi-turn method are different instruments pointed at the same list. 4. **The target configuration.** Raw open weights with a plain chat template answer differently from the same weights behind a product system prompt, a retrieval corpus and an output filter. Benchmarks usually measure the former; decoding settings (temperature, top-p, max tokens) belong here too. ### What a run costs Multiply it out: `behaviours x attempts x methods` target generations, plus roughly one judge call per transcript when the judge is itself a model. A 400-behaviour list at 25 samples under three methods is 30,000 generations and about 30,000 judge calls — on a metered chat API at a few dollars per million tokens and a thousand-odd tokens per exchange, tens to low hundreds of dollars and several hours of wall clock, mostly bounded by rate limits rather than compute. Two lines dominate and surprise people: the judge roughly doubles the bill, and an optimisation-based attack that searches per behaviour replaces the API line item with GPU-hours, which is where a benchmark sweep goes from a coffee-break job to a weekend one. ### Where the number misleads The single most common error is reading the percentage as a model property and carrying it across contexts. Three specific misreadings follow. First, a low rate is read as safety when it may be the mark of a permissive-to-refuse model that also declines benign requests — the attack suite alone never shows that cost. Second, a rate is compared against another suite's rate as though the denominators agreed. Third, re-scoring the archived transcripts with a stricter classifier moves the number, which is proof that part of what you are reading is the judge's error floor, not the model's behaviour. And an ASR of zero means only that nothing on *that* list, under *that* method, at *that* budget, was ruled a hit. ### What you would check before believing it Open the methodology section and the harness config, not the abstract, and fill in the tuple: list identity and size, attempts per item, aggregation rule, attack method and budget, target wrapper and decoding settings, ruling procedure. If the paper publishes per-item results, recompute the headline from them — the recomputation catches aggregation-rule surprises immediately. Then hand-read fifteen or twenty transcripts the judge labelled hits and a similar number it labelled misses; the disagreement rate you find is the floor under every digit of the published figure. Finally, note what was actually tested: if it was raw weights and you ship a product, the number belongs in a model-selection conversation, not a product-risk one.
- The same transcripts are re-scored with a stricter harm classifier. Does the ASR change?Yes. The numerator is whatever the ruling procedure counted, so re-scoring the identical transcripts moves the rate without any change to the model or the attack.
- Is an ASR of zero on a suite evidence the model cannot be jailbroken?No. It says that on that behaviour list, with that attack method, that many attempts and that judge, nothing was ruled a hit. It says nothing about behaviours or methods outside the list.
- Where in a benchmark run is the most actionable information, if not the headline rate?In the per-category and per-method breakdown, plus the failing transcripts, because those point at a specific defence gap rather than a single scalar.
saying these in an interview costs you the question
- Quoting an ASR percentage without being able to say what the denominator counted
- Treating ASR as a property of the model rather than of model plus list plus attack plus judge
- Assuming every suite counts one attempt per behaviour
- Ignoring that the target may be raw weights rather than the deployed product
- Believing a lower ASR always means a safer system