skip to content

Attack Success Rate

Attack-success rate is a fraction whose denominator every suite fixed differently, so two published rates rarely belong on one axis. Interviewers use it to see whether you read a number or quote it.

on this pageshow

explore

questions

5

A red-team benchmark reports an attack-success rate (ASR) for a model. What are the numerator and the denominator of that fraction, and what must the results table state before the number can be read at all?

level: juniorimportance: must knowfreq 78%

answer

  1. hits over what
  2. behaviour list vs attempt vs pair
  3. any-of aggregation rule
  4. attack method is part of the metric
  5. raw weights vs deployed endpoint

basics

~20 s

The numerator is the count of attempts the suite judged harmful; the denominator is the unit the suite chose to count over. That unit may be behaviours on a fixed list, individual attempts, or behaviour-and-attack pairs. Before reading the rate you need the behaviour list, the attempts per item, the attack method, and who ruled an attempt a hit.

solid answer

~50 s

ASR is a fraction, and almost all of the meaning sits in how the fraction was formed: - **Numerator** — attempts (or behaviours) that some judging procedure ruled a success. A string match on refusal phrases, a trained classifier and a judge model give different counts on identical transcripts. - **Denominator** — the counting unit. Three common ones: *per behaviour* on a fixed list, *per attempt* across all samples, and *per behaviour-times-attack-method* cell. - **Aggregation** — whether a behaviour counts once it falls to any attempt or any method, or is averaged. So a bare percentage is unreadable. Demand the behaviour list and its size, attempts per behaviour, the attack method(s) held fixed, the decoding settings, the target wrapper (raw weights or a deployed endpoint with its system prompt), and the judging procedure. Change any of those and the same model produces a different rate with no change in its safety.

go deeper

for a junior

Should say ASR is hits divided by attempts or behaviours and ask which one before quoting it.

for a middle

Should name the denominator options, note that the attack method and the judging procedure are part of the metric, and explain why the same model yields different rates.

for a senior

Should insist on the full tuple behind the number and point out that a suite usually targets raw weights rather than the deployed stack.

for a principal

Should argue for reporting ASR as a tuple across an organisation and for treating per-category breakdowns, not the scalar, as the decision input.

An attack-success rate (ASR) is a measurement produced by an instrument, never a property the model carries around with it. Reading one means reconstructing the instrument. ### The fraction written longhand `ASR = hits / opportunities`. Neither half is fixed by the name. The **numerator** is the count of model responses that some *ruling procedure* declared a win for the attacker. The ruling procedure is a concrete piece of software: a substring match against a list of refusal phrases ("I can't help with that"), a trained harm classifier, or a judge model prompted with a rubric. These disagree on identical transcripts, so the numerator is a property of the judge as much as of the model. The **denominator** is whatever the suite decided one *opportunity* is. Three are in common use: | denominator | one unit is | the question it answers | |---|---|---| | per behaviour | one item on a fixed harmful-behaviour list, counted broken if any sampled attempt lands | how much of the list is reachable | | per attempt | one generation against one item | how often an attack lands | | per behaviour x method cell | one item under one named attack method | which method beats which category | The same transcripts yield three different percentages under these three rules, with no change to the model. ### The four knobs behind any published rate 1. **The item list.** HarmBench, JailbreakBench and AdvBench each ship their own curated behaviour list, and they differ in size (roughly a hundred behaviours for JailbreakBench's list, several hundred for HarmBench and AdvBench), in category mix, and in difficulty. TrustLLM aggregates across several dimensions again. A different list is a different population, not a second reading of one quantity. 2. **Attempts per item.** One generation per behaviour, or twenty-five with the behaviour marked broken on the first hit, are different statistics from the same harness. 3. **The attack held fixed.** Every ASR is the rate *of some method*. A single-turn template and an optimisation-based or multi-turn method are different instruments pointed at the same list. 4. **The target configuration.** Raw open weights with a plain chat template answer differently from the same weights behind a product system prompt, a retrieval corpus and an output filter. Benchmarks usually measure the former; decoding settings (temperature, top-p, max tokens) belong here too. ### What a run costs Multiply it out: `behaviours x attempts x methods` target generations, plus roughly one judge call per transcript when the judge is itself a model. A 400-behaviour list at 25 samples under three methods is 30,000 generations and about 30,000 judge calls — on a metered chat API at a few dollars per million tokens and a thousand-odd tokens per exchange, tens to low hundreds of dollars and several hours of wall clock, mostly bounded by rate limits rather than compute. Two lines dominate and surprise people: the judge roughly doubles the bill, and an optimisation-based attack that searches per behaviour replaces the API line item with GPU-hours, which is where a benchmark sweep goes from a coffee-break job to a weekend one. ### Where the number misleads The single most common error is reading the percentage as a model property and carrying it across contexts. Three specific misreadings follow. First, a low rate is read as safety when it may be the mark of a permissive-to-refuse model that also declines benign requests — the attack suite alone never shows that cost. Second, a rate is compared against another suite's rate as though the denominators agreed. Third, re-scoring the archived transcripts with a stricter classifier moves the number, which is proof that part of what you are reading is the judge's error floor, not the model's behaviour. And an ASR of zero means only that nothing on *that* list, under *that* method, at *that* budget, was ruled a hit. ### What you would check before believing it Open the methodology section and the harness config, not the abstract, and fill in the tuple: list identity and size, attempts per item, aggregation rule, attack method and budget, target wrapper and decoding settings, ruling procedure. If the paper publishes per-item results, recompute the headline from them — the recomputation catches aggregation-rule surprises immediately. Then hand-read fifteen or twenty transcripts the judge labelled hits and a similar number it labelled misses; the disagreement rate you find is the floor under every digit of the published figure. Finally, note what was actually tested: if it was raw weights and you ship a product, the number belongs in a model-selection conversation, not a product-risk one.

  • The same transcripts are re-scored with a stricter harm classifier. Does the ASR change?
    Yes. The numerator is whatever the ruling procedure counted, so re-scoring the identical transcripts moves the rate without any change to the model or the attack.
  • Is an ASR of zero on a suite evidence the model cannot be jailbroken?
    No. It says that on that behaviour list, with that attack method, that many attempts and that judge, nothing was ruled a hit. It says nothing about behaviours or methods outside the list.
  • Where in a benchmark run is the most actionable information, if not the headline rate?
    In the per-category and per-method breakdown, plus the failing transcripts, because those point at a specific defence gap rather than a single scalar.

saying these in an interview costs you the question

  • Quoting an ASR percentage without being able to say what the denominator counted
  • Treating ASR as a property of the model rather than of model plus list plus attack plus judge
  • Assuming every suite counts one attempt per behaviour
  • Ignoring that the target may be raw weights rather than the deployed product
  • Believing a lower ASR always means a safer system

context

open as a page

You have one attack-success rate published by HarmBench and one published by JailbreakBench for the same open-weights model. What has to line up before you can put both figures in a single comparison table, and what usually does not?

level: middleimportance: must knowfreq 66%

basics

~20 s

Both rates must count over the same population and be produced the same way: the same behaviour list, the same attempts per behaviour and aggregation rule, the same attack method, the same target configuration, the same procedure for ruling a hit. Independently published suites share none of these, so their headline rates belong on different axes.

open as a page

One jailbreak suite marks a behaviour as broken if any of n sampled attempts is ruled harmful and reports the fraction of behaviours broken; another reports the fraction of individual attempts ruled harmful. How do those two rates relate, and can you convert between them?

level: middleimportance: should knowfreq 52%

basics

~20 s

They are different statistics. The any-of-n behaviour rate is never lower than the per-attempt rate on the same transcripts, and it rises as n rises even though the model is unchanged. You cannot convert one to the other from the headline numbers, because the conversion needs the per-behaviour distribution of hits, which the aggregate discards.

open as a page

A stakeholder points at a published safety-benchmark attack-success rate for the base model behind your product and asks why your own red-team scan of the deployed assistant produced a very different rate. What does the published figure's denominator actually enumerate, and why do the two numbers not sit on one axis?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The published denominator enumerates that suite's fixed behaviour list, attacked by that suite's method, against the model as the suite configured it - usually raw weights with a plain chat template, no product system prompt, no retrieval, no filters. Your scan enumerates your own attack set against the deployed stack. Different population, target and ruling.

open as a page

Your organisation wants to track one safety benchmark's attack-success rate as a quarterly figure across model upgrades. What has to be frozen for that series to mean anything, and what silently breaks it?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Freeze everything except the model: the pinned item list, the attack method and its budget, attempts per item and the aggregation rule, decoding settings, the target wrapper, and the exact judging procedure. Silent breakers are a refreshed item list, an updated judge, a changed default sampling setting, and the list leaking into training data.

open as a page