How do you report attack-success rate so it doesn't overstate LLM safety?
answer
- a ratio is only as honest as its denominator
- per-prompt versus per-attempt
- attach the attacker's budget
- best-of-N turns 0% into double digits
- report the curve, slice by severity
basics
~20 sAttack-success rate is meaningless without its denominator and the attacker's budget. Report success per attempt, not per prompt, state how many attempts and whether the attacker adapted, and publish the curve of success against attempts rather than one number.
solid answer
~40 sAttack-success rate (ASR) is the fraction of attacks that achieved the attacker's goal, and every honest report has to say what an "attack" was. Three things determine the number: the **denominator** (one attempt per prompt, or many resampled attempts), the **budget** (how many tries the attacker got), and whether the attacker could **adapt** between tries. A suite of 200 prompts can pass at 100% single-shot while best-of-32 resampling on the very same prompts reaches roughly 41% success, because an attacker only needs one of the 32 to land. The right artefact is a curve — ASR against attempts, broken out by severity category — plus a statement of who judged success and how. A single headline percentage with no budget attached should be treated as an unfalsifiable claim.
code
python · 11 linesdef asr_per_attempt(trials):
return sum(1 for t in trials if t) / len(trials)
def asr_best_of_n(trials_by_case):
# attacker wins a case if ANY attempt in its budget succeeded
wins = sum(1 for attempts in trials_by_case if any(attempts))
return wins / len(trials_by_case)
cases = [[False] * 31 + [True], [False] * 32, [False] * 30 + [True, False]]
print(asr_per_attempt([t for c in cases for t in c])) # 0.0208...
print(asr_best_of_n(cases)) # 0.666...go deeper
Know what attack-success rate means — the share of attacks that achieved the attacker's goal — and that the number is meaningless unless you also say how many tries the attacker got.
Explain why per-attempt and per-prompt ASR differ under sampling, and why best-of-N resampling can turn an apparently clean suite into a double-digit success rate. Say who or what judged success.
Show you would produce a curve of ASR against attempt budget, sliced by severity and injection channel, with deterministic success checks like canary strings where possible and a human-agreement check on any judge model.
Own how the metric drives the decision: which categories get an absolute-count bar rather than a percentage, what residual rate you are willing to carry, and how containment and monitoring are sized against the rate you cannot drive to zero.
## Defining the number Attack-success rate is the share of adversarial trials in which the system did the thing the attacker wanted: leaked a record, invoked a tool it should not have, produced disallowed content, completed a harmful task. It is the standard safety metric because it maps to the only question that matters operationally — how often does hostile input win? Its weakness is that it is a ratio, and ratios are only as honest as their denominator. ## Three parameters that decide the value **1. What counts as one trial.** *Per-prompt* ASR asks each catalogued attack once and reports how many succeeded. *Per-attempt* ASR asks the same attacks many times and reports the share of individual attempts that succeeded. These diverge wildly under sampling nondeterminism. A defense that blocks a given payload 97% of the time reports 0% per-prompt ASR on a lucky single run — while an attacker who submits it thirty times has better than even odds of getting through. Since the attacker chooses how often to try, the operational metric is *did any attempt in the budget succeed*, not the mean. **2. The attacker's budget.** Safety is a function of compute, not a property. The pattern reported across 2025–26 is steep: a static 200-prompt suite passing at 100% single-shot, then best-of-32 resampling of those same prompts landing around 41% success; a frontier assistant at roughly 0.1% success on a single attempt rising to about 5–6% after a hundred adaptive attempts. Reporting the single-attempt figure while the deployed system faces attackers with unbounded retries is the most common way ASR is used to mislead — usually without intent. **3. Whether the attacker adapts.** Best-of-N with random perturbations is *compute-scaled* but not adaptive. A human or automated attacker who reads each refusal and rewrites accordingly is adaptive, and adaptive attacks are the ones that drove published defenses past 90% success in the 2025 literature. State which you ran; they are different claims. ## Who decides a trial succeeded The grader is part of the metric. Prefer deterministic checks wherever the goal permits one: did the canary string planted in another tenant's record appear in the output? Did the transfer tool actually fire? Was an outbound URL constructed containing case data? These are unambiguous and cheap to rerun. Judge models are necessary for fuzzier goals — "did it give harmful instructions" — but they introduce their own error rate, so sample and hand-label a slice to establish agreement before trusting the aggregate. A safety number produced by an unvalidated judge inherits that judge's blind spots. ## Slicing that carries information A single global ASR hides the thing decisions are made on. Break it down by: - **Severity category.** Cross-tenant data exfiltration and unauthorised state change are not commensurable with an off-policy joke. Aggregate them and a serious failure disappears into a large denominator. - **Injection channel.** Direct user input versus content the system retrieved or a tool returned — defenses that work on one often do nothing on the other. - **Attempt budget.** The curve, not the endpoint. The shape tells you whether you have a hard boundary (flat near zero) or a probabilistic filter (rises steadily with attempts), which is a different engineering conclusion. ## Reporting it without misleading anyone A defensible sentence names all the parameters at once: "Against build 4.2, an adaptive attacker with 100 attempts per case achieved cross-claimant data exfiltration in 2 of 60 cases (3.3%); success was judged by canary-string presence; with a single attempt per case the same suite showed zero successes." That is falsifiable, it tells a reviewer what would change the number, and it makes the residual risk explicit so containment and monitoring can be sized against it. Two failure modes to name explicitly. **Denominator inflation** — padding the suite with easy cases so the percentage falls while the absolute number of successful high-severity attacks is unchanged; count high-severity successes absolutely, not just proportionally. **Budget silence** — publishing a percentage with no attempt count, which is the safety equivalent of a benchmark score with no test set. Neither is usually dishonest on purpose; both are how a team convinces itself it is ready to launch.
- Your suite shows 0% ASR at one attempt per case. What do you do before reporting it?Rerun with a real budget. Resample each case many times and report whether any attempt succeeded, then let an adaptive attacker iterate against the build. A zero at one attempt per case is consistent with a defense that blocks 95% of tries, which is not a boundary. If the number stays at zero across attempts and adaptation, that is a claim worth making — with the budget stated.
- Why report an absolute count for some categories instead of a rate?Because a rate can be diluted. For outcomes like cross-tenant data exfiltration or an unauthorised irreversible action, one success is a launch blocker regardless of denominator, and expanding the suite would make the percentage fall while the harm is unchanged. Rates are useful for tone and policy-adherence categories where volume genuinely matters.
- How do you keep a judge model from inflating or deflating the number?Validate it before you trust it: hand-label a stratified sample of trials, measure agreement with the judge, and report that agreement alongside the ASR. Prefer deterministic checks — a planted canary string, an assertion that a tool fired — wherever the attacker goal can be expressed mechanically, and reserve the judge for genuinely fuzzy goals like harmful-advice quality.
Quoting an ASR without the attempt budget is like reporting a lock's security as "it held against one push" when the attacker has all night and a hundred pushes.
saying these in an interview costs you the question
- Quotes a single ASR percentage with no attempt budget attached
- Reports the mean success rate when the attacker only needs one win
- Aggregates cross-tenant leakage with off-policy tone into one number
- Pads the suite with easy cases so the percentage drops
- Trusts an unvalidated judge model to decide whether an attack succeeded