skip to content

An Attacker Who Knows

A defense is worth what the best attack designed against it achieves, not what a stock attack suite reports. Interviewers ask because it is the one methodological rule this whole field runs on.

on this pageshow

explore

questions

4

Why isn't a robust-accuracy number from a standard attack suite evidence that a defense works?

level: juniorimportance: must knowfreq 74%

answer

  1. the suite predates your defense
  2. a high score has two causes
  3. it may have broken the search
  4. what is being measured is unfamiliarity
  5. the attack must be written afterwards

basics

~20 s

A standard suite runs attacks written before the defense existed, so a high score can mean the defense disturbed the attacker's search rather than stopped the attack. Evidence needs an attack designed by someone holding the defense's description.

solid answer

~50 s

A published attack suite is a fixed set of searches tuned against undefended models. Running it on a new defense tells you one thing: those particular searches, in their default form, did not find inputs the model reads wrongly. That is equally consistent with the defense being strong and with it having broken the search — a randomized stage, a non-differentiable preprocessing step or an ensemble vote can send a gradient-following search wandering without changing which inputs are actually misread. The convention the field settled on is adaptive evaluation: the adversary is *granted* the defense's full design by assumption and writes the attack against it. Anything weaker is measuring how unfamiliar your defense is, which is obscurity, not robustness. So the defensible wording of a result is `not broken by these attacks, written with the defense in hand, at this effort` — never simply `robust`.

code

text · 7 lines
text
Table 3 - Robustness of the shipped classifier
  attack ............ published attack suite, default configuration
  suite authored .... before this defense was designed
  defense-aware ..... no
  access ............ query only (verdict + coarse confidence band)
  robust accuracy ... 60.2%
  ...

go deeper

for a junior

Be ready to say, in one sentence, that a stock suite was written before your defense existed, so a good score may mean its searches broke rather than that the model is safe.

for a middle

An interviewer expects you to name the concrete ways a defense can disturb a search - randomization, a non-differentiable step, an ensemble vote - and to explain why that produces a high number without changing which inputs are misread.

for a senior

Show the asymmetry in production terms: a break is decisive, a failure is bounded by who tried. Then say who owns producing the defense-aware attack and how a claim gets worded so it survives review.

for a principal

Own the standard itself. Decide what evidence your organisation will accept before a robustness claim goes to a customer, who funds the attack effort, and what the claim is allowed to say when nobody funded it.

### What the number actually is Robust accuracy is the share of a test set a model still classifies correctly when an adversary is allowed to modify each input inside a stated budget. Two things determine that number, not one: the model-and-defense being measured, and **the attack that was run against it**. A published attack suite fixes the second half in advance, with searches designed for plain, undefended models — usually written before the defense under test was invented. The number that comes out is a joint property of your system and somebody else's assumptions. ### Two explanations for a high score, and they are not equally good news When a defense scores well against a stock suite, exactly two things could be true. 1. Inputs that the defended system reads wrongly, inside the stated budget, are genuinely hard to find. 2. Such inputs are still there, but the stock search cannot locate them because the defense perturbed the mechanics of searching — a randomization stage makes the signal the search follows noisy, a non-differentiable preprocessing step means there is no usable derivative to follow, an ensemble vote flattens the quantity the search was reading, a detector rejects the exact shape the suite happens to produce. The second case is common enough that the field has diagnostic signals for it: an attack given weaker access outscoring one given stronger access, or an attack with its budget removed entirely still not driving the model to total failure. Neither of those is possible if the search is working. A stock number cannot distinguish case 1 from case 2, because in both cases the same thing is observed — the suite returned without a break. ### Obscurity, measured What an unaware attack really measures is the *gap between the attacker's assumptions and your system*. That gap is real, and in the field it does buy time. But it is not a property of the model, it is not something you can promise a customer, and it shrinks every time your design is documented, inferred from behaviour, or reconstructed from a shipped artifact. That is why the adaptive convention grants the adversary the design rather than waiting for them to obtain it: a number that improves when your documentation gets worse is not measuring the defense. ### The evidence is asymmetric, which is why the burden sits with the claimant A successful break is decisive: something exists, someone found it, the claim is over. A failure to break is bounded by whoever tried, how hard, and from what vantage. This asymmetry has a direct consequence: producing the strongest defense-aware attack is the *author's* job. A reviewer who declines to spend a week attacking your design has not endorsed it, and a suite that never knew your design existed has not either. The reverse direction is worth stating too. If a generic, defense-unaware attack *does* break the model inside the stated budget, the claim is refuted immediately, and no adaptive work is required. Stock suites are weak evidence of strength and strong evidence of weakness — useful as a cheap screen, useless as a warrant. ### Why this rule exists at all This is the central methodological finding of the whole defense literature, and it was learned the expensive way: a long run of published defenses reported large robustness gains, and a large fraction of them were reduced to roughly the undefended level once somebody wrote an attack with the defense's description in hand. In almost none of those cases was the original evaluation dishonest. The evaluations ran real attacks and reported real numbers; the attacks simply had not been written against the thing being measured. ### What it looks like in a real engagement Take a static malware classifier shipped inside an endpoint agent, deciding block-or-allow at execution time. The vendor whitepaper describes the defense fully — a randomized feature-hashing stage plus an ensemble vote — and reports a robustness figure from a published suite. Now a contracted red-teamer arrives with a trial licence, an offline copy of the engine they can run locally as often as they like, and the whitepaper. They get back a verdict and a coarse confidence band, and no derivatives through the ensemble. That person is a categorically different adversary from the suite, because they can aim at the thing the deployed pipeline actually decides on. When such an attack is written, the whitepaper number frequently collapses — and the genuinely useful output is not merely *the number was wrong*, but what erasing it cost in analyst-days, because that cost is what tells the vendor which adversaries the defense still stops. ### How to phrase a result you can defend Name the attacks that were written against this defense, the access the attacker was granted, the effort actually spent, and the date. Report the number as what those attacks achieved. Say plainly that a better-resourced adversary is not covered. A claim in that shape can be argued with, re-scoped and re-tested. `60% robust` cannot.

  • The authors say they ran ten different published attacks instead of one. Does that fix the problem?
    No. Ten searches written without knowledge of the defense are still ten unaware searches, and they tend to fail for the same reason — whatever the defense did to the first one it also did to the other nine. Adaptivity is a property of how an attack was designed, not a count. One attack aimed at the deployed mechanism outweighs any number of generic ones.
  • If a stock suite drives the model to near-zero accuracy inside the stated budget, is that number worth anything?
    Yes, and more than the high one. Failure found by any attack is a real failure, so the robustness claim is refuted on the spot and nobody needs to fund a defense-aware attempt. The asymmetry is the point: an unaware suite is weak evidence of strength and strong evidence of weakness, which makes it a useful cheap screen and a useless warrant.
  • Isn't keeping the defense's design private a legitimate control?
    It is a legitimate cost control and an illegitimate evidence base. Non-disclosure raises the effort an outsider spends before they can aim at your mechanism, which has real value in the field. It cannot appear in a robustness claim, because the claim would then depend on a secret that documentation, a shipped binary, or careful probing of behaviour can each remove.

A lock that survives every key already on the locksmith's ring has shown that none of those keys fit. It has not shown that no key fits, and least of all against someone holding its schematic.

saying these in an interview costs you the question

  • Reads a high stock-suite score as proof of robustness
  • Thinks running more standard attacks makes an evaluation adaptive
  • States a robustness number without saying who attacked it
  • Assumes an adversary will not read the published design
  • Confuses breaking the attacker's search with stopping the attack

context

open as a page

In evaluating a model defense, what makes an attack adaptive rather than stock?

level: middleimportance: must knowfreq 58%

basics

~20 s

An adaptive attack is designed after reading the defense. The adversary is assumed to know the mechanism and its parameters, and aims the attack at the quantity the defended pipeline actually decides on, rather than at the model underneath it.

open as a page

A nine-day adaptive attack on a shipped malware classifier failed - what can the red-team report honestly claim?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Only that these designs, from this vantage, at this effort, did not break it. A failure to break is bounded by who tried and how hard, so the report states access, days spent and designs abandoned - never robust.

open as a page

Your team wants to publish a robustness claim backed only by a stock attack suite - what do you require before it ships?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

That the authors attack their own defense first and publish what they tried, from what access, at what cost. The burden sits with whoever makes the claim, and the wording must carry those bounds and a date.

open as a page