skip to content

An adversarial-attack library reports two numbers after an evaluation run: accuracy on the adversarial examples, and a mean perturbation size. Which rows does each of those two numbers average over?

level: juniorimportance: must knowfreq 55%

answer

  1. two denominators
  2. conditional on success
  3. failures dropped, not infinite
  4. success rate + mean, as a pair
  5. stronger attack raises the mean

basics

~20 s

Adversarial accuracy is over every example you evaluated, flipped or not. The mean perturbation size is over only the examples the attack actually flipped; rows it failed on are dropped, not counted as huge perturbations. Two different denominators, so the two numbers are not about the same population.

solid answer

~50 s

There are **two different denominators**, and that is the whole answer. - **Accuracy on adversarial examples** (robust accuracy) uses every row you fed the run: still-correct-after-attack divided by all evaluated examples. - **Mean perturbation size** is *conditional on attack success*. The helper measures the norm of the perturbation only for inputs the attack managed to flip, then averages those. Failures are excluded from the average — they are not folded in as an infinite or maximal distance. So the perturbation mean answers "when this attack worked, how much did it have to move the input?", not "how hard is this model to break?". The simple reading breaks as soon as the two numbers move together: a stronger attack that flips more rows pulls in harder examples, which *raises* the mean perturbation while *lowering* adversarial accuracy. Read alone, the mean looks like the model improved.

code

python · 4 lines
python
flipped = [i for i in range(n) if predict(x_adv[i]) != predict(x[i])]
if not flipped:
    return 0.0
return sum(norm(x_adv[i] - x[i]) / norm(x[i]) for i in flipped) / len(flipped)

go deeper

for a junior

Should say the perturbation mean covers only flipped examples while adversarial accuracy covers all of them.

for a middle

Explains that the mean is conditional on attack success, that failures are dropped rather than penalised, and why that makes it non-comparable across runs.

for a senior

Insists on reporting the success count alongside the mean, and points out that a stronger attack can raise the mean by pulling in harder examples.

for a principal

Frames it as an instrument-reading discipline: define which pair of numbers the org quotes and on what frozen example set, so nobody ships a conditional mean as a robustness headline.

Almost every reporting mistake made with these libraries comes from handing one number the other number's denominator. The cure is to be exact about which rows each helper actually counted. **The two numbers, and who computes them.** *Adversarial accuracy* (also called robust accuracy) is the fraction of evaluated rows the model still gets right after the attack has perturbed them. In the Adversarial Robustness Toolbox you assemble it yourself — run `classifier.predict()` over the perturbed batch and compare against the labels; Foolbox packages the same thing as `foolbox.accuracy(fmodel, x_adv, labels)`. Either way the denominator is *every row you fed the run*. *Mean perturbation size* is a different animal. The Adversarial Robustness Toolbox exposes it as `art.metrics.empirical_robustness(classifier, x, attack_name, attack_params)`. Internally it (1) instantiates the named attack with the parameters you passed, (2) generates perturbed inputs for the whole batch, (3) builds a boolean mask of the rows whose predicted label changed, (4) measures the norm of `x_adv - x` for the masked rows only — normally divided by the norm of the original input, so the figure is a *relative* perturbation — and (5) returns the mean over that mask. Foolbox does not average anything for you: an attack call returns the triple `(raw_advs, clipped_advs, success)`, and `success` is exactly that mask. If you average distances without applying it, you have manufactured a number the library never claimed. | | adversarial accuracy | mean perturbation size | |---|---|---| | denominator | all evaluated rows | rows the attack flipped | | defined when nothing flips | yes | no — degenerate | | direction | lower = weaker model | ambiguous (see below) | **Why it is built that way.** A distance to a decision boundary exists only where you crossed one. For a row the attack failed on the library holds no measurement, only a lower bound of the form "further than the budget I was granted". Folding that in as infinity, or as the maximum allowed perturbation, would turn your own flag setting into data and make the average a function of the configuration rather than of the model. Dropping the row is the honest choice, and it is precisely why the two numbers cannot share a denominator. **What the run costs.** An iterative gradient attack costs one forward and one backward pass per step per example: a thousand images at fifty steps is fifty thousand gradient evaluations — a few minutes on a GPU for a small vision model, hours for a large one. A decision-based black-box attack costs thousands of *queries* per example, which against a metered endpoint is an invoice and a rate-limit wall-clock cost, not just compute. That expense is the reason this mistake survives: teams run the sweep once, take whichever floats came out, and paste them onto a slide. The re-run that would disambiguate the two numbers is the expensive thing, so reading the pair correctly the first time is the actual skill being tested. **Where the number misleads.** The two figures move in *contradictory* directions as you change attack strength, and neither moves purely with the model. A weak configuration flips only the rows already sitting nearest a boundary, so the surviving success set is composed of small distances: the mean prints low — the fragile end of this scale — while adversarial accuracy prints high and reassuring. Strengthen the attack and the mask grows to include rows that needed larger perturbations: the mean *rises*, which reads as "the model got harder to break", at exactly the moment adversarial accuracy falls because more rows were broken. Anyone tracking the mean alone across runs, models or library versions is watching the composition of a subset, not a property of a classifier. The same defect makes the mean useless as a headline: it is a conditional expectation whose conditioning event nobody has stated. **What I would check before quoting either.** How many rows were evaluated. How many were flipped — the success count, which is also the denominator of the perturbation mean, and which turns a bare float into evidence or a rumour. Whether rows the model already misclassified before any attack were filtered out. Which norm the distance is in, and whether it is absolute or normalised by the input. What attack and what perturbation budget produced it. Then quote the *pair* — success rate at a stated budget, plus the conditional mean with its count — and never the mean on its own.

  • If a run flips more rows than a previous run and the mean perturbation goes up, did the model get more robust?
    No conclusion. The stronger run pulled harder examples into the average. Compare on a fixed example set at a pinned perturbation budget, using success rate, not the conditional mean.
  • Why do libraries not record failures as an infinite perturbation?
    A failure gives only a lower bound set by the search budget you granted, not a measured distance. Recording infinity would turn a budget choice into data and make the mean undefined.
  • Which of the two numbers is safe to quote on its own?
    Adversarial accuracy, provided you also state the attack and its perturbation budget, because its denominator is the full evaluated set.

It is like a locksmith who reports his average time to open a lock, counting only the locks he actually got open; the ones that defeated him are left out rather than recorded as taking forever. As he gets better and starts cracking harder locks, his average time goes up even though no lock got tougher.

saying these in an interview costs you the question

  • Treating the mean perturbation as a whole-test-set average.
  • Assuming unflipped rows enter the average as a maximal or infinite distance.
  • Reading a rise in the mean perturbation as an improvement without checking how many rows were flipped.
  • Quoting the perturbation mean with no attack, norm, or success count attached.

context