skip to content

Two image classifiers are scored with the same library metric — the mean perturbation size over the examples an attack flipped — using the same attack, the same norm and the same settings. Model A's number is larger than model B's. Why does that not settle which model is more robust?

level: seniorimportance: should knowfreq 44%

answer

  1. E[distance | success, model]
  2. different success sets, same units
  3. hardened model loses only borderline points
  4. thin denominator, huge variance
  5. compare success rate on a frozen set

basics

~20 s

Because each number averages over a different set of rows: whichever examples that model happened to lose. If the attack flipped only a handful of A's examples, A's average covers a narrow subset, while B's may cover most of its test set including hard cases. Same units, different populations, so the comparison is not like-for-like.

solid answer

~50 s

Units and settings being identical does not make the two averages comparable, because the metric is **conditional on attack success** and the two models have different success sets. Suppose the attack flips 8% of A's examples and 85% of B's. A's mean is computed over eight rows in a hundred — and the composition of that subset is not random; it is whatever A was weakest on. B's mean spans nearly its whole test set, easy and hard rows alike. A conditional mean over a small, selected subset can sit anywhere relative to one over a broad subset, so their ordering carries no ordering of robustness. The usable comparison is over a **common denominator**: attack success rate on the same frozen example set at a pinned perturbation budget, and if you still want a distance, the distribution of distances restricted to examples *both* models lost — plus the row counts, so a reader can see how thin either average is.

go deeper

for a junior

Should at least notice the two averages come from different sets of examples.

for a middle

Names the conditioning on attack success and explains that success-set composition, not robustness, drives the difference.

for a senior

Constructs the inversion case, flags the variance of a thin success set, and proposes a fixed-budget success-rate comparison on a frozen example set instead.

for a principal

Decides what the organisation compares at all: one frozen evaluation set, a pinned budget, reported denominators, and no headline ranking from a conditional mean.

**The statistical shape.** The quantity the helper returns is `E[distance | the attack succeeded on this row, model M]`. It is a conditional expectation, and the conditioning event is model-specific: it is "whichever rows *this* model happened to lose". Identical attack settings fix the *procedure* applied to both models; they do not fix the *sample* each average is taken over. Two conditional means whose conditioning events differ are not on a common scale, and no amount of matching norms, budgets or random seeds repairs that. This is why the answer is not "run it again more carefully" — it is a property of what the helper computes. **The concrete inversion, which is the part worth holding onto.** Suppose the attack flips 8% of model A's rows and 85% of model B's. A hardened model resists nearly everything and loses only on genuinely borderline points — points that by construction already sit near a decision boundary and therefore need *small* perturbations. A weak model loses almost everywhere, including on points that were far from any boundary and required *large* perturbations to move. Averaged, the hardened model can score **lower** on mean perturbation while being unambiguously harder to break. Selection into the average runs against the property you are trying to measure, so the ordering of the two floats carries no ordering of robustness — not a weak signal, no signal. **Second effect: thin denominators.** Eight successes out of a hundred rows produce an average with enormous variance; one outlier moves it visibly. Eight hundred successes produce a stable one. A confidence interval on eight points and one on eight hundred are wildly different evidence, and the helper hands you a bare float with no indication of either. Comparing A's eight-row mean to B's eight-hundred-row mean is comparing an anecdote to a measurement. **What I would actually run.** - Freeze one example set and use it for both models, filtered to rows each model classifies correctly at baseline — and record how that filter differs per model, because it has already changed the population before the attack starts. - Compare **attack success rate at a fixed perturbation budget**. The denominator is every eligible row, identical in construction for both models, and the quantity is monotone in the direction you care about: more successes means less resistant, always. - If a distance comparison is genuinely wanted, restrict it to the *intersection* of the two success sets and compare paired distances there. Small, but genuinely like-for-like. - Carry `n_evaluated`, `n_eligible`, `n_flipped` and the budget beside every number so a reader can see how thin either average is. **What that costs.** You are now paying for two full attack runs over a frozen set rather than reusing whatever each team ran last quarter, and the intersection trick shrinks your usable sample — if A loses only 8% of rows, the paired comparison rests on at most those rows, so obtaining a stable estimate may mean growing the evaluation set several-fold and paying several times the GPU-hours or metered queries. A defensible pinned-budget comparison of two models on a vision workload is realistically hours of accelerator time plus a day of engineer time to build and check the harness; a "just diff the two floats we already have" comparison costs nothing, which is exactly why it is the one people do. **Where it still misleads even after you fix the denominator.** A success rate under a single attack family at a single budget is an *upper bound* on robustness under that procedure only. Two models can trade places under a different attack family, and a model tuned against the family you chose will look better than it is. So the honest output is never "A is more robust than B". It is: under this named attack, at this named budget, on this named frozen example set, A's success rate was x% over n rows and B's was y% over m rows — with the explicit note that a different attack family could reverse it. A ranking claim that outlives its attack, its budget and its example set is the failure mode this whole question exists to catch.

  • Give the mechanism by which a more robust model scores lower on this metric.
    It only loses on points already near a decision boundary, which need small perturbations. The weak model also loses far-from-boundary points needing large ones, which inflate its average.
  • What comparison would you substitute?
    Attack success rate at a pinned perturbation budget on one frozen example set, with row counts reported — the denominator is then identical for both models.

Think of two goalkeepers who each report the average distance of the shots that beat them. The excellent one is only beaten from close range, so his average is small; the poor one is beaten from everywhere, including the halfway line, so his average is large. The bigger number belongs to the worse keeper.

saying these in an interview costs you the question

  • Concluding 'A is more robust' straight from the larger number.
  • Believing identical attack settings make the two averages comparable.
  • Ignoring how many rows entered each average.
  • Ranking two models from a single attack family with no budget stated.

context