skip to content

Why doesn't training against small per-pixel changes stop an attacker who changes a few pixels a lot?

level: middleimportance: must knowfreq 56%

answer

  1. the allowed set is a shape
  2. worst case is searched inside it
  3. thin box versus a few extremes
  4. a rotation is in no norm ball

basics

~20 s

Because adversarial training makes a model resist the worst case inside one fixed set of allowed changes. "Every coordinate moves a little" and "a few coordinates move a lot" are different sets, and points in the second sit far outside the first.

solid answer

~50 s

Adversarial training is a min-max procedure: at each step it searches for the worst input inside a stated allowed set against the current weights, and learns from that. The allowed set is a geometric object the defender picked. A small per-coordinate bound is a thin box around the input — everything may move slightly, nothing may move far. A sparse budget is the opposite shape — almost everything is frozen, a handful of coordinates may move to the extremes. The worst case in the second set is not in the first set at all, so it was never in the training signal. Empirically, cross-family robust accuracy sits far below the headline figure, and pushing hard on one family can cost ground on another. Geometric changes are worse still: a rotation or a crop is not a magnitude budget at all.

go deeper

for a junior

Know that adversarial training targets one stated kind of change with one stated size, and that an attacker is not obliged to use that kind. Being able to name two different kinds of change is enough here.

for a middle

Explain the min-max mechanic: the inner search finds the worst input inside a fixed allowed set, so the resistance learned has the shape of that set. Be able to contrast a per-coordinate bound with a sparse one concretely.

for a senior

Be ready to audit an evaluation for which families it covered and which it never had a unit for, and to say what you would measure instead for your own deployment's attacker.

for a principal

Weigh whether widening the trained union is worth its accuracy and compute bill against your actual adversary, and be willing to say the real gap is geometric and will not be closed by adding another norm.

## What the procedure actually optimises Adversarial training solves an inner search and an outer fit at the same time. At every training step it asks: **given the weights as they stand right now, what is the most damaging input inside the allowed set?** — and then updates the weights to be right on that input. The important word is *allowed set*. It is chosen by the defender, before training, and it is a specific geometric object. That choice, not the training procedure, is what the resulting claim is about. The model becomes stable over the shapes of change it was repeatedly shown at their worst. It does not become stable over "change" in the abstract, because there is no such object to train against. ## The families are different shapes, not different amounts Stated informally: | allowed set | what it permits | what it forbids | | --- | --- | --- | | small per-coordinate bound | every coordinate moves a little | any coordinate moving far | | sparse / few-coordinates budget | a handful of coordinates move to the extremes | any change to the rest | | total-energy bound | a fixed amount of change distributed as the attacker likes | exceeding the total, however spread | A point that is worst-case for the second row — a few coordinates driven hard — violates the first row by a wide margin, so it never appeared in that training's inner search. The reverse also holds: a diffuse change spread over everything is not reachable inside a sparse budget. These are not "stronger" and "weaker" versions of one adversary; they are **different adversaries with different real-world counterparts** (sensor-level noise floors versus a few dead or overwritten elements, for instance). This is why measured cross-family results are so much lower than headline numbers, and why some evaluations show a genuine trade-off: capacity spent flattening the loss surface in one geometry is not free, and can leave the model more exposed in another. ## The changes that live in no family at all The larger gap is not between norms. It is between norms and everything that is not one. Rotating an input slightly, shifting the crop, changing the viewing angle, or covering part of the scene are **geometric and area-bounded** changes. Measured with any pixel metric they are gigantic — far outside any radius a defense would train against — and yet they are semantically trivial: a person sees the same object, the same face, the same sign. A defense priced in magnitude has no unit for these. It cannot say whether it is robust to a fifteen-degree pose change, because degrees are not in the budget it trained on. So the correct statement is not "weakly robust" but **"not measured"** — and an attacker choosing where to work reads that gap directly off the published claim. ## Can you just train against the union? You can, and it helps inside the families you named. Three things to be honest about: 1. **It costs more.** Worst-case search over a union is harder, and the clean-accuracy bill is larger than for a single family. That bill does not land evenly — it falls hardest on the examples the model already found difficult. 2. **It closes nothing you did not enumerate.** A union of three families is still a list. Anything not on the list, notably geometric and area-bounded change, remains unpriced. 3. **It does not become a guarantee.** These are empirical numbers against attacks that were run. They bound the effort spent attacking, not the attacker. ## How to spot the gap in someone else's evaluation Read the columns, not the headline. An evaluation that reports a single family and a single radius has measured a single adversary; ask what was evaluated *and rejected*, and what was never evaluated at all. If the answer is "one norm, one radius," the claim covers one adversary, and the deployment's real attacker may not be that one — especially if the real attacker is standing in front of a camera rather than editing a file. ## The interview answer in one line Robustness is a property of a model **and** a set. Change the set and you have changed the claim, not merely stretched it.

  • Which perturbation family does a small rotation or a re-crop belong to?
    None. It is not a magnitude budget at all — it is parameterised geometrically, by angle or offset. In pixel terms it is an enormous change, which is exactly why a defense trained on a small per-coordinate radius neither resists it nor has a way to report on it.
  • Why does a single-step attack sometimes report higher robust accuracy than an iterative one?
    Because it is a weaker attack. A single step takes one move along the direction read off the loss with respect to the input; an iterative attack takes many small steps, re-projecting into the allowed set each time, and finds worse points. If the single-step attack ever looks stronger, that inversion is a warning sign about the evaluation, not a compliment to the defense.
  • Does training against a union of families make the claim general?
    No — it makes it broader and more expensive. It buys resistance inside the families that were enumerated, at a larger clean-accuracy cost, and leaves everything unenumerated untouched. The claim is still a list, and an attacker reads the list to find what is not on it.

saying these in an interview costs you the question

  • Thinks the norms are just different amounts of the same attack
  • Believes robustness in one norm implies robustness in another
  • Calls a rotation or crop a small perturbation
  • Assumes union training turns an empirical number into a guarantee
  • Ignores that the clean-accuracy bill falls on hard examples

context