skip to content

A vendor calls their face-verification model "adversarially trained, so robust" — what does that leave an attacker free to do?

level: juniorimportance: must knowfreq 74%

answer

  1. robust to what, exactly
  2. a family and a radius
  3. the attacker never agreed to it
  4. pose and occlusion are not pixel budgets

basics

~20 s

Ask: robust to what? Adversarial training buys resistance only inside the perturbation family and radius it trained on. An attacker stays free to work in a different metric, or to change pose and occlude part of the face, which no per-pixel budget bounds.

solid answer

~50 s

"Robust" is not a claim until a threat model is attached: which perturbation family (which norm), what radius, which attack was run inside it, and what access that attack was given. Adversarial training regenerates the worst case inside exactly that set at every step, so the resistance it learns is shaped like that set. An attacker who never agreed to the set simply steps outside it — moving a few coordinates a lot instead of every coordinate a little, or changing nothing per-pixel at all and instead altering head pose, framing, or an occluded area. That last change is enormous in any pixel metric and tiny to a human observer, so a magnitude budget does not even measure it. The honest reading of the vendor's number is: the attack we ran, inside the budget we chose, failed this often.

go deeper

for a junior

Recall that a robustness claim needs a threat model attached: which kind of change, how much of it, and what the attacker could see. Be ready to say "robust to what?" out loud as your first response.

for a middle

Explain why the trained-against family bounds the claim: the training keeps regenerating the worst case inside one fixed set of allowed changes, so what the model resists is that set and not change in general.

for a senior

Show you would map the vendor's family onto your deployment's real attacker before accepting the number, and would ask explicitly what was not evaluated rather than only how high the figure is.

for a principal

Own the framing that a robustness figure is a scoped statement your organisation signs, not a property you inherit. Decide what you will state, what you scope out, and who carries the families nobody measured.

## The claim, and what is missing from it "Adversarially trained, so robust" states a training procedure and then asserts a property. The property does not exist on its own. Robustness is always **relative to a threat model**, and a usable threat model names at least four things: - **the perturbation family** — the geometry of changes the attacker is allowed to make (commonly stated as a norm); - **the radius** — how much change that family permits; - **the access assumed** — whether the attacker was given the weights and gradients, only returned scores, or only a top-1 decision; - **the attack actually run inside that budget**, and how hard it was run. Drop any of those and the number is not comparable to any other number, and not transferable to your deployment. ## Why the trained family is the boundary of the claim Adversarial training works by repeatedly asking, during training, "what is the worst input inside this allowed set, against the weights as they stand right now?" and then learning from it. The allowed set is a fixed geometric object chosen by the defender. What the model ends up resistant to is *that object*: it learns to keep its decision stable over the shapes of change it was repeatedly shown at their worst. Different families are different objects, and they are not nested in any convenient way: - a **small per-coordinate bound** lets every pixel move a little and none of them much; - a **few-coordinates budget** lets a handful of pixels move as far as they like and forbids the rest to move at all; - an **energy bound** limits the total size of the change but not how it is distributed. A point that is extreme in one of these can sit far outside the others. So resistance bought against one buys much less against another, and published evaluations repeatedly show cross-family robustness far below the headline figure — sometimes with a visible trade-off, where pushing hard on one family costs ground on another. ## The change that is not in any family The sharper gap is that some attacker moves are not magnitude-bounded at all. Turning the head, moving closer, changing the crop, or covering part of the face with an ordinary object are **geometric and area-bounded** changes. Measured in pixels, they are enormous — vastly larger than any radius a defense trains against — and yet a person looking at the result sees the same face. A defense priced in per-pixel magnitude does not resist these; it does not even have a unit in which to state them. This is why a robustness figure earned against a magnitude budget says essentially nothing about an attacker who is physically present in front of a camera and constrained instead by how far they can turn their head and how much of their face they may cover. ## Reading the number in the right direction A high robust-accuracy figure proves that **the attack that was run failed** at the budget it was run under. It does not prove that no attack succeeds. Two attackers with the same nominal "strength" but different families are simply different adversaries, and only the one that was tested was priced. So the first question to a robustness claim is not "how high is it" but "robust to what": which family, which radius, which access, which attack and how hard. The second is "what was *not* evaluated" — the families deliberately or accidentally left off the sheet are exactly where an attacker who has read the claim will go, because the claim tells them where the training effort was not spent. ## What to do with a vendor claim Treat the figure as characterising one adversary. If your deployment's realistic adversary is constrained differently — by physical access and a viewing angle rather than by a pixel budget — the vendor's number is true and irrelevant, and you need a measurement stated in the units your attacker actually works in.

  • They add that they also trained on random noise — does that broaden the claim?
    Barely. Random noise of the same magnitude essentially never flips a trained classifier, so training against it prices an adversary who does not exist. An adversarial perturbation is a chosen direction read off the loss with respect to the input, not noise; resistance to noise says nothing about resistance to a direction picked to hurt.
  • If they trained against two perturbation families instead of one, is the claim general?
    Broader, still bounded. Training over a union raises an attacker's cost inside the families that were named, generally costs more clean accuracy, and gives no coverage at all over a family nobody enumerated — most obviously a geometric or area-bounded change, which respects no magnitude budget.
  • What one line would you add to the claim to make it usable?
    The perturbation family and radius, the access the attack was given, the attack and how hard it was run, and the metric reported. With those, the number is comparable and scopeable; without them it cannot be compared to any other robustness number or applied to your deployment.

It is like a lock rated against picking. The rating is real, and it says nothing about someone who removes the hinges.

saying these in an interview costs you the question

  • Treats "adversarially trained" as a property of the model itself
  • Assumes robustness is one scalar that transfers across attack types
  • Believes a small per-pixel bound also covers occlusion or pose
  • Confuses an adversarial perturbation with random noise
  • Quotes a robust-accuracy number with no norm and no radius

context