skip to content

An image classifier is adversarially trained inside an L-infinity ball of a chosen radius — which attackers does that cover?

level: middleimportance: should knowfreq 52%

answer

  1. not a hyperparameter
  2. each norm is a different adversary
  3. a little everywhere versus a lot in a few places
  4. outside the radius it is a cliff
  5. a camera respects no ball

basics

~20 s

Exactly the attackers whose whole move is a per-pixel change no larger than the chosen radius. That robustness transfers poorly to sparse or energy-bounded perturbations, barely at all to a camera-captured artefact, and falls off steeply just outside the radius.

solid answer

~50 s

The norm and the radius are not hyperparameters, they are the threat model, and the trained model is fitted to that region and nothing else. An L-infinity ball says every coordinate may move a little; an L-2 ball bounds the total energy of the change; a sparse budget says a few coordinates may move a lot. Those are three different adversaries, and training against one buys little against the others - a model fitted to small per-pixel changes can still be broken by rewriting twenty pixels outright. Nothing in the ball describes an attacker who re-photographs or re-encodes the image, because a capture-level change respects no norm. And robustness does not plateau outside the trained radius; it falls away quickly. The honest coverage statement: a digital adversary who edits every pixel of an upload by at most r, evaluated with a search at least as strong as the one used in training.

go deeper

for a junior

Know that a robustness claim is meaningless without the norm and the radius it was trained and measured at, and that a small per-pixel budget and a few-pixels-changed-a-lot budget are different attackers.

for a middle

Explain what each family lets the adversary do and why a model fitted to one shape of change was never asked about another. Be able to describe the steep fall-off just outside the trained radius.

for a senior

Demand the full set of columns before accepting a result — access assumption, norm, radius, steps, restarts — and be able to say which two rows in a table are not comparable and why.

for a principal

Own the question of whether the chosen region describes an adversary the product actually faces, given that widening it costs compute and clean accuracy for every user while an attacker who re-photographs the input pays none of it.

## The radius is a claim, not a knob When a team says "we adversarially trained the classifier", the sentence is incomplete until they say **which norm and which radius**. The training loop fits the model to the worst point inside that region; everything the method delivers is a statement about that region. Quoting a robust-accuracy figure without the norm and the radius quotes nothing at all, because the same model can look strong or useless depending on which region you evaluate over. ## What each norm actually says about an adversary The families are not interchangeable, and each corresponds to a different real capability. | Region | What the adversary may do | Who that resembles | |---|---|---| | L-infinity, radius r | move **every** coordinate, each by at most r | someone editing a file they upload, spreading a small change everywhere | | L-2, radius r | move coordinates freely subject to a **total energy** budget | a change concentrated where it helps, still small overall | | Sparse (few coordinates) | change **a handful** of coordinates by any amount | someone overwriting a small region outright | | Area and viewpoint bounded | place something in the scene the camera captures | a physical adversary, who respects no norm ball at all | The last row is the one candidates most often mishandle. A digital perturbation is bounded in magnitude; a physical change is bounded in **area and viewpoint** and must survive angle, lighting, printing and re-encoding. Robustness bought in a magnitude ball is not evidence about it. That is a different threat model with different machinery. ## Why robustness does not transfer across the families Fitting a neighbourhood teaches the model that a specific *shape* of change is uninformative. An L-infinity ball's shape is "a little bit, everywhere", and its extreme points are perturbations that touch every coordinate. A sparse adversary's extreme points sit nowhere near those — most coordinates untouched, a few pushed as far as the domain allows — so the model was never asked about them. Empirically, a model trained for one family typically holds up far better inside that family than outside it, and the fall-off is largest against sparse and physical adversaries. You can train against a **union** of families — solve the inner search over several regions and fit the worst case across all of them — but the compute rises with the number of families searched, and the robustness attained within any one of them is usually lower than training for that one alone. Breadth is bought with depth. ## Inside the norm, the radius is a cliff and not a wall The second half of the coverage question is what happens at r + a little. Adversarial training does not produce a plateau that decays gently; robust accuracy typically drops steeply once the evaluation radius passes the trained one. That is the expected shape, not a defect — the model was fitted to a region and the evaluation left it. The defect is reporting a single number at the trained radius and letting a reader hear "the model is robust". So a robustness result is only readable with all of: the **access assumption** (weights and gradients available, or query-only), the **norm**, the **radius**, and the **strength of the search** — steps and restarts. Two rows with different values in any of those columns are not comparable. ## The adversary is not obliged to agree The deepest version of the point: the threat model is a guess about what an attacker will accept, made by the defender, before the attacker showed up. An attacker with a goal — get a listing photo accepted in a category it does not belong to — has no reason to confine themselves to imperceptible per-pixel edits. They can re-shoot the photograph, composite it, compress it, or change the object. None of those are in the ball, and none of them are addressed by the number the training produced. The trained region is a real cost imposed on one specific family of moves; treating it as coverage of the attack surface is the mistake the question is testing for. ## Answering it crisply Name the region as a threat model; say it covers a digital, magnitude-bounded adversary of that norm at or below that radius, evaluated with a comparable search; say the transfer to other norms is weak and to physical or capture-level changes essentially nil; and say the number falls off steeply just outside the radius, so the honest artefact is a curve across radii and at least one other norm, not a point.

  • Which real adversary does an L-infinity ball model well, and which does it model badly?
    It models well someone who holds a digital file and can nudge every pixel a little before submitting it — an in-file, magnitude-bounded editor. It models badly anyone who changes a few coordinates a lot, anyone whose change is captured through a camera or a re-encode, and anyone who simply supplies a different image. Those adversaries were never in the region the model was fitted to.
  • Robust accuracy is 62% at the trained radius and 3% at twice it. Is the defence broken?
    No — that is the expected shape. Adversarial training buys a fitted neighbourhood, not a wall, and the claim was only ever about that neighbourhood. What is broken is a report that gives the 62% without the radius it was measured at, because a reader will hear it as coverage. The correct artefact is the curve across radii, and the decision is whether the radius describes an adversary you face.
  • Can you train against several perturbation families at once?
    Yes — the inner search can maximise over a union of regions, so the model is fitted to the worst case across all of them. It costs more compute, since each family has to be searched, and the robustness reached in any single family is typically lower than training for that family alone. You buy breadth by spending depth, which is a real decision rather than a free upgrade.

saying these in an interview costs you the question

  • Treats the radius as a tuning knob rather than a threat model
  • Assumes robustness in one norm implies robustness in another
  • Reads a robust-accuracy number as protection against any change
  • Expects a magnitude ball to cover a camera-captured change
  • Quotes a robust number without the norm, radius or access assumption

context