A year after shipping a classifier adversarially trained at one L-infinity radius, what is that radius choice still worth?
answer
- a price list, not a wall
- the toll gets cheaper as search gets cheaper
- just outside the radius it collapses
- public weights mean white-box by default
- report the curve, never a point
basics
~20 sA price, not a wall. Inside the trained radius an attacker pays far more search and far more visible distortion for the same misread; a short step outside, that price collapses. Re-measure and report the curve.
solid answer
~50 sThink of what the radius bought as a price list. Take a marketplace classifier checking whether an uploaded photo matches the declared category, whose weights were open-sourced: the attacker solves against final fixed parameters, an easier problem than the moving target the training loop chased. Inside the trained radius, forcing the same targeted misread costs them many more search steps and restarts, and a distortion large enough that a reviewer might notice. Just outside it, the cost falls back toward an undefended model, and against a sparse edit or a re-shot listing it was never charged at all. So a year on the honest statement is a curve, not a percentage: robust accuracy across radii at and above the trained one, in at least one other norm, measured with the strongest search you can afford today.
code
text · 11 linesshipped: adversarial training, L-inf, radius 4/255, 10 inner steps
re-measured: white-box (weights public), 50 steps, 5 restarts
access norm radius steps restarts accuracy
- - - - - 0.91 (clean)
white-box L-inf 4/255 50 5 0.62
white-box L-inf 6/255 50 5 0.44
white-box L-inf 8/255 50 5 0.21
white-box L-2 0.5 50 5 0.33
white-box sparse 20 pixels 50 5 0.08
...go deeper
Remember that a robustness result describes one region and one attack that was run, so it ages: the same radius is worth less each year as the search an attacker can afford gets stronger and cheaper.
Be able to explain why the number falls off steeply outside the trained radius and why an attacker holding final weights has an easier problem than the moving target the training loop chased.
Show the re-measurement discipline: white-box by default when weights are public, a stronger search than last time, a curve across radii and at least one other norm, and an explicit list of the moves that were never priced.
Own the framing that the defence is a toll whose worth is set by what one misread is worth to the attacker, and be ready to argue that widening the region is not automatically the right next spend.
## The question behind the question Somebody chose a radius before the model shipped. A year later the person who has to answer for the defence is usually not the engineer who tuned it, and the question they are asked is not "is the model robust" but "what is that decision still buying us". Answering it well means converting a training-time choice into a statement about **an adversary's cost today**. ## What was actually purchased Adversarial training bought a fitted neighbourhood: within a stated norm and radius, the model was trained on the hardest point the inner search could find, so points in that region stop flipping the answer cheaply. The purchase is real and it is measurable, and it has three properties worth stating plainly. **It is a price, not a barrier.** An attacker inside the region is not stopped; they are made to spend more — more search steps, more restarts, and typically a larger, more visible distortion to achieve the same targeted misread. Whether that price matters depends entirely on what one misread is worth to them. **The price collapses at the boundary.** Robust accuracy does not decay gently past the trained radius; it falls steeply. So the defence's value is concentrated on adversaries who, for their own reasons, want to stay imperceptible. **It was never charged outside the norm.** A sparse edit, a re-encode, a re-shot photograph or a substituted object were not in the region and were never priced. ## Reading the re-measurement A year on, the right artefact is a table, and its columns are the argument. Access assumption, norm, radius, steps, restarts, accuracy. The row at the trained radius is the claim that was made; every row below it is an attacker who declined to stay in the ball. If the shipped model's weights are public, the access column reads white-box by default and no part of the remaining number may be attributed to obscurity. Two readings are commonly wrong. A steep fall-off just outside the trained radius is **not** evidence that the training failed — it is the expected shape of what was bought. And an unchanged number from a year ago is not evidence that nothing has changed, if it was produced by re-running the same search at the same settings; the thing that moves over a year is the strength and cheapness of the search an attacker will run, so a re-measurement that does not increase steps and restarts has re-measured nothing. ## What you can honestly say - **What it still stops:** an adversary who edits the file they submit within the trained magnitude, evaluated with a search of at least the strength used in training. Give the current number with its columns. - **What it never stopped:** anyone whose move is outside the region — a few coordinates changed a lot, a capture-level change, or a genuinely different input. Those need their own measurement or their own answer. - **What decayed:** not the region, which is fixed, but the assumption that an attacker's search is expensive. Cheaper compute and better search make the same radius a smaller toll each year. - **What has not been established:** anything about the model as a whole. An empirical robust-accuracy figure bounds the attack that was run. It is not a guarantee, and it should never be repeated in a deck as one. ## The decision that follows The tempting move is to retrain at a larger radius. That is only right if the larger region describes an adversary you face, because widening it costs compute on every retrain and clean accuracy on every prediction — bills paid by all traffic, including the overwhelming majority that is not adversarial — while an attacker who re-photographs the listing pays none of it. The comparison worth putting in front of whoever owns the budget is: what does the same spend move the attacker's price by, here versus elsewhere in the pipeline, given what one successful misread is worth to them. On a marketplace listing, if a misread is worth a few dollars and the defence forces a fifty-step white-box search plus visible distortion, the economics are decisive and the radius earned its keep. If a misread is worth thousands, the same toll is a rounding error and the honest recommendation is that this control is not where the next unit of effort belongs. ## Saying it in a room Open with the reframe — it bought a cost, not immunity — then the current numbers with their columns, then the explicit list of moves that were never priced, then the recommendation with the bill attached. That sequence is what separates someone who owns a defence from someone who reports a percentage about it.
- The weights were open-sourced. Does that change what the trained radius is worth?It removes any contribution from obscurity and makes white-box the default assumption for every re-measurement. The attacker solves against final, fixed parameters, which is strictly easier than the moving target the training loop faced, so whatever robustness remains has to come from the fitted neighbourhood alone. It does not shrink the region that was bought; it just means you may not quietly count secrecy as part of the defence.
- What single reported number would you refuse to put on a slide?A bare robust-accuracy percentage. Without the access assumption, the norm, the radius and the search strength it bounds nothing and cannot be compared with any other figure, and a reader will hear it as coverage of the attack surface. Replace it with the curve across radii plus at least one row in another norm, and state the search settings beside it.
- Re-measurement shows the cliff sits just outside the trained radius. Do you retrain at a larger one?Only if the larger region describes an adversary you actually face. Widening the region raises training cost and lowers clean accuracy for all traffic, while buying nothing against someone who edits a few pixels heavily or re-shoots the photograph. Put the comparison in economic terms: what one successful misread is worth to the attacker, and whether the same spend moves their price further somewhere else in the pipeline.
A lock is rated in minutes against a named tool, not declared unbreakable — and the rating quietly falls as the tools get better and cheaper.
saying these in an interview costs you the question
- Reads robust accuracy as a wall rather than a cost imposed
- Assumes robustness degrades gently outside the trained radius
- Counts unreleased weights as part of the robustness
- Repeats last year's number from an unchanged, unstrengthened search
- Calls an empirical robustness figure a guarantee