skip to content

You build an evasion attack from an adversarial-robustness library such as Foolbox, torchattacks or the Adversarial Robustness Toolbox and pass only the wrapped model — no iteration count, no step size, no number of random restarts. Where do those values come from, and what were they chosen for?

level: juniorimportance: must knowfreq 60%

answer

  1. defaults = docs example, not assessment
  2. constructor values differ per library
  3. iterations, step size, restarts
  4. log the config next to the metric
  5. re-run at two effort levels

basics

~20 s

They come from the attack class's own constructor defaults, baked in so the library's documentation example runs fast on a laptop. They are demo settings, not assessment settings. A model that resists them has resisted only a short, weak search. Choose and record every strength argument yourself before reporting a robustness number.

solid answer

~50 s

Every attack class in these libraries ships literal default values for its search arguments — how many iterations to run, how big a step to take, how many random restarts to try. Those numbers exist so the README snippet finishes in seconds on one image, and they differ from library to library and from attack to attack even for the same published attack. So "the model survived the attack" from a defaults run means only "the model survived *that* configuration", which nobody outside your process can reconstruct from the number alone. The practical rule: pass every strength argument explicitly in the run script, even where your value equals the library default, and store the resulting configuration next to the metric. That turns a number into a reproducible claim and makes it obvious when someone later raises the iteration count and the reported robustness collapses.

code

yaml · 15 lines
yaml
evaluation:
  library: <adversarial robustness library + release>
  attack_class: <iterative gradient evasion attack>
  threat_model:
    norm: <from the threat model, not from compute>
    bound: <from the threat model, not from compute>
  effort:
    iterations: <explicit, even if equal to the library default>
    step_size: <explicit>
    random_restarts: <explicit>
  examples_attacked: <n>
  model: <name + weights hash>
  result:
    attack_success_rate: <value>
    second_effort_level_success_rate: <value>   # did the number move?

go deeper

for a junior

Knows the values came from the attack class's defaults and that they are tuned for a quick example, and knows to pass them explicitly.

for a middle

Separates threat-model arguments (bound, norm) from effort arguments (iterations, step size, restarts) and logs the whole configuration with the metric.

for a senior

Sets an effort level from the compute budget, re-runs at two levels to show the number has converged, and refuses to publish a defaults-only 'robust' verdict.

for a principal

Makes the recorded configuration a required field of any robustness claim the organisation publishes, so results across models and teams stay comparable.

### What "the default" physically is An attack in these libraries is a Python object, and its search arguments are literals sitting in that class's `__init__` signature. Construct the Adversarial Robustness Toolbox's `ProjectedGradientDescent` with nothing but a wrapped estimator and Python fills in `eps`, `eps_step`, `max_iter` and `num_random_init` from that signature. Nothing about your model, your dataset or your threat model is consulted. The wrapper you passed does declare the input range through its own `clip_values`, but that tells the attack where to clip a perturbed input, not how hard to search. `torchattacks.PGD(model)` behaves the same way with `eps`, `alpha`, `steps` and `random_start`; so does `foolbox.attacks.LinfPGD()` with `steps` and `rel_stepsize`. Those literals were chosen by the library author so the README snippet and the unit tests finish in seconds on one machine. That is the entire selection criterion. They are not "the recommended settings", they were never validated against your model class, and two libraries implementing the same published attack do not agree on them. | library | attack class | its effort arguments | the default that surprises people | |---|---|---|---| | Adversarial Robustness Toolbox | `ProjectedGradientDescent` | `max_iter`, `eps_step`, `num_random_init` | `num_random_init` defaults to **no random restart at all** — one deterministic start | | torchattacks | `PGD` | `steps`, `alpha`, `random_start` | `steps` is a small count sized for a laptop demo; `eps` and `alpha` are written as image-scale fractions of 255 | | Foolbox | `LinfPGD` | `steps`, `rel_stepsize` | the step is a *fraction of epsilon*, so it rescales when you change the bound — ART's absolute `eps_step` does not | There is a second confusion sitting under the first. The **perturbation bound** — the norm and the epsilon — is a threat-model decision: it states what an attacker is assumed able to change. The **effort arguments** — iterations, step size, restarts — are a budget decision: they state how hard you looked. Only the second group trades against compute, and only the second group is yours to raise freely. ART's default `eps=0.3` is a threat-model value sized for MNIST-scale images living in `[0, 1]`; carried unchanged onto a normalised ImageNet tensor it describes a completely different attacker. ### What a run costs Effort is multiplicative, not additive. The model calls a bounded gradient attack makes are roughly `examples × restarts × iterations`, and each call is a forward pass plus a backward pass. A thousand examples at a hundred iterations with five restarts is half a million forward-backward passes: tens of minutes to a couple of hours on one modern GPU for a mid-size vision model, and GPU-days once you are at ImageNet resolution with a large backbone. At rented rates of roughly a couple of dollars per GPU-hour that is trivial for one model and a visible line item across a portfolio evaluated every quarter. The defaults are small precisely because the library author was paying that bill in continuous integration on every commit. ### Where the number misleads A defaults run that produces no adversarial examples reads in a report as "the model is robust" and actually means "this short, single-start search did not find anything". Those two are indistinguishable in the output. An optimisation that ran out of iterations, and one that started once in a flat region of the loss and never got traction, return exactly what a genuinely hard target returns: an empty result set. With ART's `num_random_init` left at its default the reported success rate is a single draw from a lottery, presented as a property of the weights. The cross-library version is worse. Two teams both report "PGD, L-infinity, epsilon 8/255, attack success rate 12%". One ran forty steps with a step that rescaled itself when the bound was set; the other ran a hundred steps with an absolute `eps_step` now coarse for that epsilon and no restarts at all. The two numbers are not comparable, and nothing in either report tells a reader that. ### What you would check before believing it - Pass every strength argument explicitly in the run script, **including where your value equals the library default**. Defaults move under a dependency bump, and a reviewer should not have to read library source to learn how hard you searched. - Persist the configuration beside the metric: library and release, attack class, norm, epsilon, iterations, step size, restarts, examples attacked, model version and weights hash. - Re-run at two effort levels before publishing anything. If the success rate moves between them, the lower run measured your budget rather than the model. - Never rank a number produced by one library's defaults against another's; the defaults are not the same search. The sentence an interviewer is listening for: a robustness number without its attack configuration is not a result, it is an anecdote.

  • You keep the library defaults deliberately for a fast screening pass. How should that number be labelled?
    As a screening result that can only prove weakness, never strength: a hit at defaults is a real finding, but the absence of a hit means nothing until the run is repeated at assessment effort.
  • Why write out an argument explicitly even when your value equals the library default?
    Because the default can change under you when the dependency is upgraded, and because the reviewer of the report should not have to read library source to learn how hard you searched.
  • Which arguments are not yours to tune for convenience?
    The threat-model ones — the perturbation bound and its norm, and the allowed input constraints. Those come from what the attacker is assumed able to do; only the search-effort arguments trade off against compute.

saying these in an interview costs you the question

  • Treating a defaults run as evidence the model is robust.
  • Reporting an attack-success rate with no record of iterations, restarts or the bound.
  • Assuming two libraries' defaults are the same because they implement the same published attack.
  • Tuning the perturbation bound downward until the attack fails, and calling that robustness.
  • Believing defaults are 'the recommended settings' because they are what the docs show.

context