skip to content

For an iterative gradient evasion attack driven from an adversarial-robustness library such as the Adversarial Robustness Toolbox or torchattacks, what does raising each of these three arguments buy you — the iteration count, the step size relative to the perturbation bound, and the number of random restarts — and what does each cost?

level: middleimportance: must knowfreq 55%

answer

  1. iterations = refinement
  2. step size coupled to the bound
  3. restarts = start-point lottery
  4. cost multiplies: iters x restarts x examples
  5. raise until the success rate plateaus

basics

~20 s

More iterations let the search refine longer inside the same bound. A step size too large for the bound overshoots and one too small never reaches its edge. More restarts re-launch from fresh random starting points, so one unlucky start is not read as robustness. Each knob multiplies GPU time roughly linearly.

solid answer

~50 s

Think of the three as buying different things. - **Iterations** buy refinement: the search takes more, smaller corrections inside the allowed region. Cost is linear in GPU time per example. - **Step size** is not a strength knob on its own — it has to be set relative to the bound and the iteration count. Too large and the search bounces off the boundary without settling; too small and it never traverses the region within the iterations you allowed. This is the argument people most often leave at a default that no longer matches a bound they changed. - **Restarts** buy independence from where the search began. A single start can land in a flat region and report failure; several independent starts make that far less likely, at a cost multiplied by the restart count. A sound practice is to raise iterations and restarts together until the measured success rate stops moving, then report the configuration at which it plateaued.

go deeper

for a junior

Knows the three arguments exist, that they control how hard the attack tries, and that raising them costs compute.

for a middle

Explains what each one buys, that step size is coupled to the bound, and that restarts address unlucky starting points rather than search depth.

for a senior

Runs an effort sweep, reports the plateau configuration, and trades sample size for search strength when the budget is fixed.

for a principal

Sets the plateau-sweep procedure as the house standard so the effort level behind every published number is defensible and comparable.

### The loop these three arguments drive An iterative gradient evasion attack is a short optimisation loop: start somewhere inside the region the perturbation bound allows around an input, compute the gradient of a loss with respect to that input, take a step along it, project back inside the region, repeat. The three arguments in the question are the only handles you have on that loop. Each one closes off a distinct way the loop can fail for a reason that has nothing to do with the model — which matters because the library reports a failed optimisation and a genuinely hard target identically, as "no adversarial example found". ### Iterations — refinement `max_iter` in the Adversarial Robustness Toolbox's `ProjectedGradientDescent`, `steps` in `torchattacks.PGD` and in `foolbox.attacks.LinfPGD`. More iterations mean more, smaller corrections, so the search can settle at the edge of the allowed region instead of stopping wherever the budget ran out. The tell that you have too few is direct: double the count and watch whether the attack-success rate is still climbing. Cost is linear in the count. ### Step size — coupled to the bound, never independent `eps_step` in ART, `alpha` in torchattacks, `rel_stepsize` in Foolbox. This is not a strength knob on its own; it is only meaningful relative to the bound and the iteration count. The product of step and iterations has to be able to traverse the region epsilon defines, and the step has to be fine enough to settle inside it rather than bouncing off the projection on every pass. The libraries genuinely differ here, and the difference is where the silent misconfiguration lives: | library | how the step is expressed | what happens when you tighten epsilon | |---|---|---| | Foolbox `LinfPGD` | `rel_stepsize`, a fraction of the epsilon passed at attack time | the step shrinks with the bound automatically | | ART `ProjectedGradientDescent` | `eps_step`, an absolute value | the step stays put and is now coarse relative to the smaller `eps` | | torchattacks `PGD` | `alpha`, an absolute value | same as ART — it does not follow `eps` | Tighten epsilon for a stricter threat model in ART or torchattacks, leave the step alone, and every iteration now overshoots the region and is projected straight back. The model looks more robust, and the entire gain is configurational. ### Restarts — independence from where you began `num_random_init` in ART, `random_start` in torchattacks and Foolbox (with the run repeated to get several draws). Each restart is an independent draw of a starting point inside the allowed region. On an uneven loss surface some starts find a descent direction and some sit in a flat patch and go nowhere; with one start you have sampled that lottery once and reported the outcome as a property of the weights. ART's `num_random_init` defaults to zero, so an ART defaults run is a single deterministic start. Restarts are also the most expensive knob, because they multiply the whole loop rather than lengthening it. ### What it costs Model calls are roughly `examples × restarts × iterations`, each a forward and a backward pass. A thousand examples at a hundred iterations with five restarts is half a million forward-backward passes — tens of minutes to hours on one modern GPU at CIFAR scale, GPU-days at ImageNet scale with a large backbone, low single-digit dollars per GPU-hour on rented capacity. Doubling all three at once is an eight-fold budget request, which is exactly why nobody should raise them blindly. ### Where the number misleads The dangerous reading is a low attack-success rate quoted as robustness when the success-versus-effort curve was still climbing at the configuration used. That number is a statement about your compute budget wearing the clothes of a statement about the model. The second misreading is treating restarts as noise reduction: they are not averaging away variance in a measurement, they are additional independent attempts, and omitting them does not make the estimate noisier, it makes it *biased low*. The third is subtler. If you spend a fixed budget on many examples at default effort, you get a tight confidence interval — around the wrong quantity. Precision on a weak attack is not accuracy about the model. ### What you would check Run a small effort sweep on a subset before the headline run: a handful of `(iterations, restarts)` pairs, each roughly doubling the previous, plotted against attack-success rate. The point where the rate stops moving materially is your assessment configuration. Everything above it is wasted budget; everything below it is an under-strength result. Then, when budget binds, cut the number of examples rather than the effort — and if you changed the bound at any point, re-derive the step size for the new epsilon rather than inheriting the old one.

  • You tighten the perturbation bound for a stricter threat model and change nothing else. What silently breaks?
    The step size, if it came from a default matched to the older, larger bound. It now overshoots the allowed region every iteration and the attack looks weaker for a purely configurational reason.
  • How do you know you have spent enough on iterations and restarts?
    Run an effort sweep on a subset: raise the pair and watch the attack-success rate. When doubling both no longer moves it materially, you are at the plateau, and that configuration is what you report.
  • Given a fixed GPU allowance, would you rather attack more examples or search harder on fewer?
    Search harder on fewer, until the plateau. A weak search over a big sample gives a tight confidence interval around a number that measures the attack, not the model.

Restarts are lottery tickets, iterations are how long you spend working each one, and examples are how many people you buy tickets for — and the bill is the product of all three, not the sum. That is why "just double everything" is an eight-fold budget request, and why one team's PGD number is not another's.

saying these in an interview costs you the question

  • Treating step size as independent of the perturbation bound.
  • Raising the iteration count while leaving a single start, then calling the result converged.
  • Picking a round number of iterations with no sweep showing the number stopped moving.
  • Spending the whole budget on more examples rather than on stronger search.
  • Describing restarts as 'just averaging noise' rather than as independent attempts.

context