skip to content

An evasion run driven from an adversarial-robustness library finishes and reports zero successful adversarial examples at the configured perturbation bound. Before you write 'the model is robust', what checks do you run on the attack configuration and the model wrapper itself?

level: seniorimportance: must knowfreq 50%

answer

  1. zero hits = suspect the harness
  2. relax the bound: must reach near-total success
  3. wrapper input range and preprocessing
  4. sweep effort, watch the curve
  5. weak baseline beating the gradient attack

basics

~20 s

Assume the harness is broken first. Re-run with the bound relaxed until the attack must succeed; if it still fails, the wrapper or the loop is wrong. Then sweep iterations and restarts upward and watch the success rate. Also check a cheap non-gradient baseline: if it beats the gradient attack, the gradients are unreliable.

solid answer

~60 s

Zero hits is the result most likely to be an artefact, so treat it as a bug report against your own harness. 1. **Relax the bound to something absurd.** At a large enough perturbation the attack should reach near-total success. If it does not, nothing downstream is trustworthy — the wrapper, the loss direction or the label handling is wrong. 2. **Check the wrapper contract.** Input range, whether preprocessing sits inside or outside the wrapped model, and whether the model is in evaluation mode. A perturbation applied in the wrong space is not the one you asked for. 3. **Sweep effort.** Raise iterations and restarts; if the success rate climbs, your original run was under-configured. 4. **Run a cheap non-gradient baseline** at the same bound. If random or transferred inputs beat the gradient attack, the gradient signal is not usable and the run needs a different attack family before any robustness claim is made. Only a zero that survives all four is worth reporting, and it is reported with the configuration attached.

go deeper

for a junior

Knows to double-check the run before believing zero hits, and can re-run with a bigger bound to see if anything succeeds.

for a middle

Checks the wrapper's input range and preprocessing and raises iterations and restarts before accepting the result.

for a senior

Runs the whole ladder — forced success, wrapper contract, effort sweep, non-gradient baseline — and reports the zero narrowly, with the sanity artefacts attached.

for a principal

Makes the forced-success artefact a required attachment for any negative robustness result, so no team can publish an untested instrument's silence.

### Two failure modes, one output A robustness evaluation has exactly two ways to report zero successful adversarial examples: the model held, or your instrument never fired. The library cannot distinguish them for you. A gradient that never reached the input, a loss pushed in the wrong direction, a label array misaligned with the batch, and a genuinely hard target all return the same empty result set. The second class of cause is far more common than the first in a freshly written harness, so a zero is treated as a bug report against your own setup until it has survived a ladder of deliberate attempts to make the attack succeed. ### 1. Force a success Relax the bound to an absurd value, or run an unconstrained attack, on ten or twenty examples. A working evasion setup drives almost everything to misclassification when the perturbation is effectively unlimited. If it does not, the model is not robust — it is unreachable, and the fault is plumbing. Suspects, in the order they usually bite: gradients not flowing to the input, because the wrapper broke the graph or the input tensor never had gradients enabled; the objective being maximised in the wrong direction, or against labels misaligned with the batch; and the model left in training mode, so dropout and batch-norm running statistics make every call nondeterministic. This check costs seconds to minutes on a handful of examples — the cheapest rung on the ladder, and the one most often skipped. ### 2. Check the wrapper contract These libraries reach your model through an adapter that declares what space the attack is operating in. In ART, `PyTorchClassifier` takes `clip_values`, which fixes the valid input range, and `preprocessing`, a mean-and-standard-deviation pair the wrapper applies itself. Get the division of labour wrong and the bound is enforced in a space your report does not describe: the attack perturbs an already-normalised tensor while your epsilon was quoted in raw pixel units, or it perturbs raw input that your own pipeline then rescales, shrinking the perturbation away before the model ever sees it. Either way, the effective search is not the one written up. Verify empirically, not by reading code: take a produced example, subtract the original, and measure the difference **in the units and the norm your threat model uses**. It should sit at or just under the configured epsilon. Orders of magnitude off in either direction is a wrapper-contract mismatch, not a modelling result. ### 3. Sweep the effort Raise iterations and restarts together on a subset and plot success against effort. A curve still climbing at your original configuration means the number you were about to publish measured your compute budget. Note that with ART's `num_random_init` at its default of zero you never had a random restart at all, so the original run was a single deterministic draw. ### 4. Cross-check with an attack that uses no gradients Attack the same examples at the same bound with something that never touches a gradient: random perturbations sampled at the bound, or examples crafted against a different model and transferred. If the gradient-free baseline beats your gradient attack, the gradient signal is not usable on this model. That is a finding about your evaluation instrument, and it belongs in the report as such — escalated to the specialists who study that class of defence behaviour — rather than being written up as a safe model. ### What the ladder costs All four rungs run on a subset, typically tens of examples, so the whole ladder is a small fraction of the headline run's compute — usually well under an hour of engineer-supervised GPU time. The expensive part is the effort sweep, and even that is bounded by running it on a subset rather than the full evaluation sample. Against that, the cost of skipping it is a published robustness claim that a reviewer can overturn by re-running with one extra flag. ### What the surviving zero is then allowed to say Even a zero that clears every rung is narrow. It says: *this attack family, at this effort, at this bound and norm, on this sample of inputs, through this wrapper, found nothing.* Write it in exactly those terms and attach the forced-success artefact, because that artefact is the only evidence anyone has that the instrument was capable of firing. A negative result whose instrument was never tested is not a weak result — it is not a result.

  • The unbounded run still produces almost no successes. What are your first three suspects?
    Gradients not flowing to the input (wrapper or graph problem), the objective being pushed in the wrong direction or against misaligned labels, and the model not in evaluation mode so its behaviour is nondeterministic.
  • How do you check that the perturbation you asked for is the perturbation applied?
    Measure the difference between a produced example and its original in the same units and norm the bound is expressed in, and confirm it sits at or below the bound rather than orders of magnitude off.
  • Random perturbations at the same bound beat your gradient attack. What do you do with that?
    Stop treating the gradient result as a robustness measurement, report the stronger baseline's number, and escalate: the gradient signal is unusable here, which is a finding in itself and not evidence of a safe model.

A silent smoke alarm means either that there is no fire or that the battery is dead, and the readout is identical in both cases. The forced-success run is the test button, which is why you press it before writing down that the building is safe.

saying these in an interview costs you the question

  • Publishing a zero-success run without ever verifying the attack can succeed at all.
  • No check that the perturbation was applied in the same units as the stated bound.
  • Ignoring that a random-noise baseline outperformed the gradient attack.
  • Explaining the zero purely as model quality without touching the configuration.
  • Leaving the model in training mode, or preprocessing outside the wrapper, and never noticing.

context