skip to content

A robustness write-up says an evasion attack "failed" against a model, but the run was driven through a command-line wrapper that exposes only some of the underlying attack library's parameters. Why is "the attack failed" not yet a supportable claim, and what do you do before it goes in a report?

level: seniorimportance: must knowfreq 48%

answer

  1. negative result, weak evidence
  2. defaults you never chose
  3. sanity-check against a weak baseline
  4. sweep the allowance, report a curve
  5. cite library, attack, settings, budget

basics

~20 s

Because the run only tested the attack at the settings the wrapper exposed. An attack with an untuned step size, iteration count or perturbation budget failing says nothing about robustness. Before publishing, check which parameters the layer passes through, sweep the ones that matter, and record the exact settings tested.

solid answer

~50 s

A failed attack is evidence about *that configuration*, not about the model. Adversarial attacks are optimisation procedures whose success is sensitive to configuration, and a wrapper exposing only a slice of the library's parameters silently pins the rest to defaults chosen by someone with no knowledge of your model. The claim you can support is narrow: "this attack, from this library, at these settings, against this model interface, did not produce a successful example within this budget." What to do: - Enumerate which parameters the wrapper passes through and which it fixes. If that is not visible, drop to the library and call the attack yourself. - Sweep the parameters governing attack effort and allowance rather than accepting one point. - Confirm the failure is not an integration artefact: a wrong preprocessing or gradient path makes any attack "fail" for reasons unrelated to robustness. - Record the library, attack, every setting and the budget, so a reader can reproduce or contest the number.

go deeper

for a junior

Should at least say a failed attack might just be badly configured, and that the settings need to be written down.

for a middle

Explains that unexposed parameters take library defaults, and that effort and allowance parameters must be swept rather than sampled at one point.

for a senior

Adds the integration sanity check against a weak baseline, drops to the library when the layer hides configuration, and phrases the claim as bounded with full provenance.

for a principal

Sets the standard: what an evaluation must record to be citable, and who may publish a robustness claim from a run at wrapper defaults.

The most common way an adversarial-robustness evaluation goes wrong is a comfortable negative: the attack did not succeed, so the model is called robust. Driving the attack through a wrapper makes this failure mode easier to reach, because the settings the run actually used are precisely the ones that never appeared on your command line. **Why a negative is weak evidence.** An adversarial attack is a search under a constraint. Projected gradient descent, for instance, repeatedly steps in the gradient direction by `eps_step` and projects back into a ball of radius `eps`, `max_iter` times, from `num_random_init` starting points. A search that returns nothing may have failed because the target genuinely resists, or because it was given too few steps, too small an allowance, one unlucky start, or a gradient path that never reached the model. Only the first is a property of the model, and the harness's result record — outcome without configuration — cannot distinguish them. **The specific hazard of the wrapper layer.** The parameter surface is curated: what `show options` prints is settable, everything else takes the underlying library's default. Two consequences follow. First, "the same attack" run through two versions of the wrapper, or through the wrapper versus a direct library call, can differ in settings nobody chose. Second, the defaults are domain conventions, not facts about your system: ART's `eps=0.3` assumes inputs normalised to [0, 1], so against a target taking 0-255 pixels or raw tabular features the search is confined to a ball too small to alter anything, and the whole table reads "failed" for reasons that have nothing to do with robustness. **The other way the number misleads: the denominator.** Attack-success rate is a fraction, and the harness rarely tells you which fraction. If the denominator is *all* test inputs, samples the model already got wrong before any perturbation count as successes and the rate is inflated. If it is *originally correct* inputs — the robust-accuracy convention — the same run yields a different, smaller number. A reported 92% and a reported 78% can be the identical run under two accounting rules. Any cross-run or cross-team comparison that does not fix the denominator, the success criterion, and the perturbation budget is comparing nothing. **What it costs to do properly.** Sweeping rather than sampling has a price you should state up front. Six values of `eps` times 200 inputs times PGD at `max_iter=40` with five restarts is roughly 240,000 forward-and-backward passes — trivial on a GPU, an afternoon on CPU. Against a black-box hosted endpoint the same discipline is brutal: ART's `HopSkipJump` at its defaults (`max_iter=50`, `max_eval=10000`) is tens of thousands of billed calls *per input*, so a six-point sweep over 200 inputs is not a plan, it is a budget request. That constraint is itself a finding — say in the report that the black-box budget capped the sweep, rather than letting a truncated run read as a strong negative. **The procedure before publishing.** 1. *Establish what was actually run.* Which library, which version, which attack, which arguments passed through and which took defaults. `pip freeze` in the harness's environment plus a diff of `show options` against the library constructor answers this. If the layer will not tell you, import the library and call the attack yourself. 2. *Rule out integration failure first.* Re-run the identical pipeline against an undefended reference model, or hand the attack an absurdly large `eps`. If it still cannot win, the failure is in preprocessing, input scaling, `clip_values`, or the gradient path — not in the target. 3. *Sweep, do not sample.* Vary budget and effort across a range and report the curve. Robustness is a function of the allowance, not a boolean. 4. *Report the whole configuration.* Library and version, attack, every setting, access assumptions, query budget, denominator, and how success was decided. **How to word the claim.** "This attack, from this library at this version, at these settings, against this model interface, within this budget, succeeded on N of M originally-correct inputs." Never upgrade that into "the model is robust". If the report needs a general statement, the evaluation has to be designed to support one, and a single run at wrapper defaults never is.

  • What is the cheapest check that a 'failed attack' is not a plumbing bug?
    Run the same pipeline against a deliberately weak or undefended reference model. If the attack still fails there, the integration is broken — preprocessing, input scaling, or the gradient path — not the model robust.
  • The wrapper's output shows outcome but not configuration. What do you do?
    Treat its result as triage only, and re-run the reportable case by importing the library directly, where you set and can record every parameter.
  • How should a robustness result be phrased so it stays defensible?
    As a bounded statement: this attack from this library, at these settings and this budget, against this model interface, achieved this success rate. Never as an unqualified claim that the model is robust.

A metal detector that beeps at nothing is not evidence the field is empty until you have buried a coin and confirmed it beeps. Re-running the same attack pipeline against a deliberately weak model is that buried coin.

saying these in an interview costs you the question

  • Reading a failed attack as proof of robustness.
  • Not knowing which parameters the layer fixed versus passed through.
  • Never checking that the attack can succeed against a deliberately weak setup, so integration bugs read as robustness.
  • Reporting a single point instead of behaviour across the allowance.
  • Publishing a number with no record of the library, attack and settings used.

context