After you attach a defence object to a wrapped classifier in an adversarial-robustness toolkit, the robust accuracy under your configured attack rises from near zero to about three quarters. What do you check before reporting that as a robustness gain?
answer
- attack tuned on the wrong surface
- unbounded budget must reach the floor
- add a gradient-free and a transfer attack
- robust accuracy is a minimum over attacks
- clean accuracy is the price tag
basics
~20 sAssume the attack broke rather than the model became robust. Re-tune the attack against the defended pipeline, raise iterations and restarts, and check that an unbounded perturbation still drives accuracy to zero. Compare against a query-based attack and a transfer attack. Confirm the defence really sits inside the wrapper you attacked.
solid answer
~60 sThe number came from one attack instance that you tuned against the undefended model. A defence that merely blinds that instance produces exactly this jump, so the checks are about attacking the defended pipeline on its own terms. 1. **Re-tune, do not reuse.** Step size, iteration count, restarts and initialisation were chosen for a different loss surface. Re-search them against the defended chain. 2. **Sanity floor.** With the perturbation budget raised far beyond the threat model, robust accuracy must collapse to roughly chance or zero. If it does not, the harness is failing to optimise, and the result is about your gradients, not the model. 3. **Cross-check with an attack that needs no gradients.** Run a query-based or decision-based attack class, and a transfer attack from a surrogate. If a weaker attack beats your strong one, the gradient signal is being degraded rather than the model defended. 4. **Confirm placement.** Verify the defence executed on the prediction path during the attack run. 5. **Price it.** Report clean accuracy loss, and for a detector its clean false-positive rate.
go deeper
Should at least say the number needs a second attack and that clean accuracy has to be reported alongside it.
Re-tunes the attack against the defended chain and runs the budget and iteration sweeps rather than reusing the earlier configuration.
Runs the full adaptive loop: placement check, budget sweep, gradient-free and transfer cross-checks, provenance of the defence's own training data, and reports the minimum over attacks.
Turns the checks into a standing acceptance rule for defence claims, and sets what the team may and may not tell a client about a robustness number.
## Why the default reading is wrong An **attack instance** is a configured optimiser: a step size, an iteration count, a number of random restarts, an initialisation, a loss. You chose those values against one loss surface — the undefended model. Attaching the defence changed the surface. A jump from near-zero to about 75 percent **robust accuracy** is therefore ambiguous by construction: it is equally consistent with - "the model got harder to fool" and with - "my optimiser stopped working here". The burden of proof sits with whoever claims the first reading, and the difference is decided by evidence you have to go and collect. ## The checks, cheapest first 1. ***Placement.*** Confirm the defence object actually executed on the estimator's prediction path during the attack run — probe the wrapper and inspect what reached the model, rather than trusting that the constructor accepted the argument. A defence sitting outside the queried chain, or one carrying `apply_predict=False`, yields an undefended baseline dressed as a defended result. Cost: minutes. It explains a surprising share of these jumps. 2. ***Budget sweep.*** Plot robust accuracy against the perturbation budget, sweeping it far past the stated threat model. A genuine defence degrades smoothly and reaches roughly chance once the budget is large enough, because at a sufficient budget any input can be turned into any class. A curve that stays flat as the budget grows is the classic **harness-failure signature**: the attack is not finding perturbations that provably exist. Cost: a 50-step attack over 1,000 images at eight budgets is a few hundred thousand forward-and-backward passes, well under a GPU-hour on a small model. 3. ***Iteration and restart sweep.*** Push iterations and random restarts up until robust accuracy plateaus. If it is still falling when you stop, your headline number is simply an under-run attack. Cost scales linearly; a 5x restart budget is a 5x sweep. 4. ***Attack diversity.*** Add at least one class that does not need gradients through the defence chain — a decision-based or score-based black-box attack — plus a transfer attack from a surrogate model. This is the check with a real price tag, because black-box attacks are priced in queries: a decision-based attack commonly needs tens of thousands of queries per image, and a score-based one a few thousand, so 1,000 test images is millions of model calls. Against a local GPU that is hours; against a metered hosted endpoint it is a line item you must budget for and quote in the report. 5. ***Provenance of the defence.*** If the defence is a detector or a trainer that was itself fitted using adversarial data from the same attack class and budget you are now testing with, the evaluation is circular and the number means "this classifier recognises its own training generator". 6. ***Price.*** Clean accuracy after the defence, added latency, and for a detector the clean false-positive rate at the threshold used. ## Where the number misleads The single most dangerous reading is treating robust accuracy as a property of the model. It is a property of the pair (model, attack you ran), and it is a *ceiling*: a stronger attack can only lower it, never raise it. - So the correct aggregate over several attacks is the **minimum**, not the mean and not the one the defence was designed for. Averaging is actively perverse, because it rewards adding weak attacks to the set. - The second misleading reading is the **inverse-strength result**: if a black-box query attack beats your white-box gradient attack, that is not a curiosity. A white-box attack has strictly more information and should not lose; when it does, the gradient signal it relies on is being degraded, and the white-box number is an artefact of the defence's effect on your optimiser rather than a measurement of the model. ## What to write down - Every attack class attempted and its tuned hyperparameters, including the ones that failed. - The budget sweep as a curve, not a point. - Whether each attack differentiated through the defence chain or used an estimated gradient. - The clean-accuracy and latency price. - And one sentence stating that the reported robust accuracy is the minimum over the attacks actually run, at a named budget and threat model, and is an **upper bound** on true robustness. The theory of why certain defences systematically degrade gradients belongs to the adversarial-ML defence literature; the tool-side obligation is to run the checks that would have exposed it and to say plainly which ones you ran.
- The robust accuracy stays flat as you increase the perturbation budget far past the threat model. What does that tell you?That the attack is failing to optimise, not that the model is robust. At a large enough budget any classifier can be driven to the floor, so a flat curve indicts the harness or the gradients.
- Why is a black-box query attack outperforming your white-box gradient attack a warning sign rather than a curiosity?A white-box attack has strictly more information, so it should not lose. When it does, the gradient signal it relies on is degraded, and the white-box number is an artefact of the defence's effect on the optimiser.
- How should the final robust accuracy be phrased in the report?As the minimum over the attacks actually run, at a named perturbation budget and threat model, and explicitly as an upper bound on true robustness that a stronger future attack can only lower.
The attack you tuned is a key cut for one particular lock. Change the lock and the key stops turning — which tells you the key no longer fits, not that the door cannot be opened.
saying these in an interview costs you the question
- Reporting the jump as a robustness gain from a single attack configuration.
- Reusing attack hyperparameters tuned against the undefended model.
- Never sweeping the perturbation budget to confirm accuracy reaches the floor.
- Ignoring that the defence was fitted on adversarial data from the same attack class being evaluated.
- Omitting the clean-accuracy cost of the defence.