A client asks you to shortlist defences for an image classifier using an adversarial-robustness toolkit, and you have one GPU for two weeks. The candidates split into defence objects that attach to the existing weights and trainer-style defences that need a full retrain per hyperparameter setting. How do you allocate the compute, and what do you tell the client the shortlist does and does not cover?
answer
- split by artefact cost, not taxonomy
- config defences: sweep wide
- trainer variants: two or three, each a baseline
- budget attack re-tuning, not just retrains
- publish the not-evaluated list
basics
~20 sSplit by cost. Attach-only defences reuse the trained weights, so each costs one evaluation sweep and can face many attack settings. Every trainer variant costs a full retrain, so you can afford only a couple of settings. Spend most of the compute on deep adaptive evaluation of a short list, not a broad shallow sweep.
solid answer
~60 sThe cost asymmetry drives the plan. Attach-and-evaluate defences are cheap per candidate; retrain-required defences cost a training run per setting before you can evaluate anything, and their strength is hyperparameter-sensitive, so one setting is not a fair test of the family. A workable allocation: - **Screen broadly, cheaply.** Run every attach-only candidate against a fixed attack set to eliminate the obviously weak ones and to measure their clean-accuracy and latency cost. - **Spend the retrains deliberately.** Pick at most one or two trainer-style candidates and give each enough settings to be represented honestly, rather than one run apiece for four families. - **Reserve the majority of the remaining compute for adaptive evaluation** of the two or three finalists: re-tuned attacks, budget sweeps, a gradient-free class and a transfer attack. The caveat to the client is the important deliverable. The ranking holds for the attack classes in this library, at this perturbation budget, on this data subset, at the training recipe the budget allowed. It is a shortlist, not a robustness certificate, and every number is an upper bound.
go deeper
Recognises that retraining costs much more than attaching a preprocessor, and that clean accuracy must be measured too.
Builds the screen-then-deepen plan and states the per-candidate cost model correctly.
Reserves the bulk of compute for adaptive evaluation of finalists and refuses to report a single undertrained run as a family's score.
Owns the whole tradeoff: what the client can operate, what the ranking is conditional on, and what the report may not claim, including that robust accuracy is an upper bound.
**Turn the brief into an arithmetic problem first** One GPU for two weeks is about 336 GPU-hours, and realistically nearer 300 once you subtract setup, data preparation, failed runs and the hours nobody is watching the machine. Everything else follows from putting a price on each unit of work. - *An evaluation sweep.* A 50-step iterative attack with a handful of restarts over a couple of thousand test images is on the order of half a GPU-hour to an hour on a CIFAR-scale ResNet. Cheap enough to run dozens. - *A black-box cross-check.* Priced in queries, not steps: a score-based attack at a few thousand queries per image over 1,000 images is millions of forward passes, so a few GPU-hours each. If the target is a metered hosted endpoint instead of local weights, this is money rather than time, and it must be quoted before it is run. - *A trainer-style retrain.* Adversarial training with a k-step inner attack costs roughly k+1 times ordinary training. A CIFAR ResNet that trains in about 1.5 GPU-hours becomes 10 to 15 hours per hyperparameter setting. Four settings is two full days of the fortnight gone before a single robustness number exists. That asymmetry — one sweep is under an hour, one retrain is half a day — is the entire plan, and it is a fact about the *artefact* each defence kind produces, not about the defence's taxonomy label. **Three buckets** 1. *Screening, a small share (say 30 hours).* Every attach-only candidate — preprocessors, postprocessors, an already-fitted detector — against a fixed attack set at a fixed budget. Output: a provisional ranking you do not yet trust, plus each candidate's clean-accuracy loss and added latency, which are the numbers the business will actually feel. 2. *Retrains, a bounded share decided up front (say 50 to 60 hours).* Trainer-style candidates. A defence family's strength is hyperparameter-sensitive, so one run does not represent a family. If you cannot afford several settings for a family, do not run one and report it — say the family was not evaluated. 3. *Adaptive evaluation, the largest share (the remaining 150 hours or so).* Finalists only, two or three of them, attacked properly: hyperparameters re-tuned against the defended chain, a perturbation-budget sweep, at least one attack class that does not rely on gradients through the defence chain, and a transfer attack from a surrogate. **Where the intuitive allocation goes wrong** A broad shallow sweep looks like more work and produces worse information. Every candidate in it was measured with an attack that was never tuned against it, so every candidate is flattered, and the ranking that results is mostly noise about which defence happens to inconvenience your default optimiser. Depth on a short list beats breadth on a long one, which inverts what most plans do. **Where the number misleads** *It is conditional on the catalogue.* The ranking holds for the attack classes this library implements. A defence that ranks first can fail outright to a class the library does not ship, or to one published next quarter, and robust accuracy can only go down when that happens. It is an upper bound, never a certificate. *One retrain is not a family.* Reporting a single undertrained run as a defence family's score is the most common way this deliverable lies, because readers rank whatever numbers are in the table regardless of the footnote. *Robust accuracy is the wrong single column.* Clean accuracy loss, added latency and, for detectors, false positives on legitimate traffic are what the client pays daily. And there is an operability axis the number cannot show: attach-only defences ship without a retraining pipeline, while trainer-style defences require the client to own training data, training compute and the discipline to redo the defence after every model refresh. A defence the client cannot maintain is not a recommendation, even when it tops the table. **What to tell the client** Say what the shortlist covers: these candidates, these attack classes, this perturbation budget, this data subset, and for the retrained ones this training recipe, chosen because it fits the compute you bought. Say what it does not cover: the families you did not evaluate and why, the attack classes outside the library, and the fact that every number is a ceiling that a stronger attack lowers. Publishing the not-evaluated list is what separates a defensible shortlist from a number the client quotes back at you in a year as a guarantee.
- You can afford four retrains total. Two trainer-style families are on the list. What do you do?Give one family all four settings and report the other as not evaluated, or drop one family. Two settings each produces two numbers that misrepresent both families, and a reader will still rank them.
- Why can a defence with a slightly worse robust accuracy still be the right recommendation?Because it attaches to existing weights, so the client can ship and maintain it without owning a retraining pipeline. A defence that must be redone after every model refresh has an operational cost the robust-accuracy column does not show.
saying these in an interview costs you the question
- Spreading the compute evenly so every candidate gets one shallow, untuned evaluation.
- Reporting one retrain at one hyperparameter setting as a defence family's score.
- Presenting the shortlist as a robustness guarantee rather than a result conditional on the attacks run.
- Ignoring whether the client can maintain a retrain-based defence after each model refresh.
- Omitting clean accuracy, latency and false-positive costs from the comparison.