skip to content

Rather than choosing iteration counts and restarts yourself, a teammate proposes measuring robustness with a standardized evaluation suite such as RobustBench, which pins the attacks and their settings for you. What does that fix about the defaults problem, and what does it still not tell you about your own deployed model?

level: middleimportance: should knowfreq 35%

answer

  1. fixed protocol = comparable, ungameable
  2. one norm, one dataset, one interface
  3. pipeline and constraints not modelled
  4. pinned attack set ages
  5. run both, study the disagreement

basics

~20 s

It fixes comparability: everyone runs the same fixed attack set at the same effort and threat model, so numbers can be ranked and nobody quietly under-configures. It does not fix relevance. It measures one threat model on one dataset through a required interface, not your inputs, your preprocessing pipeline or the attacks your product actually faces.

solid answer

~50 s

A standardized suite is a different answer to the same problem: instead of you justifying an effort level, the suite freezes attacks, settings and threat model so that two models' numbers mean the same thing. That removes the biggest failure of a defaults run — an under-strength search dressed up as robustness — and it removes your incentive to tune until the result looks good. What it costs is scope. The suite's threat model is its own: a fixed norm and bound, a fixed dataset, a fixed model interface. Your production model has a preprocessing chain, an input distribution, domain constraints on what an attacker can actually change, and possibly a guard in front of it — none of which the suite models. A strong number there is evidence about the model's weights under that one threat model, not a statement about the deployed system. Use both: the suite for a comparable, hard-to-game headline number, your own configured runs for the threat model you actually face.

go deeper

for a junior

Knows a standardized suite exists and that it removes the need to pick attack settings yourself.

for a middle

Explains the comparability-versus-relevance trade and names what the suite's fixed threat model, dataset and interface leave out.

for a senior

Runs both, defines which number is reported where, and treats a gap between them as a signal about benchmark-shaped hardening.

for a principal

Sets policy on which number is external and which is internal, and owns the refresh cadence for a pinned attack set that ages.

### What a standardized suite actually does RobustBench is a fixed protocol wearing the clothes of a library. Its `robustbench.eval.benchmark` entry point takes your model plus a `dataset`, a `threat_model` (for example L-infinity), an `eps` and an `n_examples`, and runs AutoAttack against it. AutoAttack in its standard configuration is not one attack but an ensemble run in sequence: a step-size-free gradient attack on the cross-entropy loss, a targeted variant of the same repeated across target classes, a boundary-seeking attack, and a query-based black-box search. Examples broken by an earlier stage drop out; survivors go on to the later, more expensive ones. The suite also fixes the model interface — inputs in `[0, 1]` with any normalisation living *inside* the model — and publishes the resulting accuracy on a leaderboard. The design intent is that there is no effort argument for you to set. That is the whole point. ### What that fixes about the defaults problem Three things at once, and they are worth separating. 1. **Under-configuration.** Nobody can quietly ship a short, single-start search and call the result robustness, because the search is not theirs to choose. 2. **Per-model tuning to taste.** The incentive to nudge iterations downward until a number flatters your own model disappears when the protocol is external and pinned. 3. **The single-attack failure mode.** An ensemble means one attack failing on one model does not read as robustness; something else in the sequence still gets a turn, including a gradient-free component that keeps working when the gradient signal does not. ### What it costs Substantially more than a hand-configured run at the same epsilon — commonly one to two orders of magnitude more model calls than a single bounded gradient attack, because the targeted stage repeats the search once per target class and the query-based stage spends thousands of queries on every example that survived. Wall-clock is typically hours of single-GPU time for a full test set on a CIFAR-scale model, and the cost is *data-dependent*: a fragile model is cheap to evaluate because almost everything falls at the first stage, while a genuinely robust model is expensive because every example reaches the end of the sequence. Budget accordingly, and be aware that the models you most want to measure are the ones that cost the most to measure. ### Where the number misleads The number is an accuracy figure for **those weights, on that dataset, under that one norm at that one epsilon, behind that required interface**. Four gaps follow, and each is a place someone reads it as more than it is. - **Threat model.** The suite picks a norm and an epsilon. Your attacker may be able to change only some features, or to change them in ways no ball of that shape describes — a masked or unit-constrained attacker is simply not what was measured. - **Pipeline.** Production has resizing, quantisation, caching, maybe an input filter in front. The suite deliberately removes all of it so that models are comparable, which means the figure is about weights, not about a deployed system. - **Distribution.** The benchmark dataset is not your traffic. - **Freshness.** A pinned attack set is a snapshot of a moving field. New attacks appear and, by construction, a frozen suite does not contain them, so a strong score ages quietly. It also creates the obvious pressure: a model can be hardened against the pinned ensemble specifically. The compound misreading, and the one that gets into slide decks, is comparing your own configured run's success rate against a leaderboard accuracy as if they were the same metric. They are not even the same direction, let alone the same threat model or sample. ### What you would check Run both, and treat the disagreement as the interesting result. Use the suite for the comparable, hard-to-game external number, and your own configured runs — your constraints, your preprocessing in the loop, your input distribution, effort swept to a plateau — for the threat model you actually defend against. Before quoting a leaderboard figure, confirm your model really satisfies the suite's interface contract (inputs in `[0, 1]`, normalisation inside the model), because a model that silently violates it produces a number that is not measuring what the leaderboard column says. A model that scores well on the suite and falls quickly to your constrained, pipeline-aware run has been hardened for the benchmark's threat model rather than yours. That is a finding, not a contradiction.

  • A model scores well on the standardized suite and falls quickly to your own constrained run. What is the likely story?
    It was hardened for the suite's threat model — one norm, one bound, that dataset — while your run varies what an attacker can really change and puts the production pipeline in the loop.
  • Why is a fixed suite harder to game than a self-configured run?
    Because the effort and attack set are not yours to choose, so the usual lever — quietly under-configuring the search until the model looks robust — is unavailable.

A standardized suite is a driving test: everyone drives the same route under the same examiner, which is exactly what makes the scores comparable between candidates. It is also exactly why a pass tells you nothing about how someone handles the road they actually commute on.

saying these in an interview costs you the question

  • Treating a standardized-suite score as a statement about the deployed system.
  • Dropping your own configured runs because the suite 'already covers it'.
  • Not noticing that the suite's threat model differs from the one you defend against.
  • Assuming a fixed suite stays current without any refresh policy.
  • Comparing your custom run's number against a suite number as if they were the same metric.

context