skip to content

Your team has standardised on one adversarial-robustness toolkit and reuses the same three attack classes on every model assessment. What is the case for and against that as a program-level policy, and how would you decide the class set per engagement instead?

level: principalimportance: should knowfreq 30%

answer

  1. comparability vs fit
  2. frozen set = frozen threat model
  3. clean means those three did not apply
  4. derive from access, domain, action space
  5. small core for trend, assumptions in every finding

basics

~20 s

For: comparability across assessments, reviewable tooling, faster onboarding, predictable cost. Against: the three classes encode one fixed access level and data type, so on a differently shaped target a clean result means only that those three did not apply. Decide per engagement from granted access, input domain and attacker action space, keeping a small fixed core purely for trend comparison.

solid answer

~50 s

**For.** A fixed set is a real asset: results are comparable across engagements and over time, the classes and their arguments have been reviewed once, junior staff produce defensible runs immediately, and cost per assessment is predictable. **Against.** Each class carries an access level, a data type and a loss. Freeze three and you have frozen a threat model most targets will not match: on a label-only endpoint the white-box members cannot run; on tabular or text inputs the image-shaped members produce infeasible candidates; and a clean result reads as 'robust' when it means 'those three did not apply here'. **The policy.** Derive selection rather than defaulting: granted access picks the access-level group, the input domain picks the data-type group, and the attacker's action space sets the constraints. Keep a small core for trend lines only, and require every finding to name the assumed access level, data type and budget, so comparability comes from metadata rather than from freezing the method.

go deeper

for a junior

Should recognise that the same three attacks will not fit every target and that a clean result is not a robustness verdict.

for a middle

Names the specific mismatches — access level unavailable, wrong data type — and that class defaults change across releases.

for a senior

Derives the set from granted access, input domain and action space per engagement, and reports coverage with an explicit not-assessed list.

for a principal

Balances comparability against fit as a program policy: a small pinned core for trends, derived selection per engagement, assumptions carried in every finding, and a scheduled review of the default.

### The tension, stated honestly A fixed set of three attack classes is a program decision, not a technical one, and both sides of it are real. The question is only ever answered well by someone who can argue the case they are about to reject. ### What standardising genuinely buys - **Comparability.** Any trend claim — "this service is more robust than last quarter", "team A's models score worse than team B's" — requires a stable method. Ad-hoc per-engagement selection makes every number a one-off, and a portfolio of one-offs cannot be aggregated at all. - **Reviewability.** Three classes with agreed arguments, a fixed perturbation budget and a fixed norm can be reviewed once, versioned, and defended in a regulatory or customer conversation. An open catalogue chosen fresh each time cannot be. - **Throughput and predictable cost.** A newer engineer runs something defensible on day one. Assessment cost is estimable in advance — the same classes at the same budget on a similar model take roughly the same GPU-hours or query volume, which is what makes a fixed-price engagement possible at all. ### What it silently costs - **A frozen threat model.** Each class carries an access level, a data type and a loss. Freeze three and you have frozen one attacker position. Against a label-only endpoint the white-box members simply cannot run; against tabular or text inputs the image-shaped members run and produce infeasible candidates. Coverage ends up decided by the team's history rather than by the system in front of them. - **A misread null result.** "Clean against the standard set" is a statement about three classes under their own assumptions. Clients read it as a robustness verdict, and a report that leads with a percentage helps them do so. - **Selection by familiarity.** Once a set is the default, nobody re-reads what its members assume. The assumption check — the thing that made any of the numbers interpretable — quietly stops happening. - **Silent drift.** Class implementations, default arguments and even the meaning of a parameter change across library releases. An unpinned "fixed" set produces trend breaks that originate in the tooling, not in the target — the worst possible failure for a metric whose entire justification was comparability. ### Where the number misleads, specifically The dangerous artefact is the **coverage percentage**. "We ran 3 of 3 standard attacks; 0 succeeded" has a denominator of *the set*, and the set is not the attack space. Two different denominators get conflated: attacks-run-out-of-attacks-available (a tooling statistic) versus attacker-positions-assessed-out-of-positions-in-scope (the one the client cares about). The first is always near 100% by construction. Publishing it next to a risk rating manufactures confidence out of an arbitrary choice made two years ago by whoever set up the pipeline. ### The policy I would write 1. **Derive selection.** Granted access (weights / scores / labels) selects the access-level group; the input domain selects the data-type group; the attacker's action space sets the constraints and the transformation. Budget decides *how many examples*, never *which class*. 2. **Keep a small core, and say what it is for.** Two or three classes run wherever applicable, purely as a trend line, labelled as such — never as the coverage claim. 3. **Put the assumptions in the finding.** Every number carries access level, data type, artefact tested, perturbation budget and cost spent. That metadata is what survives a change of class set; the class name is not. 4. **Report coverage against a defensible denominator,** and attach an explicit list of attacker positions *not* assessed, with the reason (out of scope, no write path, no budget). 5. **Pin and re-review.** Pin library versions per engagement, record them in the finding, and put the core set on a scheduled review — a default nobody revisits has become an unexamined threat model. ### What tells you it is working Findings start being challenged on their assumptions rather than their percentages; two assessments of differently shaped systems stop being compared as though their numbers meant the same thing; and someone can answer "why these three?" with the current engagement's access grant rather than with history.

  • How do you keep trend lines meaningful once selection is per-engagement?
    Move comparability into metadata rather than method: every finding records access level, data type, budget and artefact, and only the small applicable core is used for cross-engagement trends, with its library versions pinned.
  • What single reporting change most reduces the harm of a default set?
    An explicit coverage statement naming the attacker positions not assessed and why, alongside the classes that were run. It converts an implied verdict into a scoped result.

saying these in an interview costs you the question

  • Defending the fixed set purely on efficiency without naming what coverage it forecloses.
  • Letting a clean result against the default set be reported as 'the model is robust'.
  • Abandoning any standard set, so no two assessments can be compared.
  • Ignoring that class defaults drift across library releases while trend lines assume they did not.

context