skip to content

Your organisation publishes robustness numbers for many models, and every increase in an attack's iteration and restart arguments costs GPU hours you have to budget. How would you set a house minimum attack-strength standard that a run must meet before its number is allowed to be published?

level: principalimportance: should knowfreq 25%

answer

  1. two tiers: screening vs publishable
  2. floor = versioned config artefact
  3. attach plateau sweep + forced-success run
  4. cut sample before cutting effort
  5. stamp every number with the config version

basics

~20 s

Define two tiers. A screening tier may use cheap settings, and its results may only report hits, never robustness. A publishable tier requires a versioned configuration — attacks, effort, threat model — plus evidence the success rate plateaued and a forced-success sanity artefact. Same configuration for every model, so numbers stay comparable.

solid answer

~60 s

Treat attack strength as a governed artefact rather than a per-engineer choice. - **Two tiers.** Screening runs are cheap and may only produce positive findings; a screening run that finds nothing is labelled inconclusive and cannot appear as a robustness claim. Publishable runs must meet the floor. - **The floor is a versioned config**, not a number in someone's head: attack families, effort arguments, the threat model's norm, bound and constraints, and the sample. Every model is evaluated against the same version so results are rankable. - **Evidence requirements.** A publishable run attaches an effort sweep showing the success rate stopped moving, and a forced-success run showing the harness fires at all. - **Spending rule.** When budget binds, cut sample size before cutting effort, and say so in the standard — a precise estimate of a weak attack is worse than a coarse estimate of a strong one. - **Refresh cadence.** Bump the config version when new attacks land, and re-baseline; record which config version produced every published number.

go deeper

for a junior

Knows that published numbers should state the settings used, and that cheap runs are not proof of robustness.

for a middle

Can propose a fixed configuration reused across models and explain why comparability requires it.

for a senior

Adds the evidence requirements — plateau sweep and forced-success artefact — and the rule to cut sample size before effort.

for a principal

Owns the whole regime: tiers, a versioned config artefact with an owner, separation of the effort choice from the model team, funded re-baselines, and a review cadence.

### The organisational failure being prevented Attack strength is not a tuning problem for one run; at portfolio scale it is a comparability and integrity problem. Left to individual engineers, each team picks the effort level that fits its deadline — and the resulting numbers get put in one table and ranked as if they measured the same thing. Worse, the effort argument is the single easiest lever for making a model look robust, and the team that owns the model is the team with the incentive to pull it. The fix is to move the effort decision out of the run script and into a governed artefact. ### Structure the standard around what a number is allowed to claim Two named tiers do most of the work. - **Screening tier.** Cheap settings, close to library defaults, run often — on pull requests, on nightly regressions, on a triage sweep across a model zoo. A screening run may report **hits only**. A screening run that finds nothing is labelled *inconclusive* and may never appear as a robustness claim anywhere. - **Publishable tier.** Must meet the strength floor, and only a run at the floor is permitted to say "we searched hard and found little". Writing it as two tiers rather than one rule is what stops the most common report defect — a cheap zero quoted three months later as evidence of robustness — without banning cheap runs, which are genuinely valuable as regression gates. ### The floor is a versioned configuration, not a number in someone's head The artefact should name: the attack families to run and their library and release; the effort arguments for each (iterations, step size, restarts) or a declaration that a fixed external protocol supplies them; the threat model — norm, epsilon, and the constraints on what an attacker may actually change; the sample construction; and the model-wrapper contract, including the input range and where preprocessing lives. Version it, keep it in the repository under a named owner, and stamp every published number with the version that produced it. That single field is what makes cross-model and cross-quarter comparison legitimate, and it is what lets you re-run history when the floor rises. ### Require evidence, not compliance Two attachments carry the weight: 1. An **effort sweep** on a subset, showing the attack-success rate stopped moving at the floor's configuration. Without it, "we used the floor" is a claim that the floor was enough, and nobody checked. 2. A **forced-success run** — the attack unconstrained on a handful of examples, breaking the model — showing the instrument fires at all. This is the evidence that separates "the model held" from "the harness never worked". ### What the regime costs Make this explicit in the standard rather than discovering it in a budget meeting. Track GPU-hours and wall-clock per publishable run and multiply by the model count and the cadence: a floor that takes a few GPU-hours per model, across forty models, quarterly, is a few hundred GPU-hours a year plus engineer time to triage. Publish the trade for when budget binds — **cut sample size before cutting effort**, because a precise estimate of a weak attack is worse than a coarse estimate of a strong one — and carve out separate capacity for the periodic re-baseline, because raising the floor retroactively invalidates published numbers unless re-runs are funded. ### Where the standard's own numbers mislead Two failure modes to design against. First, a floor expressed as prose ("run it properly", "use a strong attack") is unenforceable and drifts back into per-engineer judgement within a quarter. Second, and subtler: a fixed floor becomes a target. Teams optimise against exactly the attacks the floor names, and the published numbers stay high while real robustness does not move. That is why the floor needs a refresh cadence tied to the field moving — new attacks landing — rather than to the calendar of the team that owns it, and why an occasional off-floor probing run by a party with no stake in the result is worth its cost. ### Governance details that matter in practice Name an owner for the configuration artefact. Forbid the model's own team from choosing the effort level behind its published number. Audit by re-deriving one published figure per quarter from its stored configuration version — if it cannot be reproduced from the artefact alone, the artefact is incomplete. And keep the tier label attached to the number for its whole life, because the defect this standard exists to prevent is a screening zero that quietly loses its label on the way into a slide.

  • Why should the team that owns the model not choose the effort level for its own published number?
    Because the effort argument is the easiest lever to make a model look robust, and the incentive runs one way. Separating the choice into a shared, versioned config removes the lever entirely.
  • You raise the floor. What happens to the numbers already published under the old one?
    They stay valid only against the old config version, which is why every number carries that version. Plan and fund a re-baseline; anything not re-run is reported against the version it used, not the current floor.

saying these in an interview costs you the question

  • Letting each model's own team pick the effort level behind its published number.
  • A floor expressed as a slogan ('run it properly') rather than a versioned configuration.
  • Raising the floor without funding the re-baseline of existing numbers.
  • Allowing an inconclusive screening result to be quoted as a robustness claim.
  • Preserving sample size at the cost of search strength when budget is tight.

context