skip to content

Your organisation wants to gate model releases on one robustness number from an adversarial library — a mean perturbation size over only the examples an attack flipped. What breaks when that number is tracked release over release, and what would you gate on instead?

level: principalimportance: nice to knowfreq 26%

answer

  1. conditional mean is non-monotone
  2. zero-success degenerate case fails your best release
  3. gate becomes an optimisation target
  4. pin set, attack, budget, library version
  5. gate on a rate with a real denominator

basics

~20 s

It is a conditional mean whose population changes every release, so it is not monotone in robustness and moves for reasons unrelated to the model. It also collapses to zero when nothing is flipped. Gate instead on attack success rate at a pinned perturbation budget over a frozen example set, with row counts recorded.

solid answer

~50 s

**What breaks.** The metric averages only over the rows an attack flipped, so its population is redefined every release. A model that improves may lose only borderline examples, whose small distances *lower* the mean; a model that regresses may lose far-from-boundary examples whose large distances *raise* it. The series moves opposite to the property you are gating on. It also degenerates to zero when a release flips nothing, printing a pass as the worst score in the history. And a gate is an optimisation target: a weaker attack configuration produces a friendlier number, and nothing in the metric objects. **What I gate on.** Attack success rate at a **pinned perturbation budget**, over a **frozen** evaluation set filtered to rows the model classifies correctly, under a pinned attack configuration and library version, with the evaluated and flipped counts beside the rate. Same denominator every release, monotone direction, and a change means something. The conditional mean stays in the report as colour, never as the gate.

go deeper

for a junior

Should recognise the number depends on which examples the attack flipped and so is not stable across releases.

for a middle

Explains non-monotonicity and the zero-success degenerate case, and proposes a success rate at a fixed budget instead.

for a senior

Adds the operational pinning — frozen example set, pinned attack settings and library version, counts carried with every number — and re-baselining on change.

for a principal

Owns the policy: what the gate is, who may change its inputs, what a green gate does and does not claim, and how to resist tuning the instrument to pass.

**The three failures, in the order they will actually bite.** 1. **Non-monotonicity.** A gate must move in one direction with the property it protects; otherwise crossing it means nothing. This number does not. Its population is the set of rows the attack flipped, which is selected by each release's own weaknesses, so an improvement and a regression can both push it up or down. A model that gets harder to break may lose only borderline points, whose small distances *lower* the mean; a model that regresses may start losing far-from-boundary points, whose large distances *raise* it. A gate that can be tripped by improving the model, and passed by degrading it, will be argued around within two release cycles and then quietly ignored. 2. **The degenerate case.** When a release flips nothing, the helper divides by an empty success set and returns `0.0` — a guard-clause placeholder, not a measurement. On a scale where small means fragile, your best release scores worst in the history. Someone then adds a special case for it, and the special case becomes the piece of gate logic nobody remembers is there. 3. **Goodhart pressure.** The moment a number blocks a release it becomes an optimisation target, and the model is the hardest of its inputs to change. The attack configuration, the example set, the baseline filter and the library version are all levers that move the number without touching the model at all. If they are not pinned and reviewed as code, the gate is decorative — and it will be tuned, not maliciously, but by an engineer under deadline who genuinely believes the attack settings were too aggressive. **What a defensible gate looks like.** - **Frozen inputs.** One versioned example set; the attack family, its search settings and its perturbation budget pinned in configuration; the baseline-correct filter pinned; the library version pinned. Changing any of them is a reviewed change with a re-baseline, never a run-time flag. - **A rate, not a conditional mean.** Attack success rate at the pinned budget, with the denominator being every eligible row. Add a second, larger budget if you care about a stronger threat model — two rates read fine on one slide and bound the model between two stated adversaries. - **Counts always carried.** `n_evaluated`, `n_eligible_after_baseline_filter`, `n_flipped`. Without them nobody can distinguish a real improvement from a shrunken evaluation set, which is the single easiest way to fake a passing gate by accident. - **Direction decided in advance.** State the threshold and the consequence of crossing it before the release, so the gate is policy rather than an argument reopened every time it fails. **What it costs to run this way, and why that shapes the design.** Every release now pays for a full attack sweep over the frozen set: for an iterative gradient attack that is tens of forward and backward passes per row, so a few thousand rows is minutes to hours of accelerator time per release — schedulable, but real, and it is the reason people are tempted to shrink the set or weaken the attack. Every pin change costs a re-baseline: you must re-run the *previous* release under the new configuration too, or the series has a discontinuity nobody can interpret. Budget for that explicitly, because an un-budgeted re-baseline is how a pin quietly stops being a pin. And if the gate targets a hosted, metered endpoint rather than a local model, the run has a per-release invoice attached, which changes the argument from engineering to procurement. **Where the gate still misleads even when it is built correctly.** It measures one attack family, at one budget, on one frozen set, against a baseline-filtered population. It is an upper bound on that model's resistance to *that specific procedure* — nothing more. A green gate is not a safety statement, it is not evidence about attack families you did not run, and it says nothing about rows outside your frozen set or about drift in production inputs. Because a green result is exactly the artefact stakeholders over-read, I write that limitation into the release note itself, next to the number, rather than into an appendix: this release was not broken more than x% of the time by attack F at budget B over n rows, and that is the entire claim. The failure this whole discipline exists to prevent is a number that outlives the conditions under which it was measured.

  • Why is a rate a better gate than a mean here?
    Its denominator is every eligible row and does not change with the model's weaknesses, so it moves monotonically with the property being gated and a delta between releases is interpretable.
  • What has to be pinned for a release-over-release comparison to mean anything?
    The evaluation set, the attack family and its search settings, the perturbation budget, the baseline-correct filter, and the library version — all as reviewed config, with a re-baseline whenever one changes.
  • A team proposes lowering the attack strength because the gate is failing. What is your answer?
    That changes the instrument, not the model. Attack strength is part of the gate definition and only changes through review with a re-baseline and a note in the release history.

saying these in an interview costs you the question

  • Accepting a conditional mean as a release gate without questioning its denominator.
  • Proposing to weaken the attack when the gate fails.
  • Leaving the evaluation set, attack settings or library version unpinned across releases.
  • Presenting a green gate as evidence the model is safe rather than as one attack at one budget.

context