skip to content

An adversarial library's metric helper — the one that averages perturbation size over the examples an attack flipped — returns 0.0 after a run in which the attack flipped nothing. Why is that zero not evidence the model is trivially fragile?

level: middleimportance: must knowfreq 58%

answer

  1. mean of an empty set
  2. 0.0 is a placeholder, not a measurement
  3. polarity collision: unbreakable prints as fragile
  4. read the success count first
  5. zero successes often means broken wiring

basics

~20 s

Because it is a mean over an empty set. No example was flipped, so nothing entered the average and the helper returns zero as a degenerate default. On this scale small means fragile, so the failure case prints at the alarming end. Check the attack success count before reading the number at all.

solid answer

~50 s

The helper divides a sum of measured perturbation sizes by the number of successes. When the success set is empty there is nothing to average, and the implementation returns a degenerate `0.0` rather than raising or returning a missing value. That is a **collision of meanings on one scale**. On this metric, *small* perturbation means the model is easy to break; so "the attack broke nothing" and "the attack broke everything with a nudge" both print near zero. Nothing in the number itself separates them. The discriminator is the **attack success rate**, which you have to read separately. Zero successes plus 0.0 means the attack, as configured, never crossed a boundary within the search you granted it — a statement about your attack run, not a robustness certificate either. The right next move is to strengthen the attack and rule out broken wiring: wrong input scaling, a dead gradient path, or labels the attack never used.

go deeper

for a junior

Should recognise the number is an average over flipped examples and that nothing was flipped, so there was nothing to average.

for a middle

Explains the empty-denominator default and the polarity collision that makes an unbroken model print at the fragile end of the scale.

for a senior

Treats a zero-success run as suspect first: checks harness wiring, input scaling and gradient flow before reporting anything about the model.

for a principal

Sets the reporting rule that this metric is never published without its success count, and picks a metric with a real denominator for anything a non-specialist reads.

**The mechanism.** The helper computes `sum(measured_distances) / count(successes)` over the rows an attack flipped. When the attack flips nothing, the numerator is an empty sum and the denominator is zero. Implementations guard that division rather than raise: the Adversarial Robustness Toolbox's `art.metrics.empirical_robustness`, for example, hands back `0.0` when the success mask is empty. That `0.0` is a placeholder emitted by a guard clause. It is not a measurement, and no quantity about the model was computed to produce it. **Why this particular failure bites so hard.** The metric's polarity is "a bigger number means the attack had to work harder", so a zero renders at the *fragile* end of the scale. Two opposite situations therefore print the same value: "the attack broke everything with an imperceptible nudge" and "the attack broke nothing at all". A dashboard that plots this float over time will draw your most resistant run as your worst one, and a reader who has never opened the helper's source has no way to tell which of the two they are looking at. Nothing inside the number separates them — only the success count does, and the helper does not return it. **Three causes, in the order an operator should rule them out.** 1. **A genuinely unsuccessful attack.** The perturbation budget, step count or number of restarts you granted was too small to cross a boundary anywhere in the batch. That is a statement about *your run*, not about the model. The standard response is to escalate strength until you either obtain successes or reach a strength you can defend in writing as adequate for the threat model you were hired to test. 2. **A broken harness.** The wrapper feeds inputs in a different scaling, channel order or normalisation than the model was trained on; gradients come back identically zero because a preprocessing step is not differentiable or `requires_grad` was never set; the attack optimises against label indices the model does not emit. A *perfectly* clean zero-success run is far more often a wiring fault than an unbreakable classifier, and experienced people treat it that way by reflex. 3. **A defence that hides the signal rather than removing it.** A model whose preprocessing masks or obscures gradients can stall a gradient-based search while remaining perfectly breakable by a decision-based one. Zero successes under one attack family is silent about every other family. **What the run cost you.** By the time the `0.0` appears you have already paid the entire bill — every forward and backward pass, every metered query, every GPU-hour of the sweep — for a value that carries no information. Escalating strength is not free either: doubling the step count roughly doubles the compute, and multi-restart searches multiply it again, so a "just run it harder" reflex can turn a twenty-minute job into an overnight one. That is why the diagnostic order above matters commercially, not just intellectually. The cheapest checks come first: a single forward pass to confirm the wrapper's predictions match what the served model returns, one gradient call to confirm the tensor is not all zeros, and a control run of the same attack against a deliberately undefended reference model. Those cost minutes. If the same attack cannot break the undefended control either, the fault is in your harness and no amount of extra budget will fix it. **Where the number misleads and what I therefore report.** The seductive misreading is the optimistic one: "we ran the toolkit, nothing broke, the model resisted". That sentence upgrades an absence of evidence into a robustness claim, and it is exactly the sentence a zero-success run tempts you to write. The pessimistic misreading is the mirror image: a stakeholder sees `0.0` on a scale where small means fragile and concludes the model collapses under any nudge. Both readings come from quoting a conditional mean without its condition. So I never publish the bare float. I report the number of rows evaluated, the number flipped (zero), the attack family with its budget and search settings, the outcome of the harness sanity checks, and an explicit sentence that the metric is undefined for this run. If a stakeholder needs one number, robust accuracy at a stated perturbation budget has a real denominator — every evaluated row — and does not degenerate when the attack fails.

  • What single companion number makes the zero unambiguous?
    The count of examples the attack flipped. Zero successes tells you the metric is undefined; many successes with tiny distances tells you the model really is fragile.
  • A run flips zero examples. What do you check before claiming the model resisted the attack?
    The harness: input scaling and preprocessing match what the model serves, gradients are non-zero, the labels the attack targets are ones the model actually emits. A clean zero is usually a plumbing bug.
  • Would returning a missing value instead of 0.0 be better?
    For readability yes, since it forces the caller to handle the undefined case, but you cannot rely on it — assume the degenerate zero and always carry the success count.

saying these in an interview costs you the question

  • Reading 0.0 as 'the model is maximally fragile'.
  • Reading 0.0 as proof the model is robust, with no check of the harness.
  • Not knowing the metric is undefined when no example was flipped.
  • Plotting the number over time with no success count beside it.

context