skip to content

Before quoting a mean-perturbation robustness number from an adversarial library, what do you check about the rows that fed it — in particular, examples the model already misclassified before any attack ran?

level: seniorimportance: should knowfreq 38%

answer

  1. two stacked filters: eligible, then flipped
  2. baseline-wrong = free success at zero distance
  3. criterion: changed-from-prediction vs wrong-vs-label
  4. filter to clean-correct, record the count
  5. cluster of near-zero distances = leak or scaling bug

basics

~20 s

Check three things: how many rows were evaluated, how many the attack flipped, and whether rows the model got wrong at baseline were excluded. A baseline-wrong row can register as an instant success with near-zero perturbation, dragging the average down and inflating the success rate — a measurement of plain accuracy, not robustness.

solid answer

~50 s

The metric hands you one float and hides its inputs, so you reconstruct them. **Baseline-wrong rows are the sharp edge.** If a row was already misclassified with no perturbation, an attack that checks "did the prediction differ from the true label" counts it as an immediate success at essentially zero distance. Those zeros enter the mean and pull it toward the fragile end, and they also inflate the success rate. What you would be reporting is the model's ordinary error rate wearing a robustness label. Whether the helper filters them varies by library and by the attack's success criterion — flipped-from-original-prediction and wrong-versus-true-label behave differently here. Verify rather than assume: filter to baseline-correct rows yourself, and record how many that removed. Also record `n_evaluated` and `n_flipped`. A mean over a handful of rows is a rumour, and only the pair tells a reader whether the number means anything.

go deeper

for a junior

Should know to look at how many examples were evaluated and how many were flipped.

for a middle

Explains that already-wrong rows can count as free successes at near-zero distance and so contaminate both the mean and the success rate.

for a senior

Runs the clean-accuracy pass first, filters to correct rows, records what was removed, and reads a cluster of near-zero distances as a bug signal.

for a principal

Standardises the filter and the reported counts across teams so numbers from different engagements describe comparable populations.

**Why this specific check.** The metric's population is defined by two filters stacked on top of each other: which rows were *eligible* to be attacked, and which of those the attack then *flipped*. The helper reports a single float computed after both filters have already run, and returns neither count. Everything interesting about the number lives in the filters, so reconstructing them is the whole job. **Baseline-misclassified rows are the sharp edge.** Two success criteria are common in these libraries, and they diverge exactly here: - *Prediction changed from the model's own original output.* A row the model already got wrong can still require a real perturbation before its prediction moves, so it does not automatically become a zero-distance success. But the "attack" that succeeds on it has moved the model from one wrong answer to a different wrong answer, which is not a security event anyone should report as one. - *Prediction differs from the ground-truth label.* A baseline-wrong row satisfies this criterion **before the attack does anything**. The measured distance is essentially zero and the row is counted as a success. Under the second criterion a model with 30% clean error starts every run with 30% of the population pre-flagged as broken at zero cost. The success rate reads catastrophic and the mean perturbation collapses toward zero — the fragile end of the scale — and neither number has measured any adversarial property whatsoever. What you would be publishing is the model's ordinary error rate wearing a robustness label. Whether a given helper filters those rows for you varies by library, by attack class and by which criterion the attack was constructed with, so verify it in the source or by experiment rather than assuming. **The clean protocol.** 1. Score the clean, unperturbed set first and keep only the rows the model gets right. Record how many rows that removed; the count is itself worth reporting, because it has changed the population you are about to compare against something else. 2. Run the attack on the filtered set only. 3. Report `n_evaluated`, `n_clean_correct`, `n_flipped`, the success rate, the attack family with its perturbation budget, and only then the mean distance. 4. Sanity-check the success set for zero or near-zero distances. A cluster of them almost always means baseline-wrong rows leaked through, or that the wrapper is feeding inputs in a scaling or channel order the model was not trained on — so a nominally tiny change is a large real one. **What the check costs, and why there is no excuse for skipping it.** The clean pass is a single forward pass over the evaluation set with no gradients: seconds to a couple of minutes, against an attack run that is orders of magnitude more expensive — fifty-plus forward *and* backward passes per row for an iterative gradient attack, or thousands of metered queries per row for a black-box one. Skipping the filter saves nothing measurable and corrupts every downstream number, which makes it one of the cheapest correctness checks available anywhere in this workflow. The genuine cost is elsewhere: filtering makes two models' populations differ, since each keeps its own correct rows, and that is a real tension with cross-model comparison. The resolution is to fix one frozen example set and report both filtered and unfiltered counts, so a reader can see exactly what was excluded and why. **Where the number misleads if you skip it.** The contaminated run does not look broken. It produces a plausible-looking float on a familiar scale, a success rate that sounds alarming enough to be believed, and a report that reads as diligent security work. The failure is invisible in the output and visible only in the counts — which is precisely why the counts must travel with the number. The second-order trap is a comparison: an accurate model and a mediocre one evaluated the same unfiltered way will differ mostly by their clean error rates, so the "robustness" gap you write up is a rediscovery of the accuracy gap, and a decision made on it will be a decision about accuracy that nobody knew they were making. **What I would check before believing anyone else's number.** Which success criterion the attack used. Whether a baseline-correct filter ran and how many rows it dropped. The evaluated and flipped counts. The distribution — not just the mean — of the measured distances, looking specifically for a spike at zero. And whether the wrapper's preprocessing matches what the deployed model actually serves, because that single mismatch reproduces the same zero-distance cluster from an entirely different cause.

  • How does the success criterion change whether baseline-wrong rows matter?
    Under 'differs from the true label' they count as free successes at zero distance. Under 'changed from the model's own original prediction' they do not, but a flip between two wrong answers is still not a reportable finding.
  • You see a cluster of near-zero distances in the success set. What are the two likeliest causes?
    Baseline-misclassified rows leaked into the evaluation, or the wrapper feeds inputs in a scaling or channel order the model was not trained on, so tiny nominal changes are large real ones.

It is like counting the doors you found already standing open as doors you picked: your success rate soars and your average picking time collapses, and neither number is about lockpicking any more.

saying these in an interview costs you the question

  • Assuming the library filters baseline-misclassified rows for you.
  • Reporting a mean distance with no evaluated and flipped counts.
  • Not knowing that the attack's success criterion determines whether already-wrong rows count as successes.
  • Explaining away a cluster of near-zero perturbations as 'the model is just fragile'.

context