skip to content

As a reviewer, what does the distribution of smallest perturbations an attacker needed per input tell you that one fixed-radius success rate does not?

level: seniorimportance: should knowfreq 36%

answer

  1. a rate is one cut through a distribution
  2. plateau or cliff, you cannot tell
  3. which classes sit closest
  4. every distance is an upper bound
  5. distances live in the model's input units

basics

~20 s

A fixed-radius success rate is one threshold cut through the distance distribution. The distribution itself shows where your data actually sits, which slices sit closest, and whether the radius that was reported stood on a plateau or on a cliff.

solid answer

~50 s

One attack-success number tells you what fraction of inputs an attacker failed to break at a radius somebody chose - and nothing about how the rest of the data is arranged. The per-input distances give you the shape. You can see whether the mass sits far from any boundary or bunched just past the reported radius, which classes or traffic slices sit closest, and whether a small move in the radius would collapse the number. That last point is what makes the radius defensible or not: a figure quoted at a radius chosen after the results were in is a number selected to look good. Two cautions when reading it. Each distance is an upper bound - a stronger search only ever finds a smaller one - so the distribution shifts down under a better attacker and never up. And distances are in the units the model is fed, so they are not comparable across models with different preprocessing or input scaling.

code

text · 9 lines
text
system: call-intent classifier   access: extracted weight file

constraint on the change      attack success
perturbation at SNR >= 30 dB       46%
perturbation at SNR >= 20 dB       89%
...
median smallest change needed per clip : not reported
per-intent-class breakdown             : not reported
clips flipped below the audibility floor: not reported

go deeper

for a junior

Know that a success rate at one radius is a single cut, and that the underlying per-input distances carry information the rate throws away.

for a middle

Explain plateau versus cliff and why each per-input distance is an upper bound that a stronger search can only lower.

for a senior

Demonstrate the review reflex: ask which radius, when it was chosen, what the distribution looks like, and where the exposure concentrates before accepting any figure.

for a principal

Set the bar for what an evaluation must publish before it can support a claim, and be willing to send a result back when the radius was chosen after the numbers were seen.

## What one number leaves out An evasion evaluation reported as a single figure - accuracy under attack at one perturbation radius - is a threshold cut. Somewhere behind it is a distribution: for each input, the smallest change an attacker had to make before the model read it differently. The reported figure is just the fraction of that distribution lying above the chosen radius. Everything the cut discards is the interesting part. **Shape near the threshold.** If the mass of the distribution is bunched just above the reported radius, the number is standing on a cliff: nudge the radius and it collapses. If the mass sits far above, the number is on a plateau and is genuinely descriptive. From the single figure you cannot tell these apart, and they mean opposite things about the deployment. **Where the exposure is concentrated.** Aggregate rates hide slices. It is common for one class, one channel, or one population of inputs to sit systematically closer to a boundary than the rest, so an overall figure that looks acceptable is carried by the easy majority while the part of the traffic that matters is the part sitting close. **The tail.** A handful of inputs practically on a boundary is a different operational problem from broad mild fragility, and needs a different answer. The distribution shows the tail; the rate does not. **Whether the radius was chosen honestly.** This is the reviewer's real question. A radius fixed by convention before any results existed is one thing. A radius arrived at after seeing which value made the model look best is another, and a reader cannot distinguish them from the number alone. Publishing the distances is what makes the choice auditable, because the reader can see for themselves whether that radius was a fair place to cut. ## Two things the distribution still is not **It is an upper bound, everywhere.** Each per-input distance is the smallest change *some search found*, not the smallest that exists. A better-resourced adversary, a different formulation or more restarts can only move a given input's number down. So the whole distribution shifts left under a stronger attacker and never right, and no part of it is a guarantee. A statement in the other direction - that nothing smaller than some radius can flip this input - is a certificate, a different object with its own confidence and its own cost, and it is not what an attack produces. **It is measured in the model's input units.** A distance is expressed in whatever space the model consumes, so two models with different scaling, normalisation or resampling produce numbers that cannot be put side by side. Comparing them without accounting for that is a category error, and it is a common one in vendor material and internal decks alike. ## Using it as a reviewer When a robustness result lands on your desk, the productive questions are all about what the single number hides: - Which radius, in which norm, and was it fixed before the run or after it? - Is there a distance distribution behind it, even on a sample of a few hundred inputs? - Does the reported figure hold up when the radius moves a little in either direction? - Is the exposure even across the classes and traffic slices you actually care about? - Was the search that produced these numbers the strongest one available, or the most convenient one? A team that can answer these has measured something. A team that can only cite one figure has run one experiment and reported one cut through it. ## What it changes about a defence claim The distribution is also the cleanest way to see whether a hardening effort did anything real. A defence that moves the whole distribution to the right has genuinely bought distance. A defence that moves only the inputs that were sitting immediately below the evaluation radius, leaving the rest of the shape unchanged, has bought a headline number and very little else - and that pattern is invisible if only the rate is reported. The distance distribution turns a claim of the form 'we improved robustness' into something a reviewer can check.

  • A hardened model reports 61% under attack at one radius and 9% at twice that radius. What do you conclude?
    That the reported radius is sitting just below a cliff. Most of the data is bunched only slightly beyond it, so the headline figure is an artefact of where the cut was placed rather than a description of the model. I would ask for the distance distribution and for the reasoning behind that radius, and I would treat the 61% as uninformative until both arrive.
  • Can you use these distances to claim a minimum level of robustness to a regulator or a customer?
    No. Every distance is the smallest change a particular search found, so the honest claim is 'this adversary needed at least this much', which a stronger adversary can falsify tomorrow. Claims that hold in the other direction come from certification, which states a radius with a confidence level and its own costs, and even then it covers only the stated radius and says nothing outside it.
  • Two teams report per-input distances for two models. What must be true before you compare the numbers?
    That both are measured in the same norm, in the same input space, after the same preprocessing and scaling, and produced by searches of comparable strength. Any difference in how inputs are normalised changes the units the distance is expressed in, and a weaker search on one side inflates its distances. Without that, the comparison measures the evaluation setups rather than the models.

saying these in an interview costs you the question

  • Treats the distance distribution as a robustness guarantee
  • Reads one attack-success figure as a property of the model
  • Compares distances across models with different input scaling
  • Never asks when the evaluation radius was chosen
  • Ignores per-class and per-slice concentration behind the average

context