An attacker holding a model's weights can work inside a fixed perturbation radius or minimise the change instead - what does each measure?
answer
- one fixes the size, one fixes the outcome
- a rate versus a per-input number
- one gives you a whole distribution
- both are upper bounds, never guarantees
- shrinking costs far more work per input
basics
~20 sA fixed-radius attack answers 'can this input be broken within this much change?' and returns a success rate at a radius somebody chose. A minimising attack answers 'how much change did this input need?' and returns a distance per input.
solid answer
~50 sThey are two different measurements, not two routes to the same one. A fixed-radius attack fixes a norm and a radius up front and tries to break the model inside it; the output is an attack-success rate at that radius, and it only ever bounds the model from one side - a stronger search lowers it. A minimising attack inverts the objective: the flipped label becomes the constraint and the size of the change becomes the thing to shrink, so it returns a per-input number - how close that input already was to being read differently. That set of numbers is a distance distribution over your data, and it is what tells you whether the radius the first attack used was a fair choice or a flattering one. The attacker pays for the second in compute: shrinking costs far more work per input than filling a budget.
go deeper
Be ready to state the difference in one breath: one fixes how much change is allowed and asks whether that breaks the model; the other requires the break and asks how little change was enough.
Explain what each result is a bound on and in which direction, and why a per-input distance set is a distribution while a success rate is a single cut through it.
Show that you would ask which radius was chosen and why before reading anyone's attack-success number, and that you know the minimising result is what makes that radius defensible.
Own the reporting standard: decide what an evaluation must publish so two models are actually comparable, and refuse numbers whose radius was picked after the results were in.
## Two questions, not one measurement An adversary who can differentiate a model - because they pulled the weight file out of an appliance image, say, and run it offline - has two quite different things they can ask of a given input. **Fixed-radius (maximise damage inside a set budget).** Fix a norm and a radius first: every coordinate of the input may move by at most this much, or the total energy of the change is capped at this much. Then search inside that set for the change that hurts the model most. The result for one input is binary - flipped or not - and the result over a test set is a success rate, usually reported the other way round as accuracy under attack. **Minimum perturbation (minimise the change, subject to flipping).** Swap the roles. The flip is now the requirement, and the size of the change is the quantity being pushed down. The result for one input is not a yes or a no; it is a number - the smallest change the search managed to find that still made the model read the input differently. ## Why they are not interchangeable The common wrong answer in an interview is that these are the same experiment with the knob on a different side. They are not, and the giveaway is what each one returns. A fixed-radius run is one threshold cut through the distance distribution. It tells you what fraction of your data sits further from a boundary than the radius somebody picked - and it tells you nothing about how the rest of the distribution is shaped. Two models can report the same success rate at that radius while one has most of its data just barely on the safe side and the other has most of its data far away. The first collapses when the radius moves slightly; the second does not. A minimising run gives you the shape. You can ask where the mass sits, which classes or which slices of traffic sit closest, whether there is a thin tail of inputs that are practically on the boundary already. Crucially, it is what makes the fixed radius defensible: if you cannot say where your data actually sits, then the radius in your evaluation was chosen by convention or by whatever made the number look good. ## What each number is worth Both are empirical results produced by a search, so both are one-sided in the same direction, and this is the direction people get backwards: - A fixed-radius success rate is an **upper bound on robustness**. The attack that was run failed on those inputs. A better-resourced attacker lowers the number; nobody can raise it. - A minimum perturbation is an **upper bound on the true distance**. The search found *a* change of that size; a stronger search can only find a smaller one. It is never a guarantee that nothing closer exists - that direction requires a certificate, which is a different object entirely. So neither is a proof of safety. Both are records of what one adversary managed under one set of assumptions. ## What the attacker pays The formulations also cost differently, and that shapes where you see each one. A fixed-radius attack runs a bounded amount of work and stops; every candidate it produces is admissible by construction, because it stays inside the ball. A minimising attack has to hold two things in tension - keep the label flipped, make the change smaller - and it has to re-confirm the flip every time it shrinks. In practice that is one to two orders of magnitude more work per input. That cost is nearly free to an adversary running an extracted weight file on their own hardware, where the only bill is wall-clock time per item. It is close to prohibitive against an endpoint that meters and charges per call. This is why minimising results are typically reported by someone with the weights in hand, and why an adversary paying per query fills a budget instead. ## The payoff is different too There is also a difference in what the attacker wants. Filling a budget is about producing a flip at all. Minimising is about producing a flip **nobody notices** - an item that clears the model while the human who also looks at, or listens to, the same item registers nothing unusual. In a pipeline where a machine decision is spot-checked by a person, the smallest change is the one that survives the person, and the size of the change is measured against human perception rather than against a convenient constant.
- If a minimising search reports a very small change for an input, does that prove no smaller one exists?No. The search found one change of that size; it is an upper bound on the true distance to a different answer. A better search, more restarts or a different formulation can only push it down. A statement in the other direction - nothing smaller exists - is a certificate, produced by a completely different kind of analysis, and it comes with its own radius, confidence and cost.
- Which formulation would you expect an adversary paying per call to a hosted model to use, and why?The fixed-radius one, or a boundary-walking attack that needs only the returned label. Minimising the change means repeatedly shrinking and re-checking, so the number of model evaluations per input is large. That is cheap for someone running the weights themselves and expensive for someone billed per query, so a metered adversary takes the coarser result rather than the smallest one.
- Why does the smallest change matter more than the strongest one when a human reviews the same item?Because the human is the real detector in that path. A change big enough to be visible or audible gets the item escalated regardless of how the model labelled it, so the attack has failed its actual constraint. When the aim is that nobody looks, the binding limit is human perception, and the attacker is trading a great deal of extra compute for a change that stays under it.
One asks 'can you get through a gap this wide?' and gets a yes or no. The other asks 'how wide does the gap have to be?' and gets a measurement for every door in the building.
saying these in an interview costs you the question
- Says both formulations measure the same thing
- Reads attack-success rate as a guarantee of safety
- Calls a found minimum perturbation the exact distance
- Thinks the two differ only in reporting units
- Assumes minimising costs the attacker the same as filling a budget