A fraud model's defence cuts a weights-holding attacker's success from 90% to 2% at an unchanged perturbation budget - what has that measured?
answer
- ask what the number is about
- one attacker, one search procedure
- the map, not the territory
- the misclassified rows are still there
- masked, not fixed
basics
~20 sIt measured that one search stopped finding bad inputs, not that they stopped existing. A defence can make the signal an attacker steers by unusable while the same misclassified transactions remain reachable by a different search.
solid answer
~40 sThe number is a property of the attack that was run, not of the model. A gradient-based attacker picks a direction by reading how the loss changes with respect to the input; a defence that makes that reading noisy, discontinuous or numerically dead breaks the *search* while leaving the misclassified inputs exactly where they were. The literature calls this family obfuscated or masked gradients, and the honest reading of a 2% figure is that this attacker, with these steps and this radius, failed. Before treating it as robustness you re-run adaptively: give the attacker more steps and restarts, run a weaker adversary that never touches gradients at all, and run a control with the perturbation budget effectively removed. If the corrected failure rate comes back, the defence bought search cost, not robustness.
go deeper
Be ready to say plainly that an attack-success number describes the attack that was run. Recall that a defence can break the attacker's search while the wrongly classified inputs stay exactly where they were.
An interviewer expects you to name how the search breaks - a scrambled, random or numerically dead input gradient - and to list the extra runs that separate a hard model from a broken optimiser.
Show that you would not sign off the figure without adaptive re-runs, and that you can state what the corrected number is and what perturbation set it is scoped to.
Own the reporting standard: decide what an evaluation must include before a robustness claim may appear in a deck, and be willing to say the defence bought query cost rather than safety.
## What the number is a property of A robust-accuracy or attack-success figure is always the result of a specific procedure: this adversary, with this access, this many optimisation steps, this many restarts, inside this perturbation set. It describes what that procedure found. Turning it into a statement about the model requires the extra assumption that the procedure would have found the bad inputs if they existed - and that assumption is exactly what a masked gradient breaks. ## How a gradient-based attacker actually searches Against a card-not-present fraud model that scores a transaction as approve, decline or step-up, an adversary who holds the weights does not add noise. They read the gradient of the model's loss **with respect to the input features** - a direction in feature space along which the score moves fastest - and take small steps along it, re-projecting back into whatever perturbation set they are allowed after each step. That is why a tiny structured change flips a confident model when random noise of the same size does essentially nothing: the change is a direction, not a magnitude. Everything in that loop depends on the direction being informative. It is a local, first-order reading of a surface, and it is useful only if the surface near the current point actually points toward the region where the model is wrong. ## Three ways a defence kills the direction without removing the region - **Shattered**: a non-differentiable or wildly discontinuous step sits in front of the model, so the local reading no longer describes the function the attacker is optimising. - **Stochastic**: the model computes a random function - a randomised rounding or quantisation of the feature vector, say - so a single reading is one draw, and the difference the attacker measures between two nearby candidates is mostly the randomness. - **Vanishing or exploding through the defence**: the composed function's input gradient is numerically dead or absurdly large, so steps go nowhere or overshoot. In all three cases the set of inputs the model classifies wrongly is unchanged. Only the map the attacker was navigating by has been scrambled. This is the whole content of the phrase *masked, not fixed*. ## Why the drop looks so convincing The drop is real and it is large, because the standard evaluation runs exactly one attack: the strongest known white-box one, at a fixed step count. When that attack's directions become useless, its success collapses toward zero, and a table row reading 2% is genuinely what the harness printed. Nothing in the number itself distinguishes *the inputs are hard to find* from *my search broke*. The distinction has to be established by additional runs. ## What separates the two readings Three checks decide it, and each is a comparison rather than a single number: 1. **Give the same attacker far more budget.** More steps and more random restarts. A genuinely harder problem degrades gradually; a broken search stays flat no matter how much you spend. 2. **Run a strictly weaker adversary.** One that never reads a gradient at all - a search that only sees the returned decision. It cannot outperform an adversary holding weights against a real defence, because the weights-holder can always do whatever the weaker one does. If it outperforms, the reported white-box number was an artefact. 3. **Remove the perturbation budget as a control.** With unlimited freedom to change the input, an attacker can simply hand the model a transaction that genuinely belongs to the other class, so success should approach total. If an effectively unbounded attack still leaves a healthy-looking robust accuracy, the search never worked at any radius. ## What you can honestly claim After the checks, one of two statements is defensible. Either the corrected failure rate is close to the original 2% under every stronger and weaker adversary you ran, in which case you report robustness *against that stated perturbation set at that radius* and nothing wider. Or the corrected rate is far higher, in which case the defence's contribution is that it made the attack more expensive - which is a real but much smaller thing, and one an adversary pays once. The failure mode to avoid is publishing the 2% with no adaptive re-run, because that number is not wrong so much as it is about the wrong object: it describes an optimiser, and readers will take it as describing a model.
- What single extra run would you want before you accepted that 2% figure?The same attacker with far more steps and several random restarts, at the same radius. A model that is genuinely harder to attack degrades smoothly as you spend more optimisation budget. A masked gradient produces a flat line: ten times the compute buys almost nothing, which is the signature of a search that never had a usable direction rather than a region that is hard to reach.
- Does the same caution apply to a reported number from an attack that only saw the model's decisions?Yes, but for a different reason. A decision-only search is weaker, so its failure is much weaker evidence - it says little about what someone holding the weights could do. The asymmetry matters: a low white-box number can be an artefact of masking, while a low black-box number is simply a weak test. Neither is robustness on its own.
- Why does random noise of the same size almost never flip the model, if a perturbation of that size does?Because the perturbation is a direction, not a magnitude. It is chosen from the gradient of the loss with respect to the input, so it moves the transaction along the one axis where the score changes fastest. Random noise of equal size spreads across every direction, almost all of which barely move the decision. That contrast is why the attack is a search problem at all.
Painting over the trail markers does not move the cliff. The next hiker with a different map walks straight to it.
saying these in an interview costs you the question
- Reads a low attack success rate as robustness
- Assumes the defence deleted the misclassified inputs
- Reports one attack's result as a model property
- Thinks a failed gradient proves no bad inputs exist
- Never re-runs with a stronger or weaker adversary