skip to content

Masked, Not Fixed

Defenses that make the attacker's search fail rather than remove the weakness - unusable gradients, a bolted-on detector, a scrubbing step. Interviewers use them to see if you evaluate adaptively.

on this pageshow

explore

questions

12

A fraud model's defence cuts a weights-holding attacker's success from 90% to 2% at an unchanged perturbation budget - what has that measured?

level: juniorimportance: must knowfreq 62%

answer

  1. ask what the number is about
  2. one attacker, one search procedure
  3. the map, not the territory
  4. the misclassified rows are still there
  5. masked, not fixed

basics

~20 s

It measured that one search stopped finding bad inputs, not that they stopped existing. A defence can make the signal an attacker steers by unusable while the same misclassified transactions remain reachable by a different search.

solid answer

~40 s

The number is a property of the attack that was run, not of the model. A gradient-based attacker picks a direction by reading how the loss changes with respect to the input; a defence that makes that reading noisy, discontinuous or numerically dead breaks the *search* while leaving the misclassified inputs exactly where they were. The literature calls this family obfuscated or masked gradients, and the honest reading of a 2% figure is that this attacker, with these steps and this radius, failed. Before treating it as robustness you re-run adaptively: give the attacker more steps and restarts, run a weaker adversary that never touches gradients at all, and run a control with the perturbation budget effectively removed. If the corrected failure rate comes back, the defence bought search cost, not robustness.

go deeper

for a junior

Be ready to say plainly that an attack-success number describes the attack that was run. Recall that a defence can break the attacker's search while the wrongly classified inputs stay exactly where they were.

for a middle

An interviewer expects you to name how the search breaks - a scrambled, random or numerically dead input gradient - and to list the extra runs that separate a hard model from a broken optimiser.

for a senior

Show that you would not sign off the figure without adaptive re-runs, and that you can state what the corrected number is and what perturbation set it is scoped to.

for a principal

Own the reporting standard: decide what an evaluation must include before a robustness claim may appear in a deck, and be willing to say the defence bought query cost rather than safety.

## What the number is a property of A robust-accuracy or attack-success figure is always the result of a specific procedure: this adversary, with this access, this many optimisation steps, this many restarts, inside this perturbation set. It describes what that procedure found. Turning it into a statement about the model requires the extra assumption that the procedure would have found the bad inputs if they existed - and that assumption is exactly what a masked gradient breaks. ## How a gradient-based attacker actually searches Against a card-not-present fraud model that scores a transaction as approve, decline or step-up, an adversary who holds the weights does not add noise. They read the gradient of the model's loss **with respect to the input features** - a direction in feature space along which the score moves fastest - and take small steps along it, re-projecting back into whatever perturbation set they are allowed after each step. That is why a tiny structured change flips a confident model when random noise of the same size does essentially nothing: the change is a direction, not a magnitude. Everything in that loop depends on the direction being informative. It is a local, first-order reading of a surface, and it is useful only if the surface near the current point actually points toward the region where the model is wrong. ## Three ways a defence kills the direction without removing the region - **Shattered**: a non-differentiable or wildly discontinuous step sits in front of the model, so the local reading no longer describes the function the attacker is optimising. - **Stochastic**: the model computes a random function - a randomised rounding or quantisation of the feature vector, say - so a single reading is one draw, and the difference the attacker measures between two nearby candidates is mostly the randomness. - **Vanishing or exploding through the defence**: the composed function's input gradient is numerically dead or absurdly large, so steps go nowhere or overshoot. In all three cases the set of inputs the model classifies wrongly is unchanged. Only the map the attacker was navigating by has been scrambled. This is the whole content of the phrase *masked, not fixed*. ## Why the drop looks so convincing The drop is real and it is large, because the standard evaluation runs exactly one attack: the strongest known white-box one, at a fixed step count. When that attack's directions become useless, its success collapses toward zero, and a table row reading 2% is genuinely what the harness printed. Nothing in the number itself distinguishes *the inputs are hard to find* from *my search broke*. The distinction has to be established by additional runs. ## What separates the two readings Three checks decide it, and each is a comparison rather than a single number: 1. **Give the same attacker far more budget.** More steps and more random restarts. A genuinely harder problem degrades gradually; a broken search stays flat no matter how much you spend. 2. **Run a strictly weaker adversary.** One that never reads a gradient at all - a search that only sees the returned decision. It cannot outperform an adversary holding weights against a real defence, because the weights-holder can always do whatever the weaker one does. If it outperforms, the reported white-box number was an artefact. 3. **Remove the perturbation budget as a control.** With unlimited freedom to change the input, an attacker can simply hand the model a transaction that genuinely belongs to the other class, so success should approach total. If an effectively unbounded attack still leaves a healthy-looking robust accuracy, the search never worked at any radius. ## What you can honestly claim After the checks, one of two statements is defensible. Either the corrected failure rate is close to the original 2% under every stronger and weaker adversary you ran, in which case you report robustness *against that stated perturbation set at that radius* and nothing wider. Or the corrected rate is far higher, in which case the defence's contribution is that it made the attack more expensive - which is a real but much smaller thing, and one an adversary pays once. The failure mode to avoid is publishing the 2% with no adaptive re-run, because that number is not wrong so much as it is about the wrong object: it describes an optimiser, and readers will take it as describing a model.

  • What single extra run would you want before you accepted that 2% figure?
    The same attacker with far more steps and several random restarts, at the same radius. A model that is genuinely harder to attack degrades smoothly as you spend more optimisation budget. A masked gradient produces a flat line: ten times the compute buys almost nothing, which is the signature of a search that never had a usable direction rather than a region that is hard to reach.
  • Does the same caution apply to a reported number from an attack that only saw the model's decisions?
    Yes, but for a different reason. A decision-only search is weaker, so its failure is much weaker evidence - it says little about what someone holding the weights could do. The asymmetry matters: a low white-box number can be an artefact of masking, while a low black-box number is simply a weak test. Neither is robustness on its own.
  • Why does random noise of the same size almost never flip the model, if a perturbation of that size does?
    Because the perturbation is a direction, not a magnitude. It is chosen from the gradient of the loss with respect to the input, so it moves the transaction along the one axis where the score changes fastest. Random noise of equal size spreads across every direction, almost all of which barely move the decision. That contrast is why the attack is a search problem at all.

Painting over the trail markers does not move the cliff. The next hiker with a different map walks straight to it.

saying these in an interview costs you the question

  • Reads a low attack success rate as robustness
  • Assumes the defence deleted the misclassified inputs
  • Reports one attack's result as a model property
  • Thinks a failed gradient proves no bad inputs exist
  • Never re-runs with a stronger or weaker adversary

context

open as a page

A vision pipeline denoises and re-encodes every submitted image before the classifier sees it - which attacker does that stop, and which does it merely charge?

level: juniorimportance: must knowfreq 58%

basics

~20 s

It stops an attacker whose perturbation was fitted against the bare classifier and never had to survive the cleaning step. An attacker who knows the step is there fits a perturbation through it and pays only extra optimisation steps.

open as a page

An inbound-mail classifier sits behind an adversarial-input detector - why isn't the problem gone?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The detector is another trained model with its own decision boundary and its own mistakes. An attacker aware of it crafts a single message that reads as ordinary mail to the detector and still fools the classifier behind it.

open as a page

Why can one message within a fixed edit budget satisfy both an inbound-mail classifier and its detector?

level: middleimportance: must knowfreq 54%

basics

~20 s

The attacker searches one allowed set of edited messages for a point on the wrong side of the classifier's boundary and the ordinary side of the detector's. Two conditions in one budget means more search, not a wall.

open as a page

A decision-only random search beats a weights-holding attacker on the same fraud model at equal compute - what does that inversion mean?

level: middleimportance: should knowfreq 48%

basics

~20 s

It means the white-box result is an artefact of a broken search, not a robustness property. An adversary holding weights can always imitate a weaker one, so a strictly weaker attacker outscoring them is logically impossible unless the stronger attack's gradient signal is unusable.

open as a page

How does an attacker optimise through a cleaning stage that has no usable gradient, and what does that cost them?

level: middleimportance: should knowfreq 44%

basics

~20 s

They replace the non-differentiable stage with a smooth stand-in when computing their search direction, and the natural stand-in is the identity because a content-preserving transform barely changes its input. The cost is extra optimisation steps, offline and cheap.

open as a page

An attack with the perturbation budget removed still leaves a defended fraud model at 15% robust accuracy - what do you conclude?

level: seniorimportance: should knowfreq 44%

basics

~20 s

That the search is broken at every radius, not that the model is robust. With no limit on how far an input may move, an attacker can hand the model a transaction that genuinely belongs to the other class, so success should approach total; anything left standing is the optimiser failing.

open as a page

You turn an inspection line's input-cleaning stage up until an attacker's perturbations stop landing - what did that buy?

level: seniorimportance: should knowfreq 36%

basics

~20 s

A false-reject and re-image rate that lands on your hardest real parts, plus a cost increase for an adaptive attacker. The window between destroying a perturbation and destroying the defect signal is narrow, and the attacker works inside whatever survives.

open as a page

A supplier reports a 95% catch rate for an adversarial-input detector - what do you ask before believing it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Ask which attack produced it. A catch rate from attacks that did not know the detector existed shows those attacks were caught, not that the detector holds. Ask for the detector-aware number and the false-alarm rate.

open as a page

A fraud model randomly quantises features before scoring, so a resubmitted transaction scores differently - why does that stall an attacker?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Because each reading is one draw from a distribution, not the quantity being optimised. The difference an attacker measures between two nearby candidates is swamped by the defence's randomness, so the search follows noise instead of a direction - while the misclassified transactions stay reachable.

open as a page

You inherit a pipeline documented as 91% accurate under attack with its cleaning stage enabled - what do you ask before trusting that?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Ask whether the cleaning stage was inside the attacker's optimisation, and what norm, radius, steps and restarts the attack used. If every row was measured against an attacker who ignored the stage, the figure describes that attacker, not the pipeline.

open as a page

Your mail detector is an off-the-shelf model anyone can download - what does that do to its value?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

It removes most of the cost the detector was meant to add. Anyone can solve that half offline for free, so only the classifier still charges the attacker, and one solve serves every deployment that bought it.

open as a page