A decision-only random search beats a weights-holding attacker on the same fraud model at equal compute - what does that inversion mean?
answer
- access assumptions are nested
- the stronger one can imitate the weaker
- so the ordering cannot invert
- unless the extra information misleads
- take the weaker result as a floor
basics
~20 sIt means the white-box result is an artefact of a broken search, not a robustness property. An adversary holding weights can always imitate a weaker one, so a strictly weaker attacker outscoring them is logically impossible unless the stronger attack's gradient signal is unusable.
solid answer
~40 sThe inversion is a proof by monotonicity. Access is nested: an adversary holding the weights can run any procedure a decision-only adversary can run, so at equal compute their success should never be lower. When it is, the extra information they hold is actively misleading them - the input gradient they steer by no longer points toward the misclassified region, because the defence scrambled, randomised or flattened it. The correct triage is not to file the result as noise or to celebrate the low white-box number, but to treat the decision-only figure as a *lower bound* on what the model really does and re-run the white-box attack adaptively until it clears that bound. The corrected failure rate is the one you report; the earlier number described the optimiser.
code
text · 10 linesrobustness eval - card-not-present fraud model, build 7
threat model: L-infinity, radius 0.03 on normalised features
attack access model evals success
--------------------- ------------------ ----------- --------
iterative, 50 steps weights + gradients 100 2%
random search decision only 5,000 31%
iterative, 50 steps weights + gradients 5,000 3%
...
(seeds: 1 run each; no restart column reported)go deeper
Recall that an attacker holding the weights can always fall back to what a weaker attacker does, so a weaker one scoring higher is a signal something is wrong with the measurement.
Be ready to state the nesting argument explicitly and to check that compute was held equal in the same unit before drawing any conclusion from the two rows.
Demonstrate the triage: seeds and spread, the weaker result as a lower bound, an adaptive re-run, and a reported number carrying its radius and access assumption.
Decide what your organisation is allowed to publish. A defence that raises attack cost is worth naming as cost, and the standard should forbid presenting it as a failure rate.
## The argument that makes the inversion diagnostic Threat models on this tree are access assumptions, and they are **nested**. An adversary who holds the weights, the architecture and the gradients can, if they choose, ignore all of it and run exactly the procedure a decision-only adversary runs - probing the model, keeping what moves the decision, and searching without ever reading a gradient. Nothing forbids it. So at an equal budget of model evaluations, the weights-holder's success rate is bounded below by the decision-only adversary's. That is why the inversion is not merely surprising, it is *impossible* for a working evaluation. When you observe it, one of the two runs is not measuring what its label says. In practice it is almost always the white-box run, and the cause is that the direction it steers by has stopped describing the function. ## What equal compute has to mean The comparison only carries the argument if the budget is genuinely held equal, and this is where the check is usually done badly. A gradient-based attack at 50 steps costs on the order of 50 forward and 50 backward passes; a random search given 5,000 probes costs 5,000 forward passes. Reporting those side by side and concluding *the black-box attack is stronger* would be a budget error, not a finding. The honest comparison holds model evaluations - or wall-clock, or spend against a paid endpoint - constant across the two rows, and states which unit was held constant. ## The three tells, and how they relate This inversion is one of a family of external signatures, all of which say *the search failed* rather than *the model held*: | Observation | What it rules out | | --- | --- | | A strictly weaker adversary outscores a stronger one at equal cost | The white-box number as a property of the model | | Removing the perturbation budget still does not drive failure to total | The search working at any radius | | A single-step attack outscores an iterative one at the same radius | The local direction being informative along the path | The third deserves a word because it is the least intuitive. Iterating means taking many small steps and re-projecting into the allowed set after each one, and at the same radius it should dominate a single step - it is the same direction, refined. If one step beats many, the direction is only accidentally right at the starting point and iterating walks the search into a region where the reading is meaningless. ## Triaging the finding rather than filing it The chair here is the person who has a result that reproduces once in five tries and has to explain an inverted table to a room. Two failure modes are common. The first is to call it noise and re-run until the expected ordering appears - which is selection, not evaluation. The second is to quote the low white-box number because it is the more flattering row and the more standard attack. Both bury the finding. The defensible path is short: 1. Establish that the budget really was equal, and say in what unit. 2. Repeat both runs across several seeds and report the spread, not one draw. Non-reproducibility is itself evidence of a stochastic component in the defence. 3. Take the decision-only success rate as a floor. Whatever the model's true failure rate is, it is at least that. 4. Re-run the white-box attack with the masking accounted for - more steps, more restarts, and if the defence is random, an objective averaged over its randomness rather than read from one draw. 5. Report the highest success any adversary achieved as the model's failure rate, with the perturbation set and radius stated beside it. ## What survives about the defence Something real usually does. Masking raises the price of the attack: more evaluations, more restarts, more careful setup. Against an opportunistic attacker that is not nothing. But it is a **cost control, not a boundary** - it is paid once by anyone who reads the defence and adapts, and it does not shrink the set of inputs the model gets wrong by a single row. Presenting it as robustness is where the harm is done, because the downstream reader plans as if the failure rate is 2%.
- Someone says the inversion is just run-to-run noise. How do you settle it?Repeat both attacks across several seeds and report the spread rather than one draw. If the ordering holds across seeds, it is not noise. If the white-box run itself swings wildly between seeds, that instability is a finding too - it points at a randomised component in the defence, which is one of the mechanisms that scrambles the direction in the first place.
- Which number do you put in the report once you have both?The highest success rate any adversary achieved, with the perturbation set, the radius, the access assumption and the budget stated beside it. The decision-only run is a lower bound on the model's true failure rate, so a white-box row below it is not reportable at all. One number without those columns is not a comparable claim.
- Why is a single-step attack outscoring an iterative one at the same radius also a warning sign?Iterating is the same direction refined: many small steps with a re-projection into the allowed set after each. At a fixed radius it should dominate one large step. When it does not, the local reading is only incidentally right at the start point, and the iterations are walking the search into a region where the direction no longer describes the model.
saying these in an interview costs you the question
- Calls the inverted ordering run-to-run noise
- Quotes the flattering white-box row anyway
- Compares attacks without holding the budget equal
- Concludes the black-box attack is simply stronger
- Treats raised attack cost as raised robustness