An attack with the perturbation budget removed still leaves a defended fraud model at 15% robust accuracy - what do you conclude?
answer
- it is a control, not an attack
- you already know its answer
- unlimited freedom should mean total failure
- so the residual is the optimiser
- and it voids the other rows too
basics
~20 sThat the search is broken at every radius, not that the model is robust. With no limit on how far an input may move, an attacker can hand the model a transaction that genuinely belongs to the other class, so success should approach total; anything left standing is the optimiser failing.
solid answer
~50 sThe unbounded run is a control, and its expected result is known in advance: remove the constraint and the attacker can walk the input all the way into the target class, so failure should go to essentially 100%. A residual 15% is therefore not a property of the model - it is proof that the search never worked, and by implication every bounded number from the same harness is uninterpretable too, including the flattering one at the small radius. The finding is not *the defence is strong*; it is *the evaluation is void*. What you do next is re-run adaptively - more steps and restarts, an objective averaged over any randomness in the pipeline, and a strictly weaker gradient-free adversary as a lower bound - and report whatever failure rate survives, scoped to its stated perturbation set and radius.
go deeper
Recall that with no limit on how far an input may change, an attacker can simply supply an input that really belongs to the other class, so an attack that still fails has failed for its own reasons.
Be ready to explain why a failed unbounded control invalidates the small-radius rows from the same harness, and to name the mechanisms that make a search direction uninformative.
Walk the triage: rule out clipping in the harness first, vary seeds, spend optimisation budget, floor the result with a gradient-free adversary, then re-run bounded and report with full columns.
Own the standard that an evaluation must carry a control whose answer is known, and be prepared to tell stakeholders the review produced no robustness number rather than a flattering one.
## Why an unbounded run is a control and not an attack Every robustness number is quoted inside a perturbation set: a norm and a radius that together say how far an input is allowed to move. Most of the evaluation happens at a small radius, where the interesting question lives. The unbounded run exists for a different purpose: it is a **sanity check with a known answer**. If the attacker may change the input arbitrarily, the problem stops being adversarial at all. They can construct a transaction that is genuinely a decline - or genuinely an approve - and the model, being a reasonable classifier, will classify it that way. Success approaches total by construction. This is not a claim about any particular attack; it follows from the model being accurate on real inputs of the target class. So the unbounded row is the row you already know the answer to, and its value is entirely diagnostic: it tells you whether the harness can reach an answer it must be able to reach. ## Reading a residual 15% When 15% survives, the search failed to move the input into the target class even with unlimited freedom. There is no radius at which such a search is trustworthy, because the small-radius problem is strictly harder than the one it just failed. Two consequences follow, and the second is the one people miss: 1. The 15% is not robust accuracy. It is the fraction of inputs this optimiser could not move, at any distance. 2. **Every other number from the same harness is void**, including the impressive small-radius figure that prompted the review. They were produced by the same broken search. The conclusion to write down is about the evaluation, not the model. Reporting *the model retains 15% accuracy under unbounded attack* would be a serious error of direction: it attributes to the model a fact about the optimiser. ## What usually causes it The unbounded control fails for the same reasons the bounded runs did: the defence made the direction the search steers by uninformative. Concretely, a non-differentiable or discontinuous step in front of the model means the local reading no longer describes the composed function; a randomised step means each reading is a single draw and small differences drown in the spread; a step that saturates means the reading is numerically dead and the search does not move at all. It can also be a plain harness bug - an input clipped back into a valid feature range after every step, so the *unbounded* run was never unbounded. Rule that out first, because it is the cheapest explanation and it is common. ## The triage, in order 1. **Verify the control was really unconstrained.** Check for range clipping, validity constraints on categorical fields, or a projection step left enabled. A fraud model over tabular features has many places where a legitimate-looking constraint silently bounds the run. 2. **Vary seeds.** If the residual moves a lot between runs, there is randomness in the pipeline and the objective must be averaged over it before any number is read. 3. **Spend budget.** Ten times the steps and several restarts. Flat response confirms the direction is uninformative rather than the problem being hard. 4. **Run a strictly weaker adversary as a floor.** A search that never reads a gradient and only sees the returned decision. Whatever it achieves is a lower bound on the model's true failure rate, and if it exceeds the white-box row the ordering has inverted, which is the same diagnosis arriving by a second route. 5. **Re-run bounded, adaptively, and report that.** With the radius, the norm, the access assumption, the step and restart counts, and the seeds beside it. ## What you say to the room The uncomfortable part is that the review's output is *we cannot currently state a robustness number for this build*, which is less satisfying than either a good number or a bad one. It is also the only defensible output. The secondary finding - that the defence multiplies the cost of running an attack - can be stated honestly and separately, as cost, because that part is real: an adversary who has read the defence pays more evaluations to get where they were going. They pay it once. ## The general principle worth carrying An evaluation should include at least one run whose answer you already know. When that run returns the wrong answer, the finding is about the instrument, and no other number from the same instrument may be quoted until it is fixed. This is the ordinary discipline of a control, and adversarial robustness evaluation is unusually easy to run without one because the headline attack always produces a number that looks like a result.
- What is the cheapest explanation you should rule out before blaming the defence?That the run was not actually unbounded. Feature-range clipping, validity constraints on categorical fields, or a projection step left enabled in the harness will silently bound a run labelled unconstrained. On a tabular fraud model there are several such places. Check them before concluding anything about gradients, because a harness bug and a masked gradient produce an identical-looking row.
- Does the residual 15% mean anything at all about the model?Almost nothing. It is the fraction of inputs this particular search could not move at any distance, which is a statement about the search. The only weak inference available is that those inputs may sit somewhere the optimiser handles badly, and that is a lead for debugging the harness rather than a property to report.
- How do you report the result when the review ends without a usable number?As two separate statements. First, no robustness figure can be stated for this build because the evaluation failed its control, and here is the control that failed. Second, the defence does raise the cost of mounting an attack by some measured factor, which is a real but different claim. Merging them into one reassuring number is the failure this review exists to prevent.
saying these in an interview costs you the question
- Reports 15% as robust accuracy under unbounded attack
- Keeps the small-radius number from the same harness
- Never checks whether the unbounded run was clipped
- Concludes the defence is unusually strong
- Treats a failed control as an inconvenient outlier