You inherit a pipeline documented as 91% accurate under attack with its cleaning stage enabled - what do you ask before trusting that?
answer
- look for the adaptive column first
- norm and radius are the threat model
- attack strength is a knob
- a missing row means unevaluated
- you also need what it cost
basics
~20 sAsk whether the cleaning stage was inside the attacker's optimisation, and what norm, radius, steps and restarts the attack used. If every row was measured against an attacker who ignored the stage, the figure describes that attacker, not the pipeline.
solid answer
~50 sThree questions, in order. First: was the transform inside the attacker's loop? A robustness figure produced by an attack fitted against the bare model and then replayed through the cleaning stage measures how much the stage changed your own function, not what an adversary can do. Second: what norm and what radius? Those two together *are* the threat model, and a figure quoted without both is not comparable to any other figure. Third: how many steps and restarts, and is there a bare-model row at the same settings to compare against? Absent an adaptive row, the honest record is "unevaluated", not 91%. If re-running the attack through the chain is expensive, accept one best-effort adaptive row at the same radius with its steps reported - one is enough to tell you whether the claim survives, and it usually does not.
code
text · 9 linesattack evaluation - inspection pipeline, cleaning stage v4
pipeline attack transform in attacker's loop norm radius steps robust acc
---------------------------------------------------------------------------------------------
cleaned input single-step no Linf 8/255 1 91%
cleaned input iterative no Linf 8/255 20 78%
bare model iterative n/a Linf 8/255 20 4%
...
(no row with the transform inside the attacker's optimisation)go deeper
Recall that a number measured under attack describes the attack that was run. Know to look for whether the defence was part of what the attacker optimised against before reading the percentage.
Explain why norm, radius, steps and restarts belong to the claim, and why replaying stored adversarial inputs through a new pipeline measures your own transform rather than an adversary.
Show you would demand one adaptive row before letting the figure into a risk record, and that you would separate the engineering justification for preprocessing from the security claim attached to it.
Own what the organisation may say externally: which claims survive an adaptive evaluation, what gets recorded as unevaluated, and who signs off when a number in a customer document cannot be reproduced.
## What the table in front of you actually says Read the column that says whether the cleaning stage was inside the attacker's optimisation. If every defended row says *no*, then every attack in the table was fitted against some other function and replayed through your pipeline. The result that comes back is a measurement of one thing: how much your transform perturbs an input that was crafted for a different function. That is a real fact about your transform. It is not a fact about your security, because no deployed pipeline gets to assume an adversary who did not read it. The bare-model row is the tell that this was never checked. It shows that the same attack, run at the same norm and radius, collapses the undefended model - so the evaluators knew how to run the attack and how to make it work. What is missing is the one row that costs a bit more to produce and is the only row that matters: the same attack, with the cleaning stage inside the search. ## The questions, in the order they change your decision **1. Was the transform inside the attacker's loop?** This is binary and it dominates everything else. If the answer is no, the rest of the table is describing a non-adaptive attacker and the correct entry in your risk record is "unevaluated", not the number printed. **2. What norm and what radius?** The norm and the radius together are the threat model. A bound that lets every coordinate move a little is a different adversary from one that lets a few coordinates move a lot, and robustness bought against one family transfers poorly to another. A figure with no norm and no radius is not comparable to any other figure, including the vendor's next one. **3. How much search?** Steps, restarts, and whether the attack was iterative or single-step. Attack strength is a knob, and a weak attack produces a flattering number. An evaluation that reports one number and no search budget has not told you what adversary it modelled. **4. What did it cost you?** The clean-accuracy column, and ideally accuracy per class or per severity band, because a cleaning stage's loss concentrates on faint and rare cases while the aggregate barely moves. The table shows what you gained; you also need what you paid. ## What you can honestly write afterwards Suppose you get the adaptive row and it lands near the bare-model row, which is the usual outcome for this defence class. That is not a reason to rip the stage out. Preprocessing frequently exists for reasons that have nothing to do with an adversary: sensor-noise removal, format normalisation, matching the input size and statistics the model was trained on. Those reasons are good and should be recorded as such. What changes is the claim. You strike "robust to adversarial inputs" and replace it with something you can defend: - removes attacks that were not built against this pipeline, including reused published ones; - charges an adaptive attacker a measured multiple of the undefended attack's search cost - state the multiple, or state that you have not measured it; - costs *x* points of clean accuracy overall and *y* on the classes where the loss concentrates. ## Why this is worth pushing on when the team says it is expensive Re-running an adaptive evaluation is more work than replaying stored inputs, and teams resist it. The minimum acceptable answer is one best-effort adaptive row: same norm, same radius, same success criterion, the transform inside the search, steps reported. One row is enough, because the outcome for this defence class is rarely ambiguous - either the number holds up under an attacker who modelled the stage, or it collapses toward the bare-model row. Either result is worth more than the whole table you inherited, and the second result is the one that stops a robustness claim from reaching a customer document where somebody will rely on it. ## The general rule behind the specific one A high figure under attack proves that *the attack that was run* failed. It never proves the model is robust. Every column that describes the attacker - access, norm, radius, steps, whether the defence was in the loop - is part of the claim, and a claim missing any of them is quoting a number without its units.
- The team says re-running the attack through the chain is too expensive. What is the minimum you accept?One best-effort adaptive row: same norm, same radius, same success criterion, the cleaning stage inside the search, with steps and restarts reported. One row settles it for this defence class. Without it, record the pipeline as unevaluated rather than carrying the existing figure forward into anything anyone will rely on.
- The adaptive row comes back near the bare-model number. Do you remove the cleaning stage?Usually not. It may be there for sensor-noise removal, format normalisation or matching the training distribution, all of which are legitimate. Remove the robustness claim, keep the stage on its engineering merits, and account its clean-accuracy cost explicitly so nobody re-justifies it as security later.
- Why is a robust-accuracy figure with no norm and no radius unusable?Because those two are the threat model. Allowing every coordinate to move slightly and allowing a few coordinates to move a lot are different adversaries, and robustness against one transfers poorly to the other. Without both, the figure cannot be compared to another figure or mapped onto anything an attacker would actually do.
saying these in an interview costs you the question
- Accepts a robustness figure without asking how the attack was run
- Treats a high number under attack as proof the model is robust
- Quotes robust accuracy with no norm and no radius
- Assumes replaying stored adversarial inputs is an adaptive evaluation
- Rips out preprocessing that exists for legitimate engineering reasons