You attach two preprocessing defence objects to a wrapped classifier in an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox, then run a gradient-based attack from the same library. What decides the order they run in and whether the attack optimises through them, and why does it matter that a preprocessing defence can be configured to run only at training time?
answer
- chain in attachment order
- fit-time vs predict-time defence
- budget measured in which space
- non-differentiable link, substituted gradient
- verify the transform actually ran
basics
~20 sThey form an ordered chain applied in the order you attach them, so the second sees the first one's output. Each declares whether it runs at training, at inference, or both. A defence that runs only at training is absent when the attack queries the model, so the run measures a pipeline nobody serves.
solid answer
~60 sTwo separate switches decide what your attack is really attacking. **Order.** Stacked preprocessor objects compose as a chain in attachment order. That is not cosmetic: putting a quantiser before or after a normaliser changes what a fixed perturbation budget buys, because the budget is expressed in one space and consumed in another. **Lifecycle.** Each preprocessor declares whether it applies during training, during inference, or both. A defence that applies only at fit time is not in the path when the attack calls the model, so you evaluate an undefended pipeline while believing you tested a defended one. The reverse mistake is just as bad: a defence present at evaluation but absent from the served path inflates the reported robustness of the thing you ship. **Visibility to gradients.** A gradient attack can only optimise through a transform whose derivative the chain exposes. If a stacked defence is non-differentiable, the attack falls back to whatever estimate the library supplies, or optimises from inside the transformed space, which quietly makes the perturbation budget mean something different. Either way the reported number is about the attack's access, not the model's robustness.
go deeper
Knows defences attach to the wrapper and that the attack runs against the wrapper, not the bare model.
Explains chain order, the training-versus-inference lifecycle switch, and why a fit-time-only defence is invisible to the attack.
Adds the differentiability question, checks empirically that the chain ran, and reports the space the perturbation budget was measured in.
Insists the shipped inference path and the evaluated chain are the same artefact, and makes that equivalence a deliverable rather than an assumption.
**The mental model: the attack is handed the chain, not the model** In an adversarial-robustness toolkit the attack never touches the framework model. In ART you construct, say, `art.estimators.classification.PyTorchClassifier(model=..., preprocessing_defences=[a, b])` and then hand *that estimator* to an attack class in `art.attacks.evasion`. The attack calls the estimator's `predict` and `loss_gradient`. So every question of the form "did the defence work" reduces to two mechanical facts: which links were in the chain at the moment the attack ran, and which of them the attack could see through. **Order** `preprocessing_defences` is a list, and the estimator applies it in list order: the second object receives the first object's output. This is not cosmetic. Put a quantiser before a normaliser and a fixed perturbation budget buys a different effective perturbation than the reverse order, because the budget is defined in one space and spent after an intervening transform has rescaled or discretised it. If you compare two defences without pinning the rest of the chain, you have compared chains rather than defences, and the difference you are reporting may be entirely an artefact of where the budget was measured. **Lifecycle** Every ART preprocessor carries `apply_fit` and `apply_predict`. A defence with `apply_fit=True, apply_predict=False` is an augmentation-style object: it shaped training and is completely absent when the attack calls `predict`. The run still completes and still prints a robust-accuracy number — it is just the number for an undefended prediction path. The mirror-image error is as bad and less often caught: a defence present at evaluation time but not in the served inference path inflates the reported robustness of the artefact the client actually ships. **Differentiability, and the substituted gradient** A gradient attack needs a derivative through every link. ART handles this with a method on the preprocessor itself: framework-native subclasses (`PreprocessorPyTorch` and friends) are differentiable, so autograd flows through them normally, while a plain `Preprocessor` exposes `estimate_gradient(x, grad)` and the default implementation simply passes the incoming gradient through unchanged. That default is straight-through, or backward-pass-differentiable-approximation, and it is a guess. When it is a bad guess, the attack optimises against a surface that is not the one it is being scored on, robust accuracy comes out high, and nothing anywhere reports that a substitution happened. The resulting number is a statement about the quality of that gradient substitute, not about the model. ``` # what the attack differentiates x = input_batch for d in estimator.preprocessing_defences: # list order x = d(x) # apply_predict must be True logits = framework_model(x) # backward: differentiable defence -> real gradient # plain Preprocessor -> estimate_gradient(), identity by default ``` **What it costs** The verification is cheap and the mistake is expensive, which is the whole argument. Confirming what reached the model on a probe batch is a few minutes of an engineer's time. A perturbation-budget sweep — a 50-step attack over 1,000 test images at eight budgets — is a few hundred thousand forward-and-backward passes, well under a GPU-hour on a small model. Skipping both and shipping the conclusion costs whatever the client then spends: on a trainer-style defence that is a full retraining pipeline built on a number that was measuring an empty chain. **Where the number misleads** Two readings are wrong in opposite directions. A fit-time-only defence produces a defended-looking number for an undefended path, so the defence looks better than it is because it was never tested. A non-differentiable stacked defence produces a high robust accuracy because the attack could not optimise, so the defence looks better than it is because the attack was crippled. Both print cleanly. The distinguishing evidence is the budget curve: a real defence degrades smoothly and hits the floor once the budget is large, while a chain that blocks the optimiser stays stubbornly flat as the budget grows. Why some transforms systematically degrade gradients is theory owned by the adversarial-ML defence literature; what this leaf owns is that the chain, not the model, is what your harness measured. **What to check and what to record** Query the estimator on a crafted probe input and inspect what reaches the model, instead of trusting constructor arguments. Then record, in the report: the exact ordered chain, each object's `apply_fit`/`apply_predict` setting, whether the attack differentiated through the chain or fell back to an estimated gradient, the space the perturbation budget was defined in, and whether the evaluated chain is byte-for-byte the chain the client serves.
- How do you verify that the defence chain you configured is really in the path the attack queries?Call the wrapper on a crafted probe input and inspect what reaches the model, rather than trusting the constructor arguments. If the expected transform did not happen, the whole run is an undefended baseline.
- Why can swapping the order of two preprocessor defences change robust accuracy?Because the second transform operates on the first one's output, and the perturbation budget is defined in one space but spent in another. Order changes the effective perturbation the model finally receives.
saying these in an interview costs you the question
- Assuming defence objects commute.
- Never checking whether the configured defence was actually applied during the attack run.
- Reporting robust accuracy without saying whether the attack had gradients through the defence chain.
- Evaluating a chain that differs from the one the client serves.