In an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox, defences ship as objects of several kinds: preprocessor, postprocessor, detector, trainer and transformer. Where does each kind sit relative to a call to the wrapped model, and which of them change the model weights?
answer
- input path vs output path
- detector returns a flag, not a label
- trainer and transformer change weights
- attach-and-evaluate vs retrain-and-evaluate
- defence must live inside the wrapper
basics
~20 sA preprocessor defence transforms inputs before the model sees them. A postprocessor alters what comes back, usually the scores. A detector flags a sample as adversarial instead of classifying it. A trainer defence retrains the weights. A transformer hands you back a new, modified model object. Only trainer and transformer touch weights.
solid answer
~50 sThe kinds differ by **where in the call path they act** and by **what they cost you**. - **Preprocessor** — sits on the input path (resizing, quantising, smoothing, denoising). Cheap: it attaches to weights you already have. - **Postprocessor** — sits on the output path, editing the returned scores (rounding, adding noise, hiding all but the top label). It changes what an attacker can read back, not what the model computed. - **Detector** — a separate decision object: it answers *does this input look adversarial*, and returns a flag rather than a class. The pipeline now has a third outcome, rejection. - **Trainer** — a training procedure that produces new weights. Every hyperparameter setting you want to compare is another full retrain. - **Transformer** — consumes a model and emits a different model object, so downstream you evaluate the new one. The practical split for planning work: preprocessor, postprocessor and detector are attach-and-evaluate; trainer and transformer are retrain-and-evaluate, which is an order of magnitude more compute.
go deeper
Name the kinds and say which sit before the model, which after, and that trainer defences retrain. Knowing that a detector returns a flag is the key one.
Adds the lifecycle detail: chain order, whether a preprocessor runs at fit time or predict time, and why weight-changing defences cost a retrain per variant.
Ties the taxonomy to evaluation design: which attack classes can meaningfully test which defence kind, and how to confirm the object is inside the queried wrapper.
Uses the taxonomy as a cost model when planning an engagement, and states which defence kinds a client can adopt without a retraining pipeline.
## The chain a defence object joins An adversarial-robustness toolkit does not patch the model; it **wraps** it. In the Adversarial Robustness Toolbox (ART) the wrapper is an **estimator** such as `art.estimators.classification.PyTorchClassifier`, and its constructor takes `preprocessing_defences=` and `postprocessing_defences=` lists. Everything downstream — every attack class in `art.attacks` — talks to that estimator's `predict`, `loss_gradient` and `fit` methods, never to the raw framework model. So the five defence kinds are best read as answers to one question: at which point in the estimator's call path does this object execute, and does it leave the weights alone? ## The five defence kinds - **Preprocessor** — `art.defences.preprocessor.*` (feature squeezing, JPEG compression, spatial smoothing, total-variance minimisation). It is a callable on the input batch, invoked inside the estimator before the framework model runs. Two constructor booleans, `apply_fit` and `apply_predict`, decide whether it runs during `fit`, during `predict`, or both. Cost: no training at all; one extra tensor operation per batch, so a defended evaluation sweep costs an undefended sweep plus a few percent of wall-clock. - **Postprocessor** — `art.defences.postprocessor.*` (`Rounded`, `GaussianNoise`, `HighConfidence`, `ClassLabels`, `ReverseSigmoid`). It edits the returned score vector after the model has already computed it. It changes what a caller can read back, not what the model computed. Cost: negligible in compute; the real cost is downstream, because anything in your own stack that consumes probabilities now sees mangled ones. - **Detector** — for example `art.defences.detector.evasion.BinaryInputDetector`, which learns on raw inputs, or `BinaryActivationDetector`, which learns on a chosen hidden layer of the model being defended. It returns a per-sample adversarial-or-benign decision plus a score, not a class. Cost: you must fit it, and fitting needs a corpus of adversarial examples you have to generate first — usually the largest hidden cost on this list. - **Trainer** — `art.defences.trainer.*`, such as `AdversarialTrainer` or `AdversarialTrainerMadryPGD`. This is a training procedure: you hand it the estimator and it produces new weights. Cost: a full training run, and adversarial training with a k-step inner attack does roughly k+1 times the forward-backward work of ordinary training. On CIFAR-scale ResNets that turns a one-to-two GPU-hour job into eight to sixteen; at ImageNet scale it is days per setting. - **Transformer** — `art.defences.transformer.*`, such as defensive distillation on the evasion side or STRIP on the poisoning side. It consumes a fitted estimator and returns a *different* estimator object. Cost: another training or calibration pass, plus the part people forget — every number you measured against the original estimator now describes an artefact you are no longer shipping. ## Where the numbers mislead Three ways, in descending order of how often they actually happen. 1. First, the defence was applied by hand around the wrapper instead of being passed into it, so `attack.generate()` never saw it: the run is an undefended baseline printed under a defended heading, and nothing raises an error. 2. Second, a preprocessor was constructed with `apply_fit=True, apply_predict=False`, which is a perfectly legitimate augmentation setting and is simply absent when the attack calls `predict` — again no error, just a number. 3. Third, a postprocessor is judged with a white-box gradient attack: that attack differentiates the model's loss and never reads the postprocessed scores, so `Rounded` or `ReverseSigmoid` measures as worthless when its actual job is starving query-based, decision-based and model-extraction attacks that see nothing but `predict` output. The same object reads as useless under one attack class and valuable under another, and the taxonomy is what tells you which reading is honest. ## The planning consequence Split the five kinds into two buckets by **artefact cost** rather than by name. | Bucket | Kinds | Cost for N settings | |---|---|---| | Attach-and-evaluate | preprocessor, postprocessor, and a detector once it is fitted | N evaluation sweeps | | Retrain-and-evaluate | trainer and transformer | N training runs *plus* N sweeps | On a single GPU that is the difference between screening a dozen configurations in an afternoon and affording two configurations in a fortnight, and it is also the difference between a defence a client can adopt tomorrow and one that obliges them to own a retraining pipeline forever. ## What to check - Call the estimator on a probe input and inspect what actually reaches the framework model, rather than trusting that the constructor accepted your argument. - Confirm `apply_predict` is set for anything you expect the attack to face. - Confirm the attack was constructed against the defended estimator instance and not a copy taken before the defence was attached. - Finally, measure clean accuracy on the same wrapper: every kind here except a pure postprocessor can cost accuracy on legitimate inputs, and a defence whose price you did not measure is not a result.
- Why is a postprocessor defence a poor thing to judge with a white-box gradient attack?The gradient attack differentiates the model, it does not consume the postprocessed scores, so score rounding or noise barely touches it. Postprocessors target query-based, decision-boundary and extraction attacks, and should be judged with those.
- Two defence objects attach to the same wrapper. Which runs first?The chain runs in the order the objects were attached, each seeing the previous one's output. Reordering a normaliser and a quantiser can change both the defended accuracy and the effective size of the attacker's perturbation budget.
- Which kinds can you evaluate without retraining?Preprocessor, postprocessor and detector, since they wrap existing weights. Trainer and transformer produce new weights, so every variant is a fresh training run followed by a fresh evaluation.
saying these in an interview costs you the question
- Treating every defence as an input transform and missing that some return new weights.
- Reporting a detector's result as accuracy, with no account of rejected samples.
- Applying the defence by hand around the wrapper and still calling the run a white-box evaluation of the defended model.
- Assuming defence objects are order-independent.