skip to content

Defence Modules

Preprocessor, detector, postprocessor, trainer and transformer defences stack as objects, and the harness that attacked a model will flatter a defence that only blinds that attack.

on this pageshow

explore

questions

5

In an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox, defences ship as objects of several kinds: preprocessor, postprocessor, detector, trainer and transformer. Where does each kind sit relative to a call to the wrapped model, and which of them change the model weights?

level: juniorimportance: must knowfreq 55%

answer

  1. input path vs output path
  2. detector returns a flag, not a label
  3. trainer and transformer change weights
  4. attach-and-evaluate vs retrain-and-evaluate
  5. defence must live inside the wrapper

basics

~20 s

A preprocessor defence transforms inputs before the model sees them. A postprocessor alters what comes back, usually the scores. A detector flags a sample as adversarial instead of classifying it. A trainer defence retrains the weights. A transformer hands you back a new, modified model object. Only trainer and transformer touch weights.

solid answer

~50 s

The kinds differ by **where in the call path they act** and by **what they cost you**. - **Preprocessor** — sits on the input path (resizing, quantising, smoothing, denoising). Cheap: it attaches to weights you already have. - **Postprocessor** — sits on the output path, editing the returned scores (rounding, adding noise, hiding all but the top label). It changes what an attacker can read back, not what the model computed. - **Detector** — a separate decision object: it answers *does this input look adversarial*, and returns a flag rather than a class. The pipeline now has a third outcome, rejection. - **Trainer** — a training procedure that produces new weights. Every hyperparameter setting you want to compare is another full retrain. - **Transformer** — consumes a model and emits a different model object, so downstream you evaluate the new one. The practical split for planning work: preprocessor, postprocessor and detector are attach-and-evaluate; trainer and transformer are retrain-and-evaluate, which is an order of magnitude more compute.

go deeper

for a junior

Name the kinds and say which sit before the model, which after, and that trainer defences retrain. Knowing that a detector returns a flag is the key one.

for a middle

Adds the lifecycle detail: chain order, whether a preprocessor runs at fit time or predict time, and why weight-changing defences cost a retrain per variant.

for a senior

Ties the taxonomy to evaluation design: which attack classes can meaningfully test which defence kind, and how to confirm the object is inside the queried wrapper.

for a principal

Uses the taxonomy as a cost model when planning an engagement, and states which defence kinds a client can adopt without a retraining pipeline.

## The chain a defence object joins An adversarial-robustness toolkit does not patch the model; it **wraps** it. In the Adversarial Robustness Toolbox (ART) the wrapper is an **estimator** such as `art.estimators.classification.PyTorchClassifier`, and its constructor takes `preprocessing_defences=` and `postprocessing_defences=` lists. Everything downstream — every attack class in `art.attacks` — talks to that estimator's `predict`, `loss_gradient` and `fit` methods, never to the raw framework model. So the five defence kinds are best read as answers to one question: at which point in the estimator's call path does this object execute, and does it leave the weights alone? ## The five defence kinds - **Preprocessor** — `art.defences.preprocessor.*` (feature squeezing, JPEG compression, spatial smoothing, total-variance minimisation). It is a callable on the input batch, invoked inside the estimator before the framework model runs. Two constructor booleans, `apply_fit` and `apply_predict`, decide whether it runs during `fit`, during `predict`, or both. Cost: no training at all; one extra tensor operation per batch, so a defended evaluation sweep costs an undefended sweep plus a few percent of wall-clock. - **Postprocessor** — `art.defences.postprocessor.*` (`Rounded`, `GaussianNoise`, `HighConfidence`, `ClassLabels`, `ReverseSigmoid`). It edits the returned score vector after the model has already computed it. It changes what a caller can read back, not what the model computed. Cost: negligible in compute; the real cost is downstream, because anything in your own stack that consumes probabilities now sees mangled ones. - **Detector** — for example `art.defences.detector.evasion.BinaryInputDetector`, which learns on raw inputs, or `BinaryActivationDetector`, which learns on a chosen hidden layer of the model being defended. It returns a per-sample adversarial-or-benign decision plus a score, not a class. Cost: you must fit it, and fitting needs a corpus of adversarial examples you have to generate first — usually the largest hidden cost on this list. - **Trainer** — `art.defences.trainer.*`, such as `AdversarialTrainer` or `AdversarialTrainerMadryPGD`. This is a training procedure: you hand it the estimator and it produces new weights. Cost: a full training run, and adversarial training with a k-step inner attack does roughly k+1 times the forward-backward work of ordinary training. On CIFAR-scale ResNets that turns a one-to-two GPU-hour job into eight to sixteen; at ImageNet scale it is days per setting. - **Transformer** — `art.defences.transformer.*`, such as defensive distillation on the evasion side or STRIP on the poisoning side. It consumes a fitted estimator and returns a *different* estimator object. Cost: another training or calibration pass, plus the part people forget — every number you measured against the original estimator now describes an artefact you are no longer shipping. ## Where the numbers mislead Three ways, in descending order of how often they actually happen. 1. First, the defence was applied by hand around the wrapper instead of being passed into it, so `attack.generate()` never saw it: the run is an undefended baseline printed under a defended heading, and nothing raises an error. 2. Second, a preprocessor was constructed with `apply_fit=True, apply_predict=False`, which is a perfectly legitimate augmentation setting and is simply absent when the attack calls `predict` — again no error, just a number. 3. Third, a postprocessor is judged with a white-box gradient attack: that attack differentiates the model's loss and never reads the postprocessed scores, so `Rounded` or `ReverseSigmoid` measures as worthless when its actual job is starving query-based, decision-based and model-extraction attacks that see nothing but `predict` output. The same object reads as useless under one attack class and valuable under another, and the taxonomy is what tells you which reading is honest. ## The planning consequence Split the five kinds into two buckets by **artefact cost** rather than by name. | Bucket | Kinds | Cost for N settings | |---|---|---| | Attach-and-evaluate | preprocessor, postprocessor, and a detector once it is fitted | N evaluation sweeps | | Retrain-and-evaluate | trainer and transformer | N training runs *plus* N sweeps | On a single GPU that is the difference between screening a dozen configurations in an afternoon and affording two configurations in a fortnight, and it is also the difference between a defence a client can adopt tomorrow and one that obliges them to own a retraining pipeline forever. ## What to check - Call the estimator on a probe input and inspect what actually reaches the framework model, rather than trusting that the constructor accepted your argument. - Confirm `apply_predict` is set for anything you expect the attack to face. - Confirm the attack was constructed against the defended estimator instance and not a copy taken before the defence was attached. - Finally, measure clean accuracy on the same wrapper: every kind here except a pure postprocessor can cost accuracy on legitimate inputs, and a defence whose price you did not measure is not a result.

  • Why is a postprocessor defence a poor thing to judge with a white-box gradient attack?
    The gradient attack differentiates the model, it does not consume the postprocessed scores, so score rounding or noise barely touches it. Postprocessors target query-based, decision-boundary and extraction attacks, and should be judged with those.
  • Two defence objects attach to the same wrapper. Which runs first?
    The chain runs in the order the objects were attached, each seeing the previous one's output. Reordering a normaliser and a quantiser can change both the defended accuracy and the effective size of the attacker's perturbation budget.
  • Which kinds can you evaluate without retraining?
    Preprocessor, postprocessor and detector, since they wrap existing weights. Trainer and transformer produce new weights, so every variant is a fresh training run followed by a fresh evaluation.

saying these in an interview costs you the question

  • Treating every defence as an input transform and missing that some return new weights.
  • Reporting a detector's result as accuracy, with no account of rejected samples.
  • Applying the defence by hand around the wrapper and still calling the run a white-box evaluation of the defended model.
  • Assuming defence objects are order-independent.

context

open as a page

You attach two preprocessing defence objects to a wrapped classifier in an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox, then run a gradient-based attack from the same library. What decides the order they run in and whether the attack optimises through them, and why does it matter that a preprocessing defence can be configured to run only at training time?

level: middleimportance: must knowfreq 50%

basics

~20 s

They form an ordered chain applied in the order you attach them, so the second sees the first one's output. Each declares whether it runs at training, at inference, or both. A defence that runs only at training is absent when the attack queries the model, so the run measures a pipeline nobody serves.

open as a page

After you attach a defence object to a wrapped classifier in an adversarial-robustness toolkit, the robust accuracy under your configured attack rises from near zero to about three quarters. What do you check before reporting that as a robustness gain?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Assume the attack broke rather than the model became robust. Re-tune the attack against the defended pipeline, raise iterations and restarts, and check that an unbounded perturbation still drives accuracy to zero. Compare against a query-based attack and a transfer attack. Confirm the defence really sits inside the wrapper you attacked.

open as a page

You attach a detector defence object to a wrapped classifier in an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox. It returns an adversarial-or-benign decision per sample rather than a class label. What has to change in how you score the pipeline, and what number must you report next to its detection rate?

level: middleimportance: should knowfreq 45%

basics

~20 s

The pipeline now has three outcomes, not two: correct, wrong, and rejected. Decide up front how a rejected sample counts, for adversarial and for clean inputs separately. Report the false-positive rate on clean data beside the detection rate, because the threshold is tunable and a detection rate alone hides what it costs benign traffic.

open as a page

A client asks you to shortlist defences for an image classifier using an adversarial-robustness toolkit, and you have one GPU for two weeks. The candidates split into defence objects that attach to the existing weights and trainer-style defences that need a full retrain per hyperparameter setting. How do you allocate the compute, and what do you tell the client the shortlist does and does not cover?

level: principalimportance: should knowfreq 35%

basics

~20 s

Split by cost. Attach-only defences reuse the trained weights, so each costs one evaluation sweep and can face many attack settings. Every trainer variant costs a full retrain, so you can afford only a couple of settings. Spend most of the compute on deep adaptive evaluation of a short list, not a broad shallow sweep.

open as a page