You attach a detector defence object to a wrapped classifier in an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox. It returns an adversarial-or-benign decision per sample rather than a class label. What has to change in how you score the pipeline, and what number must you report next to its detection rate?
answer
- three outcomes: correct, wrong, rejected
- detection rate needs its false-positive twin
- threshold = operating point
- fitted on one attack, tested on the same attack
- rejection needs a downstream action
basics
~20 sThe pipeline now has three outcomes, not two: correct, wrong, and rejected. Decide up front how a rejected sample counts, for adversarial and for clean inputs separately. Report the false-positive rate on clean data beside the detection rate, because the threshold is tunable and a detection rate alone hides what it costs benign traffic.
solid answer
~60 sA detector is a binary classifier bolted onto the pipeline, with its own threshold, so it converts your two-way accuracy into a three-way outcome table. Define the accounting before you run anything: - **Adversarial input, detected** — usually counted as a defended success, since the system refused rather than misclassified. - **Adversarial input, missed** — falls through to the model and is scored as usual. - **Clean input, flagged** — a false positive, a real availability cost that has nothing to do with the attack. The number that makes a detection rate meaningful is the **clean false-positive rate at the same threshold**. Any detector can reach a high detection rate by flagging everything, so a single detection figure without its operating point is uninterpretable; report a pair, or the curve. The second thing to state is what the detector was fitted on. Detectors are usually trained on adversarial examples from one attack configuration. Tested against that same configuration they look excellent, because they learned that attack's artefacts rather than adversarial inputs in general.
go deeper
Should recognise that a detector answers a different question than the classifier and that rejected samples need their own bucket.
Writes the three-way outcome table, names the threshold as the operating point, and pairs detection rate with clean false-positive rate.
Also catches fitting leakage between the attack used to build the detector and the attack used to test it, and re-validates the detector against the full deployed chain.
Frames rejection as a product decision: what the service does with a flagged request, who absorbs the false positives, and whether that cost is acceptable at production traffic.
## What the object actually is A **detector defence** such as ART's `art.defences.detector.evasion.BinaryInputDetector` is not a hardened model. It is a second, ordinary supervised classifier with two classes — benign and adversarial — trained on a labelled corpus that you supply, and wrapped so it can be run over the same inputs as the estimator it protects. - `BinaryInputDetector` learns on the raw input. - `BinaryActivationDetector` learns on a chosen hidden layer's activations of the defended model. Either way the output per sample is a decision plus a score, and the decision comes from comparing that score against a **threshold** you pick. Nothing about that threshold is calibrated for you, and no default is meaningful for your traffic. ## Why plain accuracy stops being valid Accuracy assumes every input receives a label. A detector introduces **abstention**, and an abstained sample is neither correct nor incorrect. If you keep computing accuracy over the estimator's outputs, one of two silent things happens: - the rejected rows drop out of the denominator, inflating the score, or - they keep whatever label the model happened to emit, which measures a system nobody is running. Two different numbers come out of the same run and neither is stated. The fix is mechanical: write the **outcome table** before the run and report every cell — 1. adversarial detected, 2. adversarial missed and then misclassified, 3. adversarial missed but still classified correctly, 4. clean passed and correct, 5. clean passed and wrong, 6. clean flagged. ## The operating point is the result **Detection rate** and **clean false-positive rate** are two coordinates on one curve, moved together by the threshold. A detector that flags everything scores 100 percent detection. So a detection figure without its clean-side twin at the same threshold is not a weak result, it is an uninterpretable one. Report the pair, name the threshold, and ideally show the curve with the chosen **operating point** marked. ## What it costs - **Fitting** is the expensive part, and it is compute you spend before you have measured anything. The adversarial half of the training corpus has to be generated: a 40-step iterative attack over 50,000 training images is roughly two million forward-and-backward passes, which is on the order of an hour on one modern GPU for a small ResNet and days at ImageNet scale. - Then there is the **standing production cost**. A second network forward pass per request is added latency on every call, not just on attacks. - And at a 1 percent clean false-positive rate against ten million requests a day, the detector refuses one hundred thousand legitimate requests every day, forever. ## Where the number misleads - ***Base rate.*** This is the reading that fools the most people. Suppose 1 request in 10,000 is genuinely adversarial, the detector catches 99 percent of them, and its clean false-positive rate is 1 percent. Per million requests that is about 99 true positives against roughly 9,999 false ones: fewer than 1 percent of everything the detector flags is actually an attack. The detection rate was excellent and the alerting queue is still almost entirely noise. A detection rate says nothing about **precision** unless you also state the base rate. - ***Fitting leakage.*** Detectors are usually fitted on adversarial examples from one attack class at one perturbation budget. Evaluated against that same configuration they look superb, because they learned the artefacts of one generator rather than adversarial inputs in general. The honest version fits on one configuration and evaluates on a different one, or holds out an attack class entirely, and reports both configurations by name. - ***Composition.*** A detector validated against the bare model does not necessarily hold once a preprocessing defence is attached in front of it. The preprocessor changes the input distribution, so the features the detector learned may no longer exist. Refit or re-validate against the exact chain you intend to ship. ## What to check - Confirm the threshold is stated and that detection rate and clean false-positive rate are quoted at that same threshold. - Confirm the attack configuration used to fit the detector is different from the one used to test it, and that both are named. - Re-validate the detector against the full deployed chain rather than the bare model. - Estimate the base rate of adversarial traffic and convert the false-positive rate into requests per day. - And ask what the deployed system actually does on a rejection: if it falls back to serving the model's answer anyway, the detection rate is describing a log line, not a defence.
- A detector reports a very high detection rate. What single extra number would most change your reading of it?The false-positive rate on clean data at that same threshold. Flagging everything gives a perfect detection rate, so without the clean-side number the result is unfalsifiable.
- How should a rejected adversarial sample be counted?Normally as a defended success, but only if you also say what the deployed system does on rejection. If rejection falls back to serving the model's answer anyway, it is not a success, and the accounting must say so.
A detection rate without its clean false-positive rate is a smoke alarm that reports how many fires it caught and never how often it went off while you were making toast. Turn the sensitivity up far enough and any alarm catches every fire.
saying these in an interview costs you the question
- Quoting a detection rate with no threshold and no clean false-positive rate.
- Silently dropping rejected samples from the denominator.
- Fitting the detector on the same attack configuration used to evaluate it and calling the result robustness.
- Assuming a detector validated on the bare model still holds once a preprocessor defence is attached in front of it.