You retrained a classifier on data produced by a backdoor-poisoning routine in an adversarial-robustness library. Which two numbers do you report, and why is the trigger success rate on its own a misleading result?
answer
- two numbers: stealth and effect
- clean accuracy vs unpoisoned baseline
- exclude target-class samples from the denominator
- always state the poison fraction
- trigger applied the same way at eval
basics
~20 sReport two: accuracy on a clean, untriggered test set against the unpoisoned baseline, and the trigger success rate measured only on samples not already in the target class. Success rate alone hides a poisoned model that lost obvious accuracy, and it is inflated by inputs the clean model already sent to the target.
solid answer
~50 sA backdoor claim has two halves and one number covers only one of them. **Clean accuracy** answers *would anyone have noticed?* Measure the retrained model on an untriggered held-out set and compare it with the same recipe trained on clean data. A backdoor that costs several points of accuracy would be caught by ordinary model validation before it ever shipped, so the success rate is describing an attack that does not survive release. **Trigger success rate** answers *does it work?* — the fraction of triggered inputs routed to the attacker's target class. The trap is the denominator. Samples whose true class is already the target count as successes with or without the trigger, so leaving them in inflates the rate by roughly the target class's share of the set. Exclude them. Both numbers need the baseline and the seed context to mean anything; a two-point accuracy gap can be run-to-run variance rather than the poison.
go deeper
Names the two numbers: accuracy on clean data and success on triggered data.
Explains why target-class samples must leave the denominator and why the poison fraction has to be quoted with the rate.
Runs the full two-by-two of model and test set, checks trigger application matches between poisoning and evaluation, and reports variance across seeds.
Decides which of these numbers the risk owner can act on, and insists the poison fraction be tied to evidence about who can influence real training data.
The instrument gives you poisoned rows. The **measurement design is entirely yours**, and it is where these assessments most often go wrong — usually in a way that flatters the attack. ### The two numbers, and what each answers **Clean accuracy** answers *would anyone have noticed?* Evaluate the retrained model on an untriggered held-out set and compare it against the same recipe trained on clean data. A backdoor that costs several points of top-line accuracy would be caught by ordinary model validation before release — so a headline success rate attached to a visibly degraded model describes an attack that never reaches production. Stealth is not a nice-to-have half of the claim; it is what makes the other half matter. **Trigger success rate** answers *does it work?* — the fraction of trigger-stamped inputs the model routes to the attacker's chosen target class. ### The denominator trap The success rate's numerator is easy; the denominator is where the number lies. Samples whose *true* class is already the attacker's target land in the target class on a clean model too. Leave them in the triggered test set and they count as successes with or without the trigger, inflating the reported rate by roughly that class's share of the set — about ten points on a balanced ten-class problem, and far more on an imbalanced one where the target is the majority class. **Exclude every sample whose ground-truth label is the target class** from the triggered evaluation, and state in the report which convention you used, because an unqualified success rate is not comparable to anybody else's. ### Run four evaluations, not two The minimal honest design is a two-by-two of {clean baseline model, poisoned model} against {clean test set, triggered test set}: | | clean test set | triggered test set | |---|---|---| | **clean baseline model** | the accuracy reference | control: does the trigger already misroute? | | **poisoned model** | the stealth number | the raw success number | The top-right cell is the one people skip and the one that saves you. If the clean baseline model *already* sends triggered inputs to the target class at some rate, part of your success number is the trigger pattern being an out-of-distribution artefact that any model handles oddly — not the backdoor. The effect attributable to the poison is the bottom-right minus the top-right, not the bottom-right alone. ### State the poison fraction, always A success rate is meaningless without the share of the training corpus you had to control to buy it. "Ninety-plus percent success at one percent poisoning" and the same rate at thirty percent are different findings, and only one of them describes an attacker anybody should plan against. The fraction is the bridge between a lab number and a threat model, and it must be a share of the **real** corpus, not of the subsample you experimented on. ### What it costs Every cell of that two-by-two above the diagonal is cheap — evaluation is minutes — but the *models* in it are not. You need at minimum two trained models (baseline and poisoned) under an identical recipe, and if you intend to claim a stealth cost you need several seeds of **both**, because a two-point accuracy gap is routinely within seed-to-seed variance for the same recipe on the same data. Four seeds of each is eight production trainings for one quotable configuration; at six GPU-hours per training that is roughly two GPU-days for a single row of the results table. Budget it before you promise the row. ### Where the number misleads Three readings, in descending order of how often they appear. First, a lone success rate with no clean-accuracy comparison — an attack that works and is obvious is not the finding it looks like. Second, the inflated denominator described above, which silently adds the target class's prior to every rate you quote. Third, a small clean-accuracy gap presented as *the stealth cost of the backdoor* when it is single-seed noise; the honest phrasing is "no accuracy cost was detected at the seeds we ran", which is a bounded, defensible statement. ### What I would check That the trigger is applied at evaluation exactly as the poisoning routine applied it during generation — same placement, same scale, same position relative to normalisation and augmentation — because a mismatch anywhere in that chain drives success to near zero and looks like robustness. That the clean and poisoned runs differ in nothing but injected data. That the target-class exclusion was actually applied to the triggered set, by checking the row count dropped by about the class share. And that the poison fraction quoted is a share of the production corpus.
- Why also evaluate the clean baseline model on triggered inputs?To show the trigger does nothing on an unpoisoned model. If the baseline already misroutes triggered inputs, part of your success rate is the trigger being an out-of-distribution artefact, not the backdoor.
- Your poisoned model is 0.4 points below the clean baseline. Is that the stealth cost of the backdoor?Not on one run. Retrain both configurations across several seeds; if the gap sits inside the seed-to-seed spread you can only say no cost was detected at the budget you ran.
Leaving target-class samples in the triggered denominator is like grading a doctor who diagnoses everyone with the flu on a ward where a tenth of patients already have flu: a tenth of the credit is the ward's composition, not the diagnosis.
saying these in an interview costs you the question
- Reporting a single success rate with no clean-accuracy comparison
- Leaving target-class samples in the triggered denominator
- Quoting a success rate without the poison fraction that bought it
- Treating a small clean-accuracy gap as a real cost without checking seed variance
- Never checking that the clean baseline model does not already misroute triggered inputs