skip to content

You run a data-poisoning attack from an adversarial-robustness library and it hands back arrays of training samples and labels. Why is that not yet a result, and what has to happen before you can say whether the model is vulnerable?

level: juniorimportance: must knowfreq 50%

answer

  1. poisoning call returns data, not a verdict
  2. retrain is the unit of cost
  3. inject, retrain, baseline, two metrics
  4. evasion = queries, poisoning = trainings
  5. did the rows survive the pipeline

basics

~20 s

The library only crafts tainted training data; it does not train anything. You have to mix those rows into the training set, retrain the model with your normal recipe, then evaluate it twice: normal accuracy on clean test data and the attack's success on the triggered inputs. The verdict costs a retrain, not one call.

solid answer

~50 s

A poisoning routine in these libraries is a **data generator**, not an assessment. It returns crafted samples (and, for a backdoor, their attacker-chosen labels) and stops there. Everything that turns that into a finding is yours: inject the rows at the fraction you want to claim, run your real training pipeline end to end, and then measure two things on the resulting model — clean-set accuracy against an unpoisoned baseline, and success on inputs carrying the trigger. That is the structural difference from an evasion run. An evasion attack perturbs an input and asks the wrapped model once, so cost scales with queries or gradient steps. A poisoning attack's unit of cost is a **retrain**, one per configuration you want to report. If you have not retrained, you have crafted data and learned nothing about the model. A common beginner failure is to score the poisoned rows against the *existing* model and call the low accuracy a result.

go deeper

for a junior

Says the call returns crafted training data and that you still have to retrain and evaluate before there is any finding.

for a middle

Adds the injection fraction, the clean baseline, and the two metrics, and can say why the cost unit is a retrain rather than a query.

for a senior

Insists the retrain use the production recipe, checks that the rows survived pipeline filtering, and budgets seeds and sweep points up front.

for a principal

Frames whether the engagement can afford a defensible poisoning claim at all, and what evidence about real ingest control would make the poison fraction plausible.

An adversarial-robustness toolkit — the Adversarial Robustness Toolbox (ART), CleverHans, Foolbox and their relatives — splits along one line: **where in the model's life the attack acts**. Understanding that line answers the whole question. ### Why the library can finish an evasion run but not a poisoning run An *evasion* attack acts at **inference time**. To use one you first build an **estimator wrapper** — in ART, a class such as `KerasClassifier` or `PyTorchClassifier` that adapts your model to a fixed interface exposing `predict` and, in the white-box case, `loss_gradient`. Because the wrapper can be *called*, the library owns the entire loop: an attack object such as ART's `FastGradientMethod` or `ProjectedGradientDescent` perturbs the input, queries the wrapper, checks the label, perturbs again, and `attack.generate(x)` hands back adversarial inputs you can score in the same breath. One call in, a verdict out. A *poisoning* attack acts at **training time**, and no library owns your training loop. Your data lives in your pipeline; your optimiser, schedule, augmentation and filtering are yours. So the API's contract ends where its knowledge ends: at producing data. A call like ART's `PoisoningAttackBackdoor.poison(x, y)` returns a pair of arrays — modified inputs and the attacker-chosen labels to file them under. Nothing has been trained. Nothing has been measured. You are holding an ingredient, not a result. ### What you must add before there is a finding 1. **An injection policy.** How many poisoned rows, expressed as a *fraction of the training corpus*, and — crucially — **where they enter**. Rows appended to an in-memory tensor in a notebook are a much weaker claim than rows that entered upstream and survived whatever dedup, label validation, outlier filtering and augmentation your real pipeline runs. 2. **A retrain.** Same architecture, same hyperparameters, same schedule as the model you intend to make a claim about. Change the recipe and the verdict describes a model nobody ships. 3. **A clean baseline.** The identical recipe trained on unpoisoned data, so "accuracy dropped" has a reference point. 4. **Two metrics.** Clean-set accuracy — *would anyone have noticed?* — and trigger success rate on inputs carrying the trigger — *did it work?* Either alone is half a claim. ### What it costs This is the sentence that separates candidates. **The unit of cost moves from queries to trainings.** An evasion sweep over a thousand test images is minutes to hours and is priced in queries or gradient steps; against a metered endpoint it is priced in fractions of a cent per call. A poisoning sweep is priced in retrains. Five poison fractions, three seeds each, plus a seeded clean baseline is roughly eighteen full trainings. If one production training is six GPU-hours, that is on the order of a hundred GPU-hours and several days of wall clock before anyone writes a sentence — and that is *before* the engineer time, which is usually the larger bill, because plumbing crafted rows into the real ingestion path (rather than a toy loader) is days of work in someone else's codebase. ### Where the number misleads The classic beginner failure is to feed the returned poisoned rows into the **already-trained** model's evaluate call, watch accuracy collapse, and report the attack as successful. That number measures nothing about poisoning. The model's weights never saw the poison; you have only shown that a classifier scores badly on inputs you deliberately relabelled, which would be true of any wrong labels at all. A second misleading reading is a healthy success rate obtained by appending rows *downstream* of the cleaning stages — it overstates risk, because the stages skipped are exactly the ones that might have deleted the poison. A third is an impressive rate obtained at a poison fraction (say thirty percent of the corpus) that no realistic attacker could ever control; the rate is real and the threat model is fiction. ### What I would check Before believing any poisoning result: log **how many poisoned rows actually reached the optimiser** after the pipeline's filtering stages, because cleaning steps drop them silently and a zero-effect result is then a pipeline fact, not a robustness fact. Confirm the clean baseline and the poisoned run differ in *nothing* but the injected data — same seed protocol, same epochs, same augmentation. Confirm the trigger is applied at evaluation exactly as the poisoning routine applied it (placement, scale, and position in the preprocessing order), since a mismatch silently drives success to zero. And state the poison fraction next to every rate you quote.

  • Why can an evasion attack return a verdict from one call while a poisoning attack cannot?
    Evasion acts at inference, so the library can perturb and query the wrapped model itself. Poisoning acts at training time, and the library does not own your training loop — it can only hand you data.
  • You injected poisoned rows and the trained model shows no backdoor. Name two non-security explanations before you conclude the model is robust.
    The rows may have been dropped by pipeline steps such as dedup, filtering or label validation, or the injected fraction may be below the rate at which the effect appears. Log the count of poisoned rows that actually reached the optimiser and sweep the fraction.
  • What does it cost to add one more configuration to a poisoning assessment?
    One full retrain per configuration, plus repeats if you need to separate the effect from seed-to-seed variance.

The library sells you flour, not bread: it hands over crafted training data and stops. The oven is yours, and every configuration you want to report costs another bake.

saying these in an interview costs you the question

  • Believing the library trains a backdoored model for you and reports success
  • Scoring the poisoned rows against the already-trained model and calling that the attack result
  • Reporting a drop in accuracy with no clean-data baseline trained under the same recipe
  • Injecting poison into a toy training script that shares nothing with the production pipeline, then presenting the number as a production risk

context