skip to content

What kind of data poisoning does a model's aggregate accuracy actually detect?

level: middleimportance: should knowfreq 46%

answer

  1. split on goal, not on technique
  2. one wants worse, one wants specific
  3. what share of traffic changes?
  4. fraction versus absolute count
  5. dilution only defends against one

basics

~20 s

Aggregate accuracy detects a degradation attack, which needs a poisoned fraction that grows with the corpus and whose success is a moved number. It is blind to a targeted behaviour, which needs roughly an absolute row count.

solid answer

~50 s

Poisoning splits on **goal**, and the two goals behave completely differently against a dashboard. An availability or degradation attack wants the model measurably worse for everyone; success *is* a moved number, so the metric you already watch is a genuine detector for it. It is also genuinely diluted by corpus size — it needs a fraction of the data, so a ten-times-larger corpus costs the attacker roughly ten times more writes. A targeted attack wants one behaviour on inputs the adversary chooses. That behaviour touches a vanishing share of ordinary traffic, so it cannot move an average even when it works perfectly; it behaves approximately as an absolute sample count rather than a fraction, so growing the corpus barely dilutes it. The practical consequence is uncomfortable: the metric is a fine detector for the loud, expensive-at-scale attack, and a null detector for the cheap, narrow one.

go deeper

for a junior

Know that poisoning comes in two flavours by goal: make the model worse for everybody, or make it behave a chosen way on chosen inputs. The second is the one metrics miss.

for a middle

Be ready to explain the scaling difference — effect set by a poisoned fraction versus roughly an absolute row count — and why that decides whether a large corpus is any defence at all.

for a senior

Expect to translate this into what your monitoring is worth: name which attack class your existing metrics genuinely cover, and stop describing coverage of one class as coverage of poisoning.

for a principal

The strategic call is where to spend: dilution and corpus scale buy down one goal only, and buying anything against the other means funding a directed search with its own limited scope. Say which you are buying and which you are accepting.

## Two goals, not one attack "Data poisoning" names a write into a training corpus, but the write can be aimed at two quite different outcomes, and everything about how it interacts with monitoring follows from which one it is. **Availability / degradation.** The adversary wants the trained model to be measurably worse — noisier, less accurate, less useful — across ordinary inputs. Cheap label flipping is the classic instance: rows whose labels are wrong, which drag the fitted boundary around. **Targeted.** The adversary wants a specific behaviour on specific inputs: one chosen input read the wrong way, or a conditional behaviour that fires on a pattern they control. Everything else about the model is to stay exactly as it was. ## Why one is visible in the aggregate and the other is not Aggregate accuracy is an average over a sample of ordinary traffic. Whether an attack shows up in it is a question about *what share of that traffic the attack changes*. A degradation attack, by definition, changes behaviour on a broad swathe of ordinary inputs — that is the goal. So it moves the average, and the number you already watch is a real detector for it. A targeted attack changes behaviour on inputs the adversary chose. Those inputs are, by construction, either a single record or a family defined by a pattern that does not occur in natural traffic. Their share of a validation sample is somewhere between negligible and zero. A behaviour that is wrong on 100% of a set that is 0.001% of your traffic moves your headline number by 0.001 points — inside the run-to-run variation of retraining the same model on the same data. ## The scaling difference, which is the part that surprises people The two goals also price differently as data grows, and the direction is counterintuitive. - Degradation is **diluted**. Its effect is roughly set by the poisoned *fraction*. Double the honest corpus and the adversary must roughly double their writes to hold the same effect. "We have a lot of data" is a real, if partial, defence here. - A targeted or conditional behaviour is **approximately absolute**. What it needs is enough examples of the association for the model to fit it, and that requirement does not grow much when you add unrelated honest rows around it. Doubling the corpus does not double the attacker's bill; it may barely move it. That is why the intuition "our dataset is enormous, poisoning is not a realistic threat for us" holds for one goal and collapses for the other. Corpus size is a dilution defence, and dilution only works against an attack whose effect is set by a fraction. ## What this means for the dashboard Put plainly, the monitoring you have is well matched to the attack that is *loudest and most expensive at scale*, and badly matched to the attack that is *quietest and cheapest at scale*. There is no adjustment to the aggregate that fixes this, because the problem is not sensitivity — a more sensitive average is still an average over a distribution that does not contain the attacker's inputs. So the honest reading of a flat curve is asymmetric, and it is worth saying in exactly these terms: | Goal | Effect scales with | Visible in aggregate quality? | | --- | --- | --- | | Degradation of the model for everyone | poisoned fraction of the corpus | yes — this is what the metric is good at | | A behaviour on attacker-chosen inputs | roughly an absolute count of rows | no — the inputs are not in the sample | ## The follow-on an interviewer usually asks Having established that the aggregate cannot see the narrow attack, the natural next question is what could. The answer is not "a better average." It is a search directed at the thing you are worried about — evaluating behaviour on inputs the ordinary distribution never produces — which is a different exercise from monitoring, costs different money, and bounds only the shapes it actually searched for. Recognising that monitoring and searching are not the same activity is most of the point of the question. ## Common wrong answers - "Poisoning shows up as an unstable loss curve." A curve wobbles for a hundred ordinary reasons and a careful write produces none of them; there is no poisoning signature in the shape of a loss trace. - "Both goals are diluted by corpus size." Only one is, and getting this backwards is the error that makes large-corpus teams complacent. - "A targeted attack is just a small degradation attack." It is a different objective, not a weaker one, and it is often *more* reliable at achieving what its author wanted than a degradation attempt is.

  • Your corpus grew tenfold this year. Which poisoning goal did that make harder, and by roughly how much?
    It made degradation about ten times more expensive, because the effect is set by the poisoned fraction, so holding the same fraction means ten times the writes. It made a targeted or conditional behaviour barely harder at all, since what that needs is enough examples of one association, which does not grow when you add unrelated honest rows.
  • Is a per-class report enough to catch the narrow case?
    No. A per-class report is still an average, just over smaller buckets, and the attacker's chosen inputs are not in any bucket in proportion. It bounds movement in the classes it lists, at the resolution their sample sizes support, and it says nothing about behaviour on inputs the ordinary distribution never produces.
  • If degradation is the visible one, why would anyone run it?
    Because the goal is sometimes the disruption itself rather than stealth — a competitor's model made unreliable, a service degraded, a retraining cycle poisoned enough to force a rollback. Visibility is only a cost if the attacker needs to persist.

saying these in an interview costs you the question

  • Poisoning always destabilises the training curve
  • A big corpus dilutes every kind of poisoning
  • Targeted poisoning is just weak degradation poisoning
  • Per-class metrics cover attacker-chosen inputs
  • Any successful poisoning must cost some clean accuracy

context