skip to content

Degrade or Aim

The split between wanting a model measurably worse and wanting one input read one way, and why only the second is worth planning against in a real deployment.

on this pageshow

explore

questions

4

In training-data poisoning, what separates degrading a whole model from bending one chosen prediction?

level: juniorimportance: must knowfreq 74%

answer

  1. two goals, two very different price tags
  2. one moves a metric, one moves none
  3. an average over the table versus one neighbourhood
  4. percent-scale share against a two-digit row count
  5. the cheap one leaves the dashboard flat

basics

~20 s

The goal, and the price. Degradation wants the model measurably worse and needs a percent-scale share of the training rows. The aimed variant wants one chosen input answered a particular way and can cost only dozens of rows.

solid answer

~50 s

Both start from the same access: the adversary writes rows into data the model will be trained on, and neither needs anything at inference time. They split on goal. An **availability** or degradation goal wants the finished model worse for everybody, and because training minimises an average over the whole table, moving that average means owning a percent-scale share of it. An **integrity** or aimed goal wants one specific input read the way the adversary chose, and that is a local change: a two-digit count of rows placed around that one case is often enough. The consequence matters more than the taxonomy. The degrade variant is loud, expensive and shows up in exactly the metric everyone watches; the aimed variant is cheap and, by construction, moves no aggregate number at all. When an interviewer asks which one you plan against, the answer is the second.

go deeper

for a junior

Be ready to state that poisoning happens in the training data, not at prediction time, and that the goal splits two ways: make the model worse for everyone, or make it answer one chosen input a particular way.

for a middle

Explain why the two goals carry wildly different row budgets — one has to move an average over the whole table, the other only has to win a local decision — and what each one looks like in monitoring.

for a senior

Show that you plan against the cheap variant. Say plainly that a successful aimed insertion leaves aggregate accuracy where it was, and that the absence of a metric change is not evidence of absence.

for a principal

Own the framing that a threat model built only around service degradation is mispriced: the attack that fits real attacker access is the one with no reportable signal, and that changes what assurance you can promise.

## Where poisoning happens Poisoning is an attack on the **training data**, not on an input at prediction time. The adversary needs write access to something the model will learn from: a contributed export, a labelling queue, a crawled corpus, a shared table. Once the rows are in and the model has been fitted, the attack is finished — nothing further is presented at inference. That is the clean separation from evasion, where the adversary changes nothing about training and instead perturbs the input the deployed model is about to read. A concrete setting makes the arithmetic visible. Take a clinical deterioration-risk model over tabular encounter rows, retrained every quarter from periodic exports contributed by several partner hospitals, on a training table of roughly two million rows. The adversary here has one narrow foothold: write access to a single site's export file before it is appended. No query access, no weights, no gradients, and no view of the result beyond the published quarterly performance note. ## The split is on goal, and the goal sets the price **Degradation (an availability goal).** The adversary wants the model to be worse — lower accuracy, more misses, a service that people stop trusting. Training fits parameters that minimise an average loss over every row in the table. To move that average you have to out-vote the honest rows, so your influence is roughly proportional to the share of the table you own. Against two million rows, a drop anybody would notice is a percent-scale insertion: tens of thousands of rows. One site's quarterly export is nowhere near that, and even if it were, the effect lands squarely on the number the owning team reports every quarter. **Aiming (an integrity goal).** The adversary wants one chosen encounter — a specific patient profile, a specific pattern of vitals and labs — scored the way they want, and wants everything else left alone. This is not an average, it is a local decision. In the small neighbourhood of input space around that one case, the honest competition is only the handful of genuinely similar encounters in the table, not two million of them. A two-digit count of rows placed in that neighbourhood can dominate locally. The rest of the model is untouched, so nothing on the dashboard moves. Those two goals are not two intensities of the same attack. They are different buys with different budgets and different signatures, and mixing them up is the most common error on this topic. ## Why the cheap one is the one to plan against The intuition most engineers arrive with is that poisoning means sabotage, and that sabotage would be visible: accuracy falls, somebody notices, the retrain is rolled back. That intuition describes the expensive variant and misses the realistic one. The aimed attack is designed to keep aggregate accuracy exactly where it was — preserving it is part of the objective, because a model that suddenly got worse gets investigated and a model that did not gets shipped. There is no metric on the dashboard that moves. The only thing that changed is the answer on inputs the adversary named in advance. It is also the variant that fits the access an attacker actually has. Owning one contributor's quarterly file is a plausible foothold — a compromised account at a partner site, an insider, a misconfigured upload path. Owning one percent of a multi-site training table is not. ## A note on what "aimed" does and does not mean here The aimed variant discussed here changes the model's answer on an input that already exists and that the adversary picked before inserting the rows. It is worth keeping distinct from a trigger-keyed conditional trained into the weights, which is a different mechanism with a different budget: there, the adversary later presents something they control at prediction time. The distinction that matters for this question is only that both are integrity goals, priced in a small absolute count of rows rather than a fraction of the corpus. It is also worth noting that an aimed insertion does not require obviously wrong labels. Rows that carry entirely correct labels, but that describe cases positioned near the one being targeted, can move a local decision — which is why "we review labels" is a much weaker answer than it sounds. ## How to say it in an interview Name the two goals, price each one in rows, and then state the consequence: the expensive goal is the one your monitoring is built to catch, and the cheap goal is the one it is structurally blind to. If you can add the access story — one contributor's export, one retrain interval — you have given the interviewer a threat model rather than a definition.

  • What does the adversary need at prediction time for either of these goals?
    Nothing. Both are paid for before or during training: the rows go in, the model is fitted, and the changed behaviour is baked into the weights. That is the structural difference from evasion, where the adversary holds no training access at all and instead perturbs the input the deployed model reads. If your threat model only covers what arrives at the endpoint, it does not cover either of these.
  • Which of the two should a threat model for a periodically retrained model treat as the realistic one?
    The aimed variant. It fits the access an attacker plausibly has — one contributor's export, one labelling queue — rather than a percent-scale share of a multi-million-row table, and it produces no signal on the metrics that get reported. Degradation is expensive to buy and loud when it lands, which makes it the benchmark experiment rather than the threat.
  • Does an aimed insertion have to use mislabelled rows?
    No. Rows can carry entirely correct labels and still move a local decision, because what matters is where they sit relative to the case being targeted, not whether a reviewer would call the label wrong. That is why label-quality review is a weak answer to this threat: there is nothing for a reviewer to flag as an error.

Degrading a model is like watering down every bottle in a warehouse; aiming is like tampering with the one bottle you know a particular person will open.

saying these in an interview costs you the question

  • Says poisoning always shows up as a drop in accuracy
  • Treats poisoning and evasion as the same attack
  • Assumes every poisoning attack needs a large share of the data
  • Thinks only obviously wrong labels can poison a training set
  • Believes the adversary must be present at prediction time

context

open as a page

Why does degrading a trained model cost percent-scale poisoned rows while bending one input costs dozens?

level: middleimportance: should knowfreq 56%

basics

~20 s

Degradation has to move an average taken over every training row, so influence scales with the share owned. Bending one input only has to win a local decision, where the honest competition is the few genuinely similar rows.

open as a page

A retrained risk model's accuracy is unchanged quarter over quarter. What does that rule out about poisoning?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Only a degradation campaign large enough to move that particular number, on the slices reported. It rules out nothing about an aimed insertion, whose defining property is that aggregate accuracy stays exactly where it was.

open as a page

An attacker writes to one quarterly training export once. How long does an aimed poisoning effect survive?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

As long as those rows stay in the data each retrain uses, and as long as newly arriving similar rows do not outweigh them. An accumulating table keeps the effect; a rolling window expires it.

open as a page