skip to content

Why does degrading a trained model cost percent-scale poisoned rows while bending one input costs dozens?

level: middleimportance: should knowfreq 56%

answer

  1. ask what the denominator is
  2. an average over everything versus one region
  3. you are out-voting the local handful, not the corpus
  4. how dense is the honest data around the target
  5. cheap per target, linear in targets

basics

~20 s

Degradation has to move an average taken over every training row, so influence scales with the share owned. Bending one input only has to win a local decision, where the honest competition is the few genuinely similar rows.

solid answer

~50 s

Fitting a model minimises an average loss over the whole training table. A degradation goal is therefore a tug-of-war against every honest row at once: your inserted rows pull the fit one way and roughly two million others pull it back, so what you achieve tracks the fraction you own. That is why a noticeable drop is a percent-scale buy. An aimed goal is not an average at all. The model's answer on one particular encounter is determined mostly by what the table says about cases resembling it, and there may only be a few dozen of those. Insert a comparable number of rows describing that region and you can dominate it locally, while every other region is decided by data you never touched. The aimed budget is set by how dense the honest data is around the target, not by how big the table is.

go deeper

for a junior

Know that a model is fitted by minimising an average over all training rows, and that a few rows cannot move an average taken over millions — but can still dominate what the model says about cases like themselves.

for a middle

Be able to name the two denominators out loud: the whole table for a degradation goal, the locally similar rows for an aimed one. That single contrast is the answer this question is looking for.

for a senior

Push it into consequences: the target must be picked in advance, the cost is linear in targets, and rare profiles are the cheap ones — which is awkward when the model exists to flag rare events.

for a principal

Frame the two goals as one axis priced by denominator, so a threat assessment can state what an attacker's realistic access actually buys rather than quoting an attack name with no budget attached.

## Two goals, two completely different arithmetic problems The row budgets differ by three or four orders of magnitude, and it is not a matter of one attack being cleverer. The two goals are quantified against different denominators. ### Degradation is a fight against the whole table Training fits parameters that minimise an average loss over all the rows. Every row gets a vote, and the votes are pooled. If you insert rows that pull the fit in some direction, honest rows describing the same regions pull it back, and the result is roughly a weighted contest in which your weight is your share of the data. Own a tenth of a percent of a two-million-row training table — two thousand rows — and you are outvoted two thousand to one in most of the space. To make the finished model visibly worse across the board, you need a share large enough that the fit is genuinely compromised, which lands at percent scale: tens of thousands of rows against a table that size. There is a second reason this is a bad buy. Degradation is measured by the same aggregate metrics the owning team reports. A campaign large enough to work is a campaign large enough to be seen in the quarterly comparison, and it is one that will be attributed and rolled back. You spend a lot and you get caught. ### Aiming is a fight over one neighbourhood Now ask what determines the model's output on one particular encounter — a specific combination of vitals, labs, age, and history. Overwhelmingly, it is what the training data says about encounters that resemble that one. Rows describing very different patients constrain the model elsewhere; they say little about this corner. So the relevant denominator is not two million. It is however many genuinely similar encounters the table contains, which for a specific and not-very-common profile might be a few dozen. Against a denominator like that, a two-digit count of inserted rows is not a rounding error — it is a comparable force. The adversary is no longer trying to out-vote the corpus; they are trying to out-vote the local handful. That is the entire reason an aimed attack is cheap. Two consequences follow, and they are the ones interviewers probe. **The target has to be chosen first.** The rows are placed relative to a specific case. An adversary who wants an effect on "some input I will pick later" is buying something quite different and paying far more, because they have to influence a region they cannot yet name. **Cheap per target is not cheap in bulk.** Twenty targets means roughly twenty separate local buys, since each has its own neighbourhood and its own local competition. The aimed variant scales linearly in targets, and by the time an adversary wants the model wrong about everything, they are back to paying the degradation price. That is the clean way to see that these are two ends of one axis, not two unrelated attacks. ### Density, not size, sets the aimed price Because the contest is local, the number of rows required is governed by how much honest data already describes the region being attacked. A target sitting in a dense, well-represented part of the distribution is expensive; an unusual profile with few comparable encounters is cheap. This is uncomfortable, because the rare and unusual cases are frequently the ones a risk model exists to catch. ### Labels are not where the leverage lives It is tempting to assume the inserted rows must be obviously wrong — a flipped outcome flag that a reviewer would spot. Flipping labels is one way and it is cheap, but the aimed variant does not require it. Rows can carry correct labels and still shift a local decision, because their leverage comes from where they sit in feature space relative to the target, not from disagreeing with the truth. A review process looking for mislabels has nothing to find, which is why "our labels are checked" answers a different question than the one being asked. ### Saying it well The compact version: degradation is measured against the corpus, aiming is measured against a neighbourhood. Give the two denominators — two million versus a few dozen — and the price difference explains itself, without any need to describe how rows would actually be constructed.

  • Does the adversary need to know which input they want misread before inserting rows?
    For the aimed variant, yes. The rows only carry local leverage because they are placed relative to a case chosen in advance. Wanting the effect on an input they cannot yet name is a much broader objective and prices accordingly — closer to the degradation end, because they must influence regions they cannot point at.
  • What happens to the price if the adversary wants twenty targets instead of one?
    It roughly multiplies. Each target sits in its own neighbourhood with its own local competition, so the buys do not share much. The aimed variant is cheap per target, not cheap in bulk — which is exactly why it and the degradation goal are two ends of one axis rather than unrelated attacks.
  • Are some targets structurally cheaper than others?
    Yes. The cost tracks how much honest data already describes the region around the target. A common, densely represented profile is expensive to bend; an unusual one with only a handful of comparable rows is cheap. Rare cases are the cheap ones, and on a risk model those are often the cases that matter most.

saying these in an interview costs you the question

  • Explains the price difference purely as one attack being more sophisticated
  • Thinks the aimed budget scales with the size of the training table
  • Assumes poisoned rows must carry visibly wrong labels
  • Claims a handful of rows could plausibly move aggregate accuracy
  • Treats twenty targets as costing about the same as one

context