skip to content

An attacker writes to one quarterly training export once. How long does an aimed poisoning effect survive?

level: seniorimportance: nice to knowfreq 26%

answer

  1. the effect is refitted, not stored
  2. ask what the next retrain trains on
  3. accumulating table versus rolling window
  4. new honest rows keep arriving nearby
  5. rows times re-buys is the real bill

basics

~20 s

As long as those rows stay in the data each retrain uses, and as long as newly arriving similar rows do not outweigh them. An accumulating table keeps the effect; a rolling window expires it.

solid answer

~50 s

The effect lives in the fitted weights, so it lasts as long as the conditions that reproduce it at each retrain. Two things govern it. First, whether the inserted rows are still in scope: a table that accumulates every export keeps them indefinitely, so one write buys the effect at every later retrain, while a rolling window drops them when the window passes. Second, whether the local balance still favours them: honest encounters resembling the target keep arriving, so the adversary's share of that neighbourhood erodes and a narrow effect can wash out. For a red-team report the cost line is therefore not just how many rows, but how often the write access must be re-bought. Persisting across four retrains from one write is a different finding from expiring next quarter at the same row count.

go deeper

for a junior

Know that a model's behaviour is re-derived at every retrain from whatever data is then in scope, so an inserted row keeps mattering only while it is still part of that data.

for a middle

Explain the two conditions — whether the rows remain in scope, and whether newly arriving similar data outweighs them — and why an accumulating table and a rolling window give completely different answers.

for a senior

Establish how the training set is assembled before claiming any durability, and report row count, access required and retrain intervals survived together, treating a thin-margin intermittent result as a positive finding.

for a principal

Recognise that durability is set by pipeline design rather than by the attack, so the same finding carries a different price depending on how training data is assembled — and that is the number a budget conversation actually turns on.

## The effect is a property of the fit, not a stored object There is no poisoned artefact sitting in the deployed system. What exists is a set of weights that were fitted on a table containing certain rows. Every retrain re-derives the weights from whatever data is then in scope. So the question "how long does it last" is really the question "at each future retrain, are the conditions that produced the effect still present?" Two conditions govern it. ### Condition one: are the rows still in the training data? This is decided by how the training set is assembled, and it is the larger factor. - **An accumulating table.** If each quarterly export is appended and the model is refitted on everything, a single successful write buys the effect at every subsequent retrain, for free. The adversary paid once for an indefinite lease. - **A rolling window.** If training uses only recent data, the inserted rows fall out of scope when the window passes them, and the effect expires without anyone doing anything. Re-buying it means re-obtaining the write access on the same cadence. - **A rebuilt or re-exported source.** If the training table is regenerated from a system of record rather than from accumulated exports, rows that were only ever in the export file may never reappear. None of this is a control anyone deployed against poisoning; it is a property of how the pipeline happens to be built, which is precisely why a red-teamer has to establish it before writing a durability line rather than assuming one. ### Condition two: does the local balance still favour the inserted rows? Even when the rows persist, their leverage is relative. The aimed effect worked because the adversary's rows were a meaningful share of the honest evidence about one small region. New encounters resembling the target keep arriving quarter after quarter, so that share erodes. An effect bought with just enough rows to cross the decision boundary can drift back over it; an effect bought with margin lasts longer. The adversary is trading row count against durability, and that is a real dial they choose on. ### Why this produces findings that reproduce intermittently Retraining is not deterministic in practice — data arrives, initialisation and ordering vary, hyperparameters get retuned. An effect sitting close to the threshold will therefore appear in some refits and not others. For the person triaging the finding, an aimed effect that reproduces in one run out of five is a real result, not a null one: it says the mechanism works and the margin was thin, and buying more margin is simply more rows. Discarding it as flaky is the mistake, because the attacker's response to intermittency is not to give up but to spend a little more. ### The cost line a report has to carry A red-team finding on this is not complete with a row count alone. It needs three numbers together: **how many rows**, **what access was required to insert them**, and **how many retrain intervals the effect survived**. Those multiply into the attacker's real bill and divide into the defender's real exposure. A director reading "forty rows, one contributor's export, persists across every subsequent retrain" is being told something very different from "forty rows, one contributor's export, gone next quarter" — and the difference is not in the attack, it is in how the training set is assembled.

  • A finding reproduces in one refit out of five. Do you report it?
    Yes, and as a positive result. Intermittency here means the effect sat close to the decision threshold, not that the mechanism failed. The attacker's answer to a thin margin is to spend more rows, which is cheap in this variant. Report it with the margin characterised rather than filing it as flaky.
  • How does durability change the cost line you write for an engineering director?
    It converts a row count into an ongoing price. One write against an accumulating table is a one-off purchase of a permanent effect; against a rolling window it is a subscription requiring the same access every cycle. Same forty rows, very different risk, and the difference comes from the pipeline rather than the attack.

saying these in an interview costs you the question

  • Assumes retraining automatically clears a poisoning effect
  • Treats the effect as an artefact stored in the deployed model
  • Reports a row count with no statement of how long the effect lasted
  • Dismisses an intermittently reproducing effect as a flaky result
  • Ignores whether the training set accumulates or rolls

context