A vendor deck claims a poison-resistant training pipeline. What do you require before crediting the claim?
answer
- removal rate is the wrong quantity
- ask what survived, not what left
- every threshold has a bill
- a multiplier needs someone unable to pay
- the register entry should name a price
basics
~10 sRequire the claim be restated as a price: how many extra records an attacker constrained to look ordinary needs, what the threshold cost on rare real data, and who can write into the corpus.
solid answer
~50 sA removal rate against injected records is the wrong quantity, and it is usually the only one on the slide. I would ask three things. What did the surviving records do to the trained model - residual effect, not fraction removed, because a screen that removes 98% of a poison set that only needed 2% has achieved nothing. What did the threshold cost on the rarest real classes, since that is where a tighter cut lands and where aggregate accuracy hides it. And who can write into this corpus and how much, because a screen's real output is a multiplier on the attacker's record count, and a multiplier only defends you if it exceeds what a writer can produce. Then the decision I actually own: whether to fund threshold work at all, or to spend the same effort rationing and attributing writes and keeping an evaluation set the corpus cannot influence.
code
text · 12 linesTraining-data hardening (vendor deck)
outlier screen drops top 1.5% by unusualness score
influence screen drops top 0.5% by effect on the fit
injected test set 500 records, well outside the data range
removal rate 98.6%
clean accuracy unchanged (within 0.1 pt)
...
no line for: what the 1.4% that survived did to the model
no line for: recall on the rarest real classes after the drop
no line for: records needed by poison built to stay in rangego deeper
Know that a removal rate describes the poison someone chose to inject, and does not show that the pipeline resists poisoning in general.
Be able to say which quantities are missing from such a claim: the effect of the records that survived, and what the threshold cost on rare real data.
Show that you would ask for the residual model effect and rare-class recall, measured on data the corpus could not influence, before crediting any pipeline claim.
Own the allocation: screens buy a multiplier with diminishing returns, so decide whether the money goes to thresholds, to rationing and attributing writes, or to independent measurement - and be willing to accept a named residual instead of claiming a control.
## What the slide is claiming, and what it can support *Poison-resistant data pipeline* is a categorical claim - a class of attack does not work here. The evidence offered for it is almost always a removal rate: poison records were injected, the screens dropped most of them. Those two things do not connect. Removal rate is a property of the injected set, and the injected set was chosen by whoever wrote the slide. The reviewer's job is to convert the categorical claim into the quantitative one it could honestly support, and then decide whether that quantity matters for the deployment in front of them. ## The three questions that do the work **1. What did the survivors do?** A screen that removes 98.6% of an injected set has left 1.4% in the training data. Whether that matters depends entirely on how many records the effect required. If the goal was a narrow conditional behaviour needing a roughly fixed record count, the survivors may be sufficient on their own, and the removal rate is then a large number attached to a failed defence. The measurable quantity is the retrained model's behaviour on whatever the poison targeted, with the screens running. Nobody puts that on a slide because it is a harder experiment and a less flattering number. **2. What did the threshold cost?** Every removal rate has a threshold behind it, and every threshold deletes real records from the tail before it reaches poison written to look ordinary. Ask for recall on the rarest real classes, measured on an evaluation set the training corpus could not influence, before and after the screens. If the deck reports unchanged aggregate accuracy, that is consistent with having deleted most of the tail, because the tail is a small share of any sample drawn from the same distribution. **3. Who can write, and how much?** This is the question that decides the others. The genuine effect of sanitization is a multiplier: an attacker constrained to stay inside the distribution needs materially more written records for the same result. A multiplier defends you only when the required count exceeds what a writer can actually produce. For a corpus assembled from a source anyone can emit into, the ceiling is high and the multiplier buys time rather than safety. For a corpus fed by a small, attributed set of contributors with metered volume, the same multiplier can genuinely put the attack out of reach. The screen is identical in both cases; its worth is not. ## The decision that is actually yours Having priced the claim, a lead has to allocate. The options are not *better screens* versus *nothing*. - **Fund threshold and scoring work.** Cheap, visible, and subject to diminishing returns, because a second screen usually re-cuts the same tail the first one did. Its ceiling is the multiplier, and it costs real-data coverage on the way up. - **Fund write-side control.** Ration who can contribute to the corpus, attribute records to a source, cap per-source volume per window, retain provenance so a later removal can be traced. This attacks the quantity the multiplier is measured against, and it is the only lever that turns a price into a bound. - **Fund independent measurement.** Keep a held-out evaluation set the corpus cannot influence, weighted toward the rare classes the model exists to catch, so that both the poison effect and the screen's own cost become observable at all. Without this, neither of the first two options can be assessed. - **Accept the residual and say so.** Sometimes the honest answer is that this corpus is fed from a source you do not control, the multiplier is what you get, and the residual risk is accepted with a named owner and a monitoring plan. That is a defensible position. *Poison-resistant* is not. ## What you tell the vendor, and what you tell your own leadership To the vendor: restate the claim. *Our screens remove poison lying outside the observed distribution; against poison constrained to sit inside it we measure a multiplier of roughly N on the record count required, at a cost of M points of recall on the rarest classes.* Every term there is checkable, and a vendor who cannot produce any of them has measured removal rate and nothing else. To your own leadership: the sentence *the pipeline filters poisoned data* should not appear in a risk register, because it describes a control that does not exist. What exists is a cost imposed on an attacker whose size you can estimate, and a decision about whether the people who can write into this corpus can pay it. Write the register entry that way and the follow-on work chooses itself. ## The failure mode this question is testing The weak answer credits the claim because the removal rate was high, or rejects it because vendors exaggerate. Both skip the reasoning. The strong answer names the quantity that was measured, names the quantities that were not, and ends on an allocation decision - which is what makes it a lead's question rather than an analyst's.
- The vendor offers to raise the drop rate to satisfy you. Is that progress?Not by itself, and it may be a loss. A higher drop rate removes more of the tail, and the tail is disproportionately the rare real behaviour the model exists to catch. I would treat the offer as a request for a measurement: show recall on the rarest classes at both thresholds, and show what the surviving poison did to the model at each. If neither number moves in my favour, the higher rate is a worse pipeline reported as a better one.
- How should this appear in a risk register?As a priced residual, not a mitigated risk. Something like: poisoning of the training corpus is possible for anyone with write access to the collection source; screening raises the required record volume by an estimated factor; the residual is accepted, owned by a named person, with monitoring on per-source write volume. The line 'the pipeline filters poisoned data' describes a control that does not exist and should be challenged wherever it appears.
- Where would you spend first if the corpus is fed from a source you do not control?On independent measurement and on write-side metering, in that order. An evaluation set the corpus cannot influence is what makes any of the other claims checkable. After that, attributing records to a source and capping per-source volume attacks the quantity the screen's multiplier is measured against - which is the only lever that can turn a price into something an attacker cannot pay.
saying these in an interview costs you the question
- Accepts a high removal rate as proof of resistance
- Never asks what the surviving records did
- Ignores the threshold's cost on rare real classes
- Calls the screen a mitigation in the risk register
- Assumes more screens multiply the protection