What does constraining injected training records to look in-distribution cost the attacker who writes them?
answer
- the record trades effect for survival
- influence per record versus record count
- the budget changes units
- corpus size dilutes one goal, not the other
- a price only binds if they cannot pay
basics
~20 sLeverage per record. A record forced to look ordinary cannot be extreme, and extremeness is what moved the fit, so the same effect needs many more written records. Sanitization sets that exchange rate rather than closing the attack.
solid answer
~50 sThink of a poisoned record as buying a small amount of influence over where the trained boundary lands. An unconstrained record can be extreme and buys a lot. Once a screen is running, the attacker must keep each record inside the observed distribution, so each one buys much less, and the shortfall is made up in volume: more records written, over more capture windows. That is the whole effect of sanitization - it converts an exclusion into a price, denominated in write volume. Two things follow. First, the price is only a control if it exceeds what someone with write access can actually produce; on a corpus fed automatically from a monitored source, that ceiling can be high. Second, the price is not the same for every goal: a general degradation attack competes against the whole corpus and dilutes as the corpus grows, while a conditional behaviour keyed to inputs the attacker chooses behaves as a roughly absolute record count and barely dilutes at all.
go deeper
Know the basic trade: a record made to look ordinary survives the screen but moves the model less. Be able to say the attacker compensates with more records.
Explain the exchange rate in both directions - what per-record leverage buys and what volume replaces it - and separate a degradation goal, which dilutes with corpus size, from a narrow conditional one, which does not.
Be ready to say how you would measure the multiplier rather than assert it, and to tie it to how much an attacker with write access to your corpus can realistically produce.
Own the conclusion the exchange rate forces: if the price does not exceed what a writer can pay, the money belongs on rationing write access rather than on tuning scoring rules.
## Reframing the filter as an exchange rate The useful mental model is not *does the screen catch the poison* but *what does the screen charge for*. A sanitization stage that drops records scoring high on unusualness imposes a constraint on any attacker who wants their records to reach training: whatever else those records do, they must look like plausible members of the corpus. That constraint has a measurable price, and the price is the answer to this question. ## Why looking ordinary reduces per-record leverage A fitted model is a compromise across the records it saw. How much any single record can pull that compromise depends on how far it sits from the rest and how the loss treats it. A record with extreme values, or a label that contradicts everything nearby, pulls hard - which is precisely why careless poison works with few records, and also why it is easy to spot. Force that record to sit inside the observed distribution and both properties change at once. It no longer stands out, so the screen keeps it; but it also no longer pulls hard, because it now resembles the records already pulling in the ordinary direction. The attacker has traded effect for survival. The shortfall is made up the only way it can be: more records. So the poison budget stops being *how extreme may my records be* and becomes *how many records can I get written*. On a network-intrusion model retrained on captures from the segment it monitors, that second budget is the interesting one, because the attacker's vantage is simply the ability to emit traffic that lands in the next capture. There is no query limit and no per-record cost beyond the traffic itself. Whether the raised bill is a real defence depends entirely on how many records that attacker can produce over how many retraining windows, and that is a number the defender is supposed to estimate rather than assume. ## The dilution asymmetry, which is where candidates go wrong A very common answer is *our corpus is enormous, so a handful of records cannot matter*. That is true for one kind of goal and false for the other. - **Degradation of general accuracy** works by shifting the overall fit, so it competes against every honest record in the corpus. Its required volume genuinely scales with corpus size, and a growing corpus dilutes it. - **A conditional behaviour keyed to inputs the attacker chooses** does not compete against the corpus. It asks the model to learn one additional, narrow association that nothing else in the data contradicts. The number of records that takes behaves as roughly an absolute count rather than a fraction, so a corpus ten times larger dilutes it very little. In-distribution constraint raises the price of both, but it raises them from different starting points, and a defender quoting corpus size as protection is implicitly assuming the first goal. ## What the filter does and does not bound The defensible statement after a sanitization stage is a conditional one: | The screen bounds | The screen does not bound | | --- | --- | | Records lying outside the observed distribution | Records constructed to lie inside it | | The cheapest form of the attack | The attack's existence at any budget | | How extreme a single record may be | How many ordinary-looking records may be written | Notice the shape: every row on the left is about the records, and every row on the right is about the attacker. That is the gap the question is aiming at. A screen is a property of your data pipeline; a budget is a property of whoever can write into it, and only the second determines whether the raised price binds. ## Where the raised price actually bites The price is a genuine defence in exactly one situation: when the write volume the attacker needs exceeds what they can produce without being noticed by some other mechanism. Two cases are worth separating. If write access is rationed - a small number of contributing sources, a per-source volume that is itself monitored, a corpus that only accepts records through a channel someone owns - then multiplying the required record count by a large factor can genuinely take the attack out of reach. If write access is effectively unmetered, as it is on a segment anyone can emit onto, multiplying the count by a large factor changes how long the attacker spends and nothing else. This is why the sanitization discussion so often ends up somewhere else: the filter's output is a price, and the interesting engineering question is who is able to pay it. That is a question about write access and volume, not about scoring rules. ## Answering it in an interview Say the cost plainly - per-record leverage - then say what it is exchanged for - volume - then name the two conditions that decide whether the exchange helps: how much the attacker can write, and which goal they hold. An answer that stops at *the poison gets weaker* has given half of it.
- How would you actually estimate the multiplier a screen imposes, rather than asserting one?Empirically and with the constraint applied: build the poison twice, once unconstrained and once restricted to stay inside the observed distribution, and measure how many records each version needs to reach the same effect on the trained model. The ratio is the multiplier. Quoting it without measurement is an assertion, and quoting the screen's removal rate instead measures something different entirely.
- Does adding a second screen with a different scoring rule multiply the price again?Only if the two screens disagree about which records are unusual. Distance-based and influence-based scores frequently rank the same records highly, so a second stage often re-cuts the same tail and adds little. The gain comes from screens whose failure modes differ, and the honest way to know is to measure how much of the second stage's removals the first had already taken.
- Why does corpus growth reassure a team more than it should?Because they are picturing an attacker trying to move overall accuracy, which does dilute with size. A conditional behaviour keyed to attacker-chosen inputs needs roughly a fixed number of consistent records, and nothing else in a bigger corpus contradicts that narrow association, so the fraction falls while the required count stays put.
saying these in an interview costs you the question
- Says in-distribution poison is simply weaker and stops there
- Claims a large corpus dilutes every poisoning goal
- Reports removal rate as the screen's effectiveness measure
- Ignores how much the attacker can write
- Assumes stacking screens multiplies the price