skip to content

Sanitization and Its Limits

Filtering the training set raises what an attack costs instead of ending it, and turning the filter up far enough starts deleting the real tail you most wanted to learn from.

on this pageshow

explore

questions

4

Why doesn't dropping outliers from a training capture stop an attacker who knows that filter runs?

level: juniorimportance: must knowfreq 62%

answer

  1. the filter scores one thing only
  2. unusual is not the same as inserted
  3. the attacker read the data sheet too
  4. rare real records live in that same tail
  5. it buys a price, not an exclusion

basics

~20 s

An outlier filter removes records that look unusual, not records placed on purpose. An attacker who knows it runs keeps injected records inside the data's normal range, so they score as ordinary. Filtering raises the attack's cost.

solid answer

~50 s

An outlier screen scores a record by how far it sits from the rest of the corpus, so it catches poison that was written carelessly - records with impossible values, or labels that contradict everything nearby. It does not catch poison chosen with the screen in mind. An attacker who can write records into the corpus and has read that the screen runs first simply constrains what they write to sit inside the observed distribution: every injected record then looks like a plausible member of the data. The filter is a distance test, not an intent test, and the two populations - rare-but-real records and in-distribution poison - overlap. Turning the threshold up until it reaches the poison starts by deleting your own rare real records, which are usually the hardest and most valuable ones you have. The honest claim is that sanitization prices the attack, not that it prevents it.

go deeper

for a junior

Be ready to say plainly that an outlier screen scores how unusual a record looks, and that an attacker who knows it runs writes records that look ordinary. Do not claim the filter removes poisoning.

for a middle

Explain why the two populations overlap on the unusualness axis, and what constraining poison to look ordinary costs the attacker in extra written records. Name the threshold tradeoff explicitly.

for a senior

Show that you would state the screen as a price rather than a control, and that you would measure what tightening it costs on your rare real data before recommending it.

for a principal

Own the framing question: what is a filter whose real output is an exchange rate worth in a threat model, and is it cheaper to raise the attacker's bill or to limit who can write into the corpus at all?

## The claim being tested Almost every team that trains on data it did not fully control says some version of the same sentence: *we drop outliers before training, so poisoned samples get filtered out.* It sounds like a control. It is really a claim about a correlation - that records an attacker inserts will look statistically unusual - and that correlation only holds for an attacker who was not thinking about the filter. Make the setting concrete. A network-intrusion model is retrained periodically on connection records captured from the segment it monitors: source and destination, port, byte counts, duration, flag patterns. Nobody hand-labels a million flows, so the corpus is whatever the segment produced. Anyone who can emit traffic on that segment can write into the next training set. Suppose the pipeline is documented publicly - a data sheet, a conference talk, a vendor deck - and it says an outlier score is computed over each capture and the top slice is dropped before training. ## What the filter actually measures A sanitization screen assigns each record a number that means *how unlike the rest of this corpus is this record*, then cuts at a threshold. Some variants score distance or density in feature space; others score how much a record moves the fit, or which cluster it falls into. Whatever the scoring rule, the output is one axis: unusualness. The filter has no access to who wrote a record, why it is there, or what it will do to the trained model. It cannot distinguish an attacker's record from a rare real one, because that distinction is not encoded in the data. So the filter partitions the corpus into two populations that were never the populations of interest: - **Unusual records**, which are a mixture of measurement errors, genuinely rare real events, and careless poison. - **Ordinary-looking records**, which are a mixture of the bulk of real data and any poison written to look ordinary. Dropping the first bucket removes careless poison and, as a side effect, removes real data from the tail. It removes nothing from the second bucket by construction. ## The attacker's answer, and what it costs them An attacker who knows the screen runs treats *stay inside the distribution* as a constraint on what they may write, alongside whatever effect they want. On the monitored segment they emit flows whose ports, byte counts, durations and timing sit comfortably within what that segment normally produces, and let those flows be captured and labelled by the same automatic process that labels everything else. Nothing about a single such record reads as anomalous, because by construction it is not anomalous. The constraint is not free. A record forced to look ordinary is a record that cannot be extreme, and extremeness is exactly what gave a careless poisoned record its leverage over the fit. Each in-distribution record therefore moves the trained model less than an unconstrained one would. To reach the same effect the attacker needs more records. That is the real output of the filter: **an exchange rate**, not an exclusion. Sanitization converts a cheap attack into an expensive one and states its price in records the attacker has to get written. That is a useful thing to buy. It is a different thing from what the original sentence claimed. ## Why you cannot simply turn the threshold up The obvious response is to cut deeper - drop the top five per cent instead of the top one. The problem is that the two populations you care about are not separated on the unusualness axis. Tightening the cut removes in-distribution poison only after it has removed the real records nearest the boundary, and on a network-intrusion corpus those rare real records are frequently the very traffic the model exists to recognise: the uncommon protocol, the one legitimate bulk transfer a month, the low-and-slow pattern. A filter tuned to keep recall on real data keeps the poison; a filter tuned to catch the poison deletes the data you most needed. There is no threshold that avoids both, because the filter is scoring the wrong quantity. There is a second-order cost worth naming: pushing a defender to tighten the screen is itself a payoff for an attacker whose goal is a degraded monitor. ## How to state it honestly A sanitization stage is worth running. It removes the low-effort case, it catches genuine data-quality problems, and it forces any attacker who wants effect to spend more. What it does not do is bound the possibility of poisoning, so it cannot be cited as the reason the corpus is trustworthy. The defensible claim is quantitative and conditional: *this screen removes poison that sits outside the observed distribution, and against poison constrained to sit inside it, our estimate is that the attacker needs on the order of N times more written records for the same effect - which matters only if N exceeds what someone with write access to this corpus can actually emit.*

  • If the screen still helps, what is the defensible way to write it up as a control?
    Write it as a cost, with a number. Say that it removes poison lying outside the observed distribution, and that poison constrained to look ordinary needs materially more written records for the same effect. Then compare that record count against how much someone with write access to the corpus can realistically produce. A control stated as a price can be checked; a control stated as an exclusion cannot.
  • Does it matter whether the injected records carry correct or incorrect labels?
    It matters for review, not for this screen. A distance or density score looks at the feature values, so a record with a flipped label but ordinary features can pass it, while a human sampling the corpus might catch the contradiction. Correctly labelled records placed to shift a boundary give a reviewer nothing to flag at all. Either way, staying inside the distribution is what defeats the outlier cut.
  • The pipeline description was published. Would keeping it private have helped?
    Marginally and temporarily. An attacker who cannot read the description can usually infer that some outlier screen runs, because nearly every pipeline has one, and staying inside the observed distribution is a safe default that costs them little to assume. Secrecy about the threshold buys some uncertainty about how far they can push, but it is a delay, not a control, and it makes your own evaluation dishonest if you count on it.

A doorman who turns away anyone dressed strangely stops the obvious gatecrasher and nobody else. Someone who checked the dress code first walks in, and tightening the code starts turning away your own guests.

saying these in an interview costs you the question

  • Says outlier removal filters out poisoned samples
  • Assumes injected records must look statistically unusual
  • Treats a higher drop rate as strictly safer
  • Confuses an unusualness score with an intent test
  • Cites the filter as proof the corpus is trustworthy

context

open as a page

What does constraining injected training records to look in-distribution cost the attacker who writes them?

level: middleimportance: should knowfreq 44%

basics

~20 s

Leverage per record. A record forced to look ordinary cannot be extreme, and extremeness is what moved the fit, so the same effect needs many more written records. Sanitization sets that exchange rate rather than closing the attack.

open as a page

How do you tell whether tightening a training-data outlier screen removed poison or your rare real records?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Not from the screen's own output - dropped and kept records are both mixtures. Inspect what was dropped, and measure what the tighter cut cost on the rare real behaviour the model exists to catch.

open as a page

A vendor deck claims a poison-resistant training pipeline. What do you require before crediting the claim?

level: principalimportance: nice to knowfreq 26%

basics

~10 s

Require the claim be restated as a price: how many extra records an attacker constrained to look ordinary needs, what the threshold cost on rare real data, and who can write into the corpus.

open as a page