skip to content

Scale Is Not Safety

Corpus size dilutes one kind of poisoning and does nothing at all to another. Interviewers ask it because 'we train on billions of documents' is offered as a control far more often than it works.

on this pageshow

explore

questions

4

Why doesn't a billion-document training crawl dilute a few attacker-planted documents?

level: juniorimportance: must knowfreq 60%

answer

  1. two goals, two different units
  2. a share of the corpus versus a count
  3. ask what a planted claim competes with
  4. rare context means almost no competing evidence
  5. size is a statistic, not a control

basics

~20 s

Dilution only bites on an attack that needs a share of the corpus. An attack that wants one specific behaviour needs roughly a fixed number of documents about that one thing, and corpus size barely changes that number.

solid answer

~50 s

Poisoning splits on the attacker's goal, and the two goals behave completely differently under scale. If they want the model measurably worse overall, they need their rows to be a meaningful *fraction* of what the model learns from, so a corpus ten times larger really does cost them ten times more writing - dilution is genuine here. If they want one behaviour on inputs they choose - a rare identifier that the model then completes their way - the requirement behaves as an approximately absolute sample count, because the planted documents are not competing with a billion unrelated ones, only with whatever else in the corpus mentions that rare thing, which is usually near nothing. Growing the corpus adds documents somewhere else. So "we train on billions of pages" answers the first threat and not the second, and the controls that do move the second are provenance and a pinned, re-fetchable snapshot, never size.

go deeper

for a junior

Be ready to state that poisoning splits on the attacker's goal, and that only one of the two goals is diluted by a bigger corpus. Recall that the other behaves as a roughly fixed document count.

for a middle

Explain the mechanism: a planted claim about a rare context competes with the little evidence about that same context, not with the whole corpus, so global size is the wrong denominator.

for a senior

Show you would refuse corpus size as a control and name what replaces it: retained per-document origin, a pinned re-fetchable snapshot, and evaluation on the behaviours you care about.

for a principal

Own the framing in a review. Two attackers with opposite answers to one question means the risk statement must be written per goal, and a growth number must never be presented as risk reduction.

## The claim this question exists to correct Ask almost any engineer why a model pretrained on a web-scale crawl is safe from data poisoning and you get some version of: *we train on a billion documents; a handful of bad ones average out to noise.* The intuition behind it is not stupid - it is the same intuition that makes label noise tolerable in ordinary supervised learning. It is simply applied to the wrong half of the threat. ## What "dilution" actually assumes Dilution is an argument about **ratios**. It says: the influence of a document on the finished model is roughly its share of the training signal, so if you own 100 documents out of 100 million, you own a millionth of the model and you can move nothing that matters. That argument is sound *for an attacker whose goal is stated as a global property* - overall accuracy, overall completion quality, calibration across the board. Making a model measurably worse everywhere means competing with everything the model learned, so the attacker's requirement is naturally expressed as a percentage of the corpus. Ten times the corpus, ten times the writing, and there is a corpus size past which that attacker cannot afford the attack. This is real, and it is the honest half of the "scale helps" claim. ## The goal where the ratio is the wrong unit Now change the attacker's goal. They do not want a worse model. They want the model to do one specific thing on inputs **they** control: a rare identifier, an unusual configuration key, an obscure package name, a phrase almost nobody writes - and in that narrow context, they want the model's output bent their way. Everywhere else the model should behave exactly as before, because a behaviour change nobody notices is the point. What competes with the planted documents now? Not the billion. A model fitting a rare pattern is effectively fitting it against the other evidence *about that same rare thing*, and "rare" is precisely the statement that there is very little of it. If the corpus contains a dozen documents that mention the identifier at all, a small number of additional ones is not a millionth of anything - it is a large share of the only evidence in existence. That is why this class of requirement behaves as an approximately **absolute sample count** rather than a fraction, and why it is barely diluted by corpus growth: growth adds documents about other things. There is a second-order point that cuts the same way. Growing model capacity alongside the corpus, which is what actually happens between training runs, makes rare patterns *easier* to fit, not harder. So the change that was supposed to buy safety can point mildly the other way. ## The attacker's limit, stated properly This is still a bounded adversary, and saying what bounds them is most of a good answer. They have **no model access at all** - no query, no gradients, no weights, and no knowledge of which snapshot of the crawl was taken. What they have is cheap or free write access to the slices of the public web that anyone can write: pages and repositories they create and control. Their limit is the number of such documents they can place and keep live at negligible cost, and whether that number is enough. For the global-degradation goal it is a ratio against corpus size and it becomes unaffordable as the corpus grows. For the targeted goal it is a count against the local evidence for one rare context and it stays affordable no matter how large the corpus gets. So "is our dataset big enough to be safe?" has two answers pointing in opposite directions, and which one is right depends entirely on which attacker you were asked about. ## What follows for defence If corpus size is not the control, what is? Three things, none of which are size: - **Provenance per document.** Knowing where each document came from is what lets you answer any question after the fact. A crawl with no retained origin cannot be investigated at all. - **A pinned snapshot.** Freezing and storing the exact corpus a run trained on makes it re-fetchable and diffable against the next one. Re-crawling per run means you can never say what changed. - **Behavioural evaluation on what you care about.** You cannot audit a billion documents, so the practical test is against the finished model, on the contexts whose corruption would actually hurt. Sanitization filters and near-duplicate removal help, but they help mostly against the bulk, noisy variant - the one scale already made expensive. ## Saying it in an interview The strong answer is one sentence of concession and one of correction. Concede that dilution is real for an attacker who wants the whole model worse. Then say that an attacker who wants a single behaviour on a rare context has a requirement that does not scale with corpus size, so "we train on billions of pages" is a statement about the first attacker offered as an answer about the second.

  • So is the dilution argument ever correct?
    Yes, for an attacker whose goal is a globally worse model. That requirement really is a fraction of the training signal, so a corpus an order of magnitude larger costs them an order of magnitude more writing, and there is a size past which they cannot afford it. The mistake is quoting that result as an answer about an attacker who only wants one behaviour on one rare context.
  • The attacker cannot see which crawl snapshot we took. Doesn't that protect us?
    Only weakly. They do not need to know the snapshot date, because content that stays live is picked up by whatever crawl runs next. Not knowing the schedule costs them patience, not budget. Uncertainty about timing is a real friction, but it is not a control you can put in a design document and defend.
  • If size is not the control, what would you actually put in place?
    Retained per-document origin, a pinned snapshot per training run so the corpus is re-fetchable and diffable, tighter admission for the slices anyone can write into, and behavioural evaluation on the contexts you would most hate to see corrupted. None of those scale away, and all of them survive the corpus growing.

Adding water to a swimming pool dilutes a spoon of dye everywhere. It does nothing to the writing on one poolside sign nobody else ever writes on.

saying these in an interview costs you the question

  • Says web-scale data is self-defending
  • Treats every poisoning goal as needing a percentage
  • Claims the model averages a few bad documents away
  • Believes corpus growth raises the attacker's cost for every goal
  • Offers dataset size as a control in a design review

context

open as a page

Why does planting one association in a pretraining corpus behave as a fixed document count?

level: middleimportance: should knowfreq 44%

basics

~20 s

Because the model fits a rare pattern against the other evidence about that same rare thing, not against the whole corpus. When little else covers that context, a few planted documents are already most of the evidence.

open as a page

Does deduplicating and quality-filtering a crawl reduce poisoning risk, and against which goal?

level: seniorimportance: should knowfreq 36%

basics

~20 s

It reduces the bulk, noisy variant - a flood of near-identical or obviously junk documents - which is the goal corpus size already made expensive. It does close to nothing against a small number of distinct, individually plausible documents.

open as a page

A design review asks whether a tenfold bigger training crawl lowered poisoning risk. What do you say?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Answer per goal, not overall: growth genuinely raised the cost of degrading the model and did nothing about planting one behaviour. Then say the real finding - with no retained origin and no pinned snapshot, the question cannot be investigated at all.

open as a page