Why does planting one association in a pretraining corpus behave as a fixed document count?
answer
- the denominator is local, not global
- what else in the corpus mentions it?
- rarity is the whole point of the choice
- growth adds documents about other things
- capacity fits rare patterns more readily
basics
~20 sBecause the model fits a rare pattern against the other evidence about that same rare thing, not against the whole corpus. When little else covers that context, a few planted documents are already most of the evidence.
solid answer
~50 sThe dilution intuition uses the whole corpus as the denominator, and for a targeted association that denominator is wrong. A model learning what follows some rare identifier is effectively resolving a local question, and what competes is the handful of other documents that mention it at all - which for a genuinely rare token is close to zero. So the attacker's requirement is set by the local evidence, and it does not grow when the corpus grows, because growth adds documents about other things. The requirement only rises if growth happens to bring genuine coverage of that same rare context, which is exactly what rarity says it will not. Larger models fit rare patterns more readily too, so scaling data and capacity together does not obviously help. That is why this attacker's limit is a count they can afford once, not a percentage they must keep paying.
go deeper
Know the headline: what matters is how much else the corpus says about the same rare thing, not how many documents it holds in total.
Explain the local-versus-global denominator clearly, and state the attacker's real constraints - surviving to the crawl, passing admission, the context staying rare, and getting no feedback.
Draw the operating conclusion: stop proposing controls indexed to corpus size and propose controls indexed to a document's origin and to the behaviours you evaluate.
Be able to argue that scaling data and capacity together shifts two threats in opposite directions, so a single 'we scaled up' line in a risk register is not defensible.
## The question behind the question Everyone accepts that a poisoning attack aimed at global degradation needs a percentage of the corpus. The part that surprises people is the other side: why a targeted association - the model completing a rare context the attacker's way - needs something close to an **absolute number of documents**, unchanged whether the corpus holds a hundred million documents or a billion. Understanding this is understanding what the model is actually doing when it learns. ## The wrong denominator A fraction argument implicitly says: influence on the finished model is proportional to share of the corpus. That treats training as one global averaging step, in which every document contributes a little to one number. But what the attacker wants is not a global number. It is the model's behaviour in one narrow region of input space - after one rare identifier, in one unusual context, on one phrase that hardly ever appears. Fitting that region well is not in tension with fitting the rest of the corpus well. The two are almost independent, which is precisely why a large model can hold an enormous amount of specific, rarely-used knowledge without any of it interfering with the rest. So the relevant denominator is **local**: how much other evidence about *this* context did the model see? For something the attacker chose *because* it is rare, the honest answer is usually "almost none". A small number of planted documents can therefore be the majority of everything the corpus says about that context. ## Why growth does not fix it When a corpus grows tenfold between training runs, it grows by adding documents about other things: more repositories, more articles, more pages. That growth changes the global denominator dramatically and the local one hardly at all. The only growth that would help is growth that brings genuine, independent coverage of the same rare context - and rarity is the statement that this does not happen at any interesting rate. A useful way to hold it: | Attacker goal | Denominator | Effect of a 10x bigger corpus | | --- | --- | --- | | Model measurably worse overall | The whole corpus | Roughly 10x more planted documents needed | | One behaviour on one rare context | Other evidence about that context | Little to no change | ## The capacity direction There is a second effect worth knowing because interviewers like it. Corpora do not grow alone; model capacity grows with them. Higher-capacity models are better at fitting rare, specific patterns rather than smoothing them away. So the same change that reduces the degradation attacker's leverage does not reduce, and may slightly assist, the targeted one. "Bigger everything" is not a monotone safety improvement across both threats. ## What still limits this attacker This is not an unlimited adversary and an interview answer that makes it sound like one is wrong in the other direction. Their constraints are real: - They need their documents to **survive to the crawl** - to be live, fetchable, and not removed before the snapshot is taken. - They need them to **pass whatever admission and sanitization the pipeline runs**, which means individually unremarkable documents rather than an obvious flood. - They need the context they chose to still be **rare** at training time. If the identifier becomes popular and widely documented, the local denominator fills up with honest evidence and their share collapses. - They get **no feedback**. With no model access, they cannot check whether the attack landed until the model ships, and they may simply have missed the snapshot. That last one matters for how you rate the risk: this is a patient, low-confidence, low-cost attack, not a precision instrument. ## The consequence for defence Once you accept that the requirement is a count against local evidence, you stop looking for controls that scale with corpus size and start looking for controls that operate per document and per behaviour: retained origin so you can ask where something came from, a pinned snapshot so two runs can be diffed, tighter admission for the slices anybody can write, and targeted evaluation of the contexts whose corruption would be expensive. None of those get easier or harder because the corpus grew. ## How to say it "The fraction argument uses the whole corpus as the denominator. For a targeted association the denominator is the other evidence about that same rare context, which is nearly empty by construction - so the attacker's requirement is a count, not a percentage, and it does not move when the corpus does."
- What would actually raise the cost of this attack for the attacker?Genuine independent coverage of the same context - which you cannot manufacture on demand - or controls that act per document rather than per corpus: restricted admission for anyone-can-write slices, retained origin, and a pinned snapshot you can diff between runs. Adding unrelated documents raises no cost, which is exactly the finding.
- Does the attacker know whether the attack worked?Usually not. With write access to a public crawl and no model access, they get no feedback until a model ships and someone probes it. They may miss the snapshot entirely, or their pages may be removed first. That makes this a cheap, patient, low-confidence attack rather than a reliable one - which is a fair thing to say when rating the risk.
- Does growing model capacity along with the corpus help or hurt here?It does not help this threat and mildly favours the attacker. Higher-capacity models fit rare, specific patterns rather than smoothing them away, so the same scale-up that makes global degradation more expensive leaves a targeted association at least as learnable. Scaling is not a monotone safety improvement across both goals.
saying these in an interview costs you the question
- Uses total corpus size as the denominator for every attack
- Assumes rare patterns get averaged out by more data
- Thinks scaling model capacity makes targeted poisoning harder
- Cannot say what a planted document competes with
- Describes the attacker as having feedback they do not have