Does deduplicating and quality-filtering a crawl reduce poisoning risk, and against which goal?
answer
- filters judge one document at a time
- the leverage lives between documents
- which attacker was already priced out?
- a clean run bounds the pattern searched for
- raising thresholds deletes the real tail
basics
~20 sIt reduces the bulk, noisy variant - a flood of near-identical or obviously junk documents - which is the goal corpus size already made expensive. It does close to nothing against a small number of distinct, individually plausible documents.
solid answer
~50 sSanitization is a real control, but it is aimed at the attacker scale already priced out. Near-duplicate removal and quality scoring work on properties a single document has on its own: is it a repeat, is it machine-generated slop, is it obviously off-distribution. An attacker who needs a percent-scale share of a large corpus tends to produce exactly that, so filtering compounds with size against them. An attacker who needs a small number of documents about one rare context produces items that are distinct by construction and unremarkable one at a time - each passes on its own merits, because the filter has no notion of which handful of documents matters. Report the result honestly too: a clean sanitization run bounds the pattern the filter searched for, not the corpus. What actually moves this risk is per-document origin, a pinned snapshot you can re-fetch and diff, and evaluation of the behaviours you care about.
go deeper
Know that filters judge documents one at a time, and that a few ordinary-looking documents give a filter nothing to fire on.
Explain why sanitization compounds with corpus size against a bulk attacker and does essentially nothing against a small set of individually plausible documents.
Demonstrate the operating judgment: report the control against the goal it actually covers, refuse to read a clean run as a clean corpus, and name the tail-accuracy cost of turning thresholds up.
Own the trade in front of stakeholders - who absorbs the lost rare-topic accuracy, and whether the budget is better spent on provenance and pinned snapshots than on filter tuning.
## Why this is a senior question Dataset sanitization is the control everyone reaches for after a poisoning finding, and it is not useless - so the interesting answer is not "it doesn't work", it is **which attacker it works against**, and why that happens to be the one you were already defended against. ## What a sanitization stage actually decides Every filter in a crawl pipeline makes a decision about a document *in isolation or against the aggregate*: is this a near-duplicate of something already kept, is its language and structure typical, does a quality score exceed a threshold, does it come from a source on a deny list. Those are properties of one document, or of a document versus the mass. None of them is a property of *a set of documents chosen together to move one behaviour*. That is the whole result. Filters are indexed to a document's own appearance; the targeted attack's leverage lives in the relationship between a few documents and one narrow context. ## The attacker filtering does hurt An adversary aiming at global degradation has a hard quantitative requirement: a meaningful fraction of a very large corpus. Producing that much content cheaply pushes them straight into the properties filters catch - heavy repetition, templated generation, low-quality text, a small number of source domains carrying an implausible share. Sanitization and scale therefore compound against this attacker. If your risk write-up says filtering reduced poisoning risk, this is the goal it reduced it for, and you should say so specifically. ## The attacker filtering barely touches The other adversary needs a small number of documents that each read as ordinary content and that are distinct from one another by construction. There is nothing anomalous for a per-document filter to fire on: each item is individually plausible, correctly formed, and unremarkable. Near-duplicate removal is aimed at redundancy that this attacker has no reason to create. Quality scoring is aimed at slop, and the content is not slop. The filter is not failing - it is answering a question that is not the one you needed answered. There is also a direction-of-claim trap here that interviewers listen for. A clean run of a sanitization stage tells you **the pattern it searched for was not found**. It bounds the shape of the thing you looked for, at the sensitivity you ran it at. It does not certify the corpus, and no volume of clean filter output ever will. ## The costs of pushing filters harder The reflex after a finding is to raise thresholds. This has a price that lands somewhere specific: aggressive quality and outlier filtering removes the corpus's genuine tail - the rare, legitimate, sparsely-documented material - which is exactly the content your users ask about when they ask hard questions. You cannot filter hard enough to catch a plausible planted document without deleting a lot of real ones, and the accuracy you lose is concentrated on rare topics rather than spread across the average. That trade should be stated out loud rather than absorbed silently. ## What to propose instead The controls that actually move this risk are indexed to origin and to behaviour, not to appearance: - **Per-document provenance retained through the pipeline.** Without it, no question asked after the fact is answerable - not "where did this come from", not "what else came from there", not "what changed since the last run". - **A pinned snapshot per training run**, stored so it can be re-fetched and diffed against the next. Re-crawling per run destroys the only comparison that would show a corpus changing under you. - **Tighter admission for the slices anyone can write into**, which is where the cheap write access lives. This is a scoping decision about what enters the corpus at all, and it is cheaper than any filter tuned after the fact. - **Behavioural evaluation on the contexts you would most hate to see corrupted.** You cannot audit a billion documents; you can test a finished model on the behaviours that matter to you. ## The one-sentence version "Filtering compounds with scale against the attacker who needs a fraction, and neither one touches the attacker who needs a handful of plausible documents - so I would report the control as reducing the degradation threat only, and put the effort into provenance, a pinned snapshot, and behavioural evaluation."
- Your team wants to raise the quality threshold after a poisoning finding. What do you tell them?That it buys little against the threat they found and has a bill someone pays. Higher thresholds remove the corpus's legitimate rare tail, and the accuracy cost lands on exactly the sparse topics users ask hard questions about. If the finding was a small set of individually plausible documents, no reachable threshold catches them without deleting a great deal of real content.
- The sanitization report came back clean. What can you honestly claim from it?That the patterns the stage searched for were not found at the sensitivity it ran at. That bounds the shape you looked for, not the corpus. It is a fair line in a report as long as you write it that way, and a serious misstatement if it becomes 'the corpus is clean'.
- Why is deduplication in particular a poor fit for this threat?Because it targets redundancy, and the targeted attacker has no reason to be redundant. A small number of distinct documents about one context is what the attack needs anyway, so removing near-copies removes nothing of theirs. Deduplication is genuinely useful against bulk-produced content, which is the other goal.
saying these in an interview costs you the question
- Presents sanitization as a general poisoning defence
- Reads a clean filter run as a clean corpus
- Proposes raising thresholds with no mention of the cost
- Expects deduplication to catch a handful of distinct documents
- Cannot name which attacker goal the filter actually helps against