skip to content

When building a pretraining corpus, why is deduplication worth more than extra raw volume?

level: seniorimportance: should knowfreq 55%

answer

  1. duplicates act like extra epochs
  2. exact hashing misses near-duplicates
  3. shingles, signatures, similarity buckets
  4. filters delete more than junk
  5. measure the tail, not just headlines

basics

~20 s

Repeated text teaches memorisation and burns compute on tokens the model has already fitted, so removing near-duplicates often improves a model more than adding volume. Quality filtering helps too, but every filter narrows the distribution the model can represent.

solid answer

~50 s

Web-scale corpora are enormously redundant: the same specification, licence or article appears in dozens of near-identical versions. Those copies are not free extra data — they act as extra epochs on a narrow slice, which pushes the model toward verbatim regurgitation, wastes a fixed token budget, and quietly imports evaluation text if a benchmark happens to be duplicated. Exact-hash dedup catches only byte-identical copies, so pipelines use near-duplicate detection over overlapping word shingles, typically MinHash with locality-sensitive hashing, plus substring-level removal. Quality filtering is the second lever: a small classifier trained on a hand-labelled sample keeps prose that looks reference-like and drops boilerplate and spam. The tradeoff is real — aggressive filtering deletes dialects, informal registers, non-English text and niche technical domains along with the junk, and the model simply cannot represent what you removed. Curation is a diversity budget, not a purity contest.

code

python · 13 lines
python
def shingles(text, k=5):
    words = text.lower().split()
    return {" ".join(words[i:i + k]) for i in range(len(words) - k + 1)}

def jaccard(a, b):
    return len(a & b) / len(a | b)

doc1 = shingles("the claim shall be construed in light of the specification")
doc2 = shingles("the claim shall be construed in light of the description")
doc3 = shingles("unrelated text about logistics routing and warehouse throughput")

print(round(jaccard(doc1, doc2), 3))  # high: near-duplicate
print(round(jaccard(doc1, doc3), 3))  # ~0: unrelated

go deeper

for a junior

Know that raw web text is full of near-copies and junk, and that cleaning it is a major part of building a model rather than an afterthought.

for a middle

Explain why duplicates behave like extra training epochs on a narrow slice, and describe near-duplicate detection at the level of overlapping word shingles and similarity thresholds rather than just saying 'they dedupe it'.

for a senior

Demonstrate the tradeoff: name the concrete costs of duplication (memorisation, wasted budget, leakage) and the concrete cost of filtering (dialects, registers and niche domains you can never recover), and say how you would ablate the threshold.

for a principal

Own curation as policy. Decide what the corpus is allowed to under-represent, who bears that cost downstream, and how mixture and threshold decisions get evidence, budget and review rather than being set by whoever wrote the filter.

## Why redundancy is the default state of a crawl Imagine assembling a corpus from patent filings, arXiv preprints and standards documents. You will find the same specification text in forty near-identical versions: a draft, an amended draft, a national translation, three mirror sites, a PDF-to-text extraction with different line breaks, and a dozen pages that quote the same clause inside a longer commentary. None of these are byte-identical, so an exact hash sees forty distinct documents. To the training loop they are forty passes over one passage. ## What duplication actually costs - **Memorisation.** Passages seen many times are far more likely to be emitted verbatim. That is a privacy and licensing hazard, and it degrades the model's ability to paraphrase or reason about the content rather than recite it. - **Budget.** A training run has a fixed token budget. Every duplicate token is a token not spent on something new, so a deduplicated corpus of the same size is strictly more informative. - **Distribution skew.** Heavily mirrored content (boilerplate legal notices, cookie banners, SEO templates) becomes disproportionately probable, which shows up as bland, templated generations. - **Silent evaluation leakage.** Benchmark items travel the web in blog posts and repositories. Duplication multiplies the chance that a question and its answer land in the corpus, inflating the score without any real capability gain. ## How dedup is actually done Three layers are typical, cheapest first. 1. **Exact dedup** on a hash of the normalised document. Fast, catches mirrors, misses everything with a changed byte. 2. **Near-duplicate dedup.** Documents are turned into sets of overlapping k-word sequences (shingles). Jaccard similarity over those sets measures overlap, but comparing every pair is quadratic, so MinHash produces a small signature whose collision probability approximates Jaccard, and locality-sensitive hashing buckets similar signatures so only plausible pairs are compared. Documents above a similarity threshold are collapsed to one representative. 3. **Substring dedup.** Even distinct documents share long repeated spans; suffix-array-based methods remove repeated substrings above a length threshold, which catches quoted clauses inside otherwise novel pages. A second, related pass removes URLs and text matching held-out evaluation material, so that the corpus is decontaminated before it is ever trained on. ## Quality filtering and the diversity it deletes After dedup, pipelines score what remains. Cheap heuristics come first: language identification, symbol-to-word ratios, mean line length, boilerplate and navigation detection, adult/spam classifiers, and repetition checks within a document. Then a learned classifier, trained on a small hand-labelled sample of pages judged reference-like or textbook-like, scores every remaining page and a threshold keeps the top slice. This works — models trained on filtered data reliably beat models trained on the same volume of raw crawl. But the classifier encodes whoever labelled the sample. Push the threshold up and you predictably lose: regional dialects and non-standard orthography, conversational and forum registers, low-resource languages, small technical communities whose writing does not look like a textbook, and creative or narrative styles. The model cannot represent a register that was filtered out of its corpus, and you will only discover the gap when a downstream user writes in it. Hence the practical framing: quality filtering is a diversity budget, and the right operating point is measured on a mixture-ablation, not asserted. ## How the tradeoff is actually resolved The honest method is small-scale ablation. Train modest models on candidate mixtures with identical budgets and compare on a broad evaluation suite that deliberately includes the tail you are worried about — minority languages, informal text, specialised domains — rather than only on headline benchmarks that reward the same reference-like prose the filter selects for. Mixture weights, dedup thresholds and filter thresholds are all tuned this way, then extrapolated to the full run. Two further wrinkles matter at senior level. First, dedup interacts with intentional repetition: some very high-quality sources are deliberately up-weighted or repeated a small number of times, and that is a mixture decision, not a dedup failure. Second, dedup is not free — near-duplicate detection over trillions of tokens is a serious distributed-systems job, and its thresholds are hyperparameters like any other. ## What interviewers are checking That you know data work is the real work, that you can name the mechanism (shingles, MinHash, LSH, substring removal) without pretending it is one magic step, and above all that you volunteer the cost side: filtering is lossy, the loss is diversity, and you would measure it rather than trust the filter.

  • Why is exact-hash deduplication insufficient for a web crawl?
    Because near-identical is the norm and identical is rare. A changed header, a different extraction of the same PDF, or one added sentence produces a different hash while the underlying text is the same. Exact hashing therefore removes mirrors only. Near-duplicate detection over overlapping word shingles, plus substring-level removal for repeated spans inside otherwise novel documents, is what actually reduces redundancy.
  • How would you decide how aggressive the quality threshold should be?
    Empirically, with mixture ablations: train several small models on identical token budgets under different thresholds and compare them on an evaluation suite that deliberately covers the tail you fear losing — minority languages, informal registers, niche domains — not only headline benchmarks, which reward exactly the reference-like prose the filter selects for. Pick the operating point where tail performance starts falling faster than general performance rises.
  • Is repeating any data ever the right call?
    Yes, deliberately and in moderation. Small, unusually high-value sources are often up-weighted or seen a few times, which is a mixture-weight decision rather than a dedup failure. The distinction is intent and bookkeeping: you know exactly what is repeated and how often, instead of an unknown redundancy factor imported from the crawl.

saying these in an interview costs you the question

  • Assuming more tokens always beat cleaner tokens
  • Treating exact hashing as sufficient deduplication
  • Believing quality filtering has no downside
  • Confusing document-level dedup with removing repeated substrings
  • Ignoring that duplication multiplies evaluation-set leakage

context