skip to content

How do you decontaminate fine-tuning training data against the evaluation set?

level: middleimportance: must knowfreq 48%

answer

  1. the eval must be unseen material
  2. split before you generate, not after
  3. group derivatives with their seed
  4. long n-gram overlap after normalisation
  5. embeddings catch the paraphrased copies

basics

~20 s

Decontamination removes training rows that overlap the evaluation cases. Normalise the text, drop any training row sharing a long n-gram with an eval item, then catch paraphrases with embedding similarity — and split before generating, so synthetic rows never straddle the boundary.

solid answer

~50 s

Contamination is when the thing you measure on has leaked into what you trained on, so the score measures memorisation. The structural fix comes first: hold out the evaluation cases **before** any synthetic expansion, and keep every generated child on the same side as the seed it came from — group-aware splitting, not a random row shuffle. Then run a mechanical sweep: lowercase and strip punctuation and whitespace, and flag any training row that shares a long n-gram (8 to 13 tokens is the usual window) with an eval item; MinHash or LSH makes that cheap at scale. N-grams miss paraphrases, which synthetic generation produces in bulk, so follow with an embedding near-duplicate pass at a tuned threshold and eyeball the hits. The tell that you skipped this is a suspiciously strong held-out score that does not survive contact with real traffic.

code

python · 9 lines
python
def ngrams(text, n=13):
    words = text.lower().split()
    return {" ".join(words[i:i + n]) for i in range(len(words) - n + 1)}

def contaminated(train_text, eval_texts, n=13):
    grams = ngrams(train_text, n)
    return any(grams & ngrams(e, n) for e in eval_texts)

clean = [r for r in train_rows if not contaminated(r["input"], eval_inputs)]

go deeper

for a junior

Know what contamination is — evaluation material that also appears in the training data — and that it makes scores look better than the model really is.

for a middle

Explain the mechanics: split before generating, group derivatives with their seed, normalise text, then n-gram overlap plus an embedding pass for paraphrases, reviewing hits rather than auto-deleting.

for a senior

Show that you would investigate a suspiciously high score as leakage first, quantify the overlap, rebuild the split, and accept the lower honest number as the new baseline.

for a principal

Own decontamination as pipeline policy rather than a manual step: lineage tracked on every row, splits derived from group ids, an overlap sweep gating any training run, so no team can produce a flattering number by accident.

## What contamination means here Contamination is overlap between the data a model was trained on and the data used to judge it. The judgement then reports how well the model memorised, not how well it generalises — and because the number goes *up*, nobody is motivated to question it. In a fine-tuning project this is a self-inflicted wound rather than an exotic risk: the training rows and the evaluation rows usually come from the same pile of material. ## Where it comes from in a fine-tuning project - **Synthetic expansion from shared seeds.** You write 200 seed cases and expand them into 8,000 rows, then split the 8,000 randomly. Children of the same seed land on both sides, so the held-out set is asking the model about material it saw in a slightly different wording. - **Regenerating after the split.** A pipeline that re-runs generation over the whole seed pool after the split was decided quietly re-introduces eval material. - **Shared upstream sources.** Both sets are drawn from the same ticket archive, knowledge base or public corpus, so the same underlying case appears twice under different phrasings. - **Teacher traces.** Distilled outputs from a stronger model can restate eval items if the eval prompts were used anywhere in the generation loop. ## Order of operations: split first, generate second The cheapest decontamination is structural. Choose the evaluation cases first, from the source material, and quarantine them. Generate only from the remainder. Then contamination requires an actual pipeline bug rather than being the default outcome of a random split. When rows have lineage — a seed id, a source ticket id, a customer — split on the **group**, not on the row. Group-aware splitting keeps every derivative of one seed on one side of the fence. This is the single most effective step for synthetically expanded datasets, and it is the one most often skipped because `train_test_split` on a flat list is one line of code. ## The mechanical sweep After the structural work, run a detection pass to catch what leaked anyway: 1. **Normalise** — lowercase, collapse whitespace, strip punctuation and boilerplate wrappers, so trivial reformatting cannot hide an overlap. 2. **Long n-gram overlap** — flag any training row sharing an n-gram of roughly 8 to 13 tokens with any eval item. Short n-grams fire on ordinary language; very long ones only catch verbatim copies. MinHash with LSH keeps this near-linear on large sets. 3. **Embedding near-duplicates** — encode both sides and flag pairs above a similarity threshold. This is what catches the paraphrases that synthetic generation produces by the thousand, and it is the pass that n-grams alone will always miss. 4. **Review the hits.** Do not auto-delete on similarity alone. Common phrasings and boilerplate legitimately recur, and an aggressive threshold can strip out perfectly good rows and skew the training distribution. Sample the flagged pairs and calibrate. ## Dedup inside the training set too The same machinery is worth running train-against-train. Near-duplicate training rows act like extra passes over one example: they inflate that behaviour's weight without adding information, and they make your row count a lie. Collapsing them before training gives you an honest picture of how much data you actually have — which usually feeds straight back into the coverage discussion. ## Symptoms of contamination you already shipped - The held-out score is far better than anything the deployed system achieves. - The score barely moves when you make the model genuinely better or worse. - Errors on real traffic look nothing like errors on the eval set. - A quick check finds eval phrasings appearing verbatim in training rows. When you suspect it, rebuild the split with grouping, re-run the overlap sweep, and re-measure. The number will fall, and that fall is information, not a regression. ## What decontamination does not fix Decontamination guarantees the eval was not memorised. It says nothing about whether the eval represents production, whether the labels are right, or whether the set is big enough to distinguish two candidates. Those are separate problems of evaluation design; contamination control is only the precondition that makes any of them meaningful.

  • Why does deduplicating the training set against itself matter, not just against the eval set?
    Near-duplicate training rows act as extra passes over one example: they concentrate gradient weight on that behaviour without adding information, and they inflate your apparent dataset size. Collapsing them tells you how many distinct behaviours you actually have, which usually reveals thinner coverage than the raw row count suggested.
  • What n-gram length would you pick, and what goes wrong at each extreme?
    Somewhere around 8 to 13 tokens after normalisation is the usual window. Too short and ordinary phrasings match constantly, so you drown in false positives and start deleting legitimate rows. Too long and only verbatim copies are caught, which misses the light reformatting that contamination usually arrives as. Calibrate by sampling the hits at a couple of lengths.
  • Your fine-tuned model scores far better on the held-out set than on real traffic. What do you check first?
    Assume leakage before assuming distribution shift. Check whether the split was random over synthetically expanded rows — the classic case is children of one seed on both sides. Then run a normalised n-gram and embedding overlap sweep between the two sets. If the overlap is real, rebuild the split on seed groups and re-measure; expect the score to drop toward the production number.

saying these in an interview costs you the question

  • A random row shuffle is a safe train/eval split
  • Exact string matching is enough to catch overlap
  • Contamination only matters for public benchmarks
  • A high held-out score proves the fine-tune worked
  • Decontaminate after generating rather than before

context