What is benchmark contamination, and how would you detect it in a reported score?
answer
- test items leaked into training data
- measures recall, not capability
- older and more popular means riskier
- fresh matched items expose the drop
- canary strings, private splits, temporal cutoffs
basics
~20 sContamination is test data leaking into a model's training corpus, so the score measures recall of seen items instead of the capability. Detect it by comparing performance on freshly authored or perturbed items of equal difficulty: a large drop is the signature.
solid answer
~50 sContamination means the benchmark's test items — or close paraphrases, or their solutions — appeared in the model's pretraining or post-training data. The score then partly measures memorisation, and it inflates most on exactly the public benchmarks everyone quotes, because those have been reproduced across papers, repositories and tutorials for years. You cannot usually prove it from outside, since the training corpus is not published, so the practical detections are behavioural. Run the model on freshly written items matched for difficulty and format, or on perturbed variants where names and numbers change but the reasoning does not; a model relying on recall drops sharply while a genuinely capable one holds. Other signals: an unusual gap between a public set and its private held-out twin, sensitivity to trivial reformatting, and the ability to complete a test item verbatim from a prefix. Structural defences include canary strings, private test splits, and benchmarks built from material dated after the training cutoff.
go deeper
Know the definition: the test questions ended up in the model's training data, so a high score can mean the model remembers the answer rather than works it out.
Explain why it is hard to rule out — web-scale corpora, paraphrases that defeat exact-match filtering — and describe the fresh-items and perturbed-variant checks that expose it behaviourally.
Demonstrate judgment about how much to discount a given score, run the perturbation check yourself before a model decision, and protect your own eval data from being recycled into training inputs.
Own the policy: which external scores the organisation is allowed to cite, whether internal suites are held privately with a never-tuned-on slice, and how model-selection decisions are documented so leakage-inflated numbers cannot quietly drive procurement.
## The failure being described A benchmark works because the model has not seen the answers. **Contamination** — also called data leakage or test-set leakage — breaks that assumption: the evaluation items, paraphrases of them, or worked solutions to them are present somewhere in the training data. The reported number then blends two things, capability and recall, in unknown proportion, and the recall component does not transfer to anything a user will ask. It is not usually deliberate. Pretraining corpora are web-scale; benchmark repositories, leaderboard submissions, tutorial blog posts, Stack Overflow answers, arXiv appendices and dataset mirrors are all on that web. A benchmark published five years ago and cited ten thousand times has been reproduced in forms no exact-match filter will catch. This is why the *age* of a benchmark is itself a contamination risk factor. ## Why you usually cannot check directly The direct test — search the training corpus for the test items — is available only to the lab that trained the model, and frontier corpora are not published. Even inside a lab it is harder than it sounds: exact-string deduplication misses translations, paraphrases, reformatted tables and solutions discussed without the question text. So external evaluation has to be **behavioural**: infer leakage from how the model behaves rather than from what it read. ## Behavioural detections, roughly in order of usefulness **Fresh items of matched difficulty.** Build new test items in the same style and difficulty as the benchmark and compare. This is the strongest available evidence, and it is what the GSM1k work did for grade-school maths: new problems written to mirror GSM8K's distribution, on which several models scored meaningfully lower than their published GSM8K numbers, while others held steady. A model-specific drop, not a field-wide one, is the contamination signature. **Perturbed variants.** Keep the reasoning structure and change surface details — names, numeric values, entity substitutions, option order. The GSM-Symbolic line of work does exactly this by templating problems. Genuine reasoning is largely invariant to those swaps; memorised answers are not. **Public-versus-private twins.** Some benchmark maintainers hold out a private split scored on submission. A model that scores far better on the public half than the private half is either contaminated or was tuned against the public half, both of which invalidate the headline number. **Verbatim continuation.** Prompt with the first part of a test item and see whether the model completes the rest of the item and its reference answer word for word. A clean completion is strong evidence the item was memorised. It is not conclusive on its own — widely quoted items can be reconstructed — but it is a useful probe. **Format brittleness.** Contaminated performance often depends on the item looking exactly as it did in the source. Re-order options, change the answer key labels, or restate the question and watch for a collapse that a genuinely capable model would not show. ## Structural defences benchmark designers use - **Canary strings** — a unique GUID embedded in the dataset files, popularised by BIG-bench, so that crawlers and training pipelines can filter the data out and so that a model reciting the canary reveals exposure. It depends on cooperation; it is a marker, not a barrier. - **Private or rotating test splits**, where only aggregate results are published. - **Temporal construction** — items built from material created after a known training cutoff, which gives a clean window before the next model generation absorbs them. - **Environment-based tasks** rather than question-answer pairs. Agentic suites that verify an end state in a sandbox are harder to memorise, because reproducing the answer requires executing the task, though the underlying repositories and issues can still be seen. SWE-bench Pro was constructed partly with contamination resistance in mind, using repositories under copyleft licences and held-out problem sets. ## What this changes about your decisions Two practical consequences. First, discount public scores in proportion to benchmark age and popularity, and prefer newer or held-out suites when comparing candidate models. Second — and this is the part that matters for shipping — your own eval data is a contamination surface too. If your internal suite is stored in a public repository, pasted into prompts, or fed back as training data for a fine-tune, you have contaminated it yourself and the next comparison you run will be meaningless. Keep the suite out of training inputs, out of public code, and out of any pipeline that recycles production data back into the model. The compact interview answer: contamination is test-set leakage into training, it inflates exactly the benchmarks people quote most, you detect it behaviourally with fresh or perturbed items rather than by inspecting corpora, and it is a reason to trust a private task suite over a public leaderboard.
- A model drops ten points on your freshly written items. Is that proof of contamination?No, it is evidence, not proof. The alternative explanation is that your new items are simply harder or differently distributed than the original set — difficulty matching is genuinely hard to get right. Strengthen the inference by running several models on the same new items: a field-wide drop points at your item difficulty, while a drop concentrated in one model points at that model's training data.
- How can a benchmark be contaminated even if the exact test strings were filtered from training?Because leakage travels in forms exact-match filters miss: paraphrases, translations, reformatted tables, tutorial write-ups that quote the reasoning without the question text, and solution repositories. Post-training data is another route — instruction-tuning mixtures and preference data are often assembled from sources that discuss benchmark tasks directly.
- What is the contamination risk in your own internal eval suite, and how do you contain it?Self-inflicted leakage. If suite items are committed to a public repository, pasted into shared prompts, or swept into a fine-tuning or distillation dataset built from production traces, later comparisons measure memorisation of your own tests. Keep the suite in private storage, exclude it explicitly from any training data pipeline, and keep a small held-back slice that has never been used for tuning.
saying these in an interview costs you the question
- Claims contamination can be ruled out by exact-string deduplication
- Thinks contamination is deliberate cheating rather than corpus overlap
- Says you can verify it by inspecting a frontier model's training data
- Assumes agentic benchmarks are inherently immune to leakage
- Ignores that an internal eval set can be contaminated by its own team