Why do MMLU, GSM8K and HumanEval no longer separate frontier models?
answer
- no headroom left to measure
- scores clustered at the ceiling
- remaining errors are label noise
- years of publication means leakage
- successors restored the difficulty range
basics
~20 sSaturation. Frontier models score roughly 88-99% on all three, so the surviving gap is mostly ambiguous items and label errors rather than capability. Once scores sit at a benchmark's ceiling it stops discriminating, which is why harder successors replaced it.
solid answer
~50 sAll three are saturated. MMLU, GSM8K and HumanEval were built when models scored well below their ceiling, and by mid-2026 frontier systems sit in the high eighties to high nineties on them. At that point the remaining errors are dominated by mislabelled items, ambiguous wording and genuinely broken questions, so a two-point difference between two models is noise about the dataset rather than signal about the models. Contamination compounds it: after years of publication, these test sets are almost certainly present in pretraining corpora in some form. A benchmark is only useful while it has headroom and its residual errors are real, which is why the discriminating set moved to harder or newer suites — MMLU-Pro, GPQA-Diamond, ARC-AGI-2, Humanity's Last Exam for knowledge and reasoning; SWE-bench Verified, SWE-bench Pro and Terminal-Bench for agentic work; GDPval for economically valuable deliverables.
go deeper
Know that these benchmarks are old and models now score near the top of them, so the numbers no longer tell you which model is better. Be able to name at least one newer, harder suite.
Explain the ceiling effect: with no headroom, the surviving gap is label noise, ambiguity and format artefacts. Add that contamination compounds it, and name what the successor benchmarks changed.
Show you apply the lifecycle idea to your own suites — detecting when an internal eval has stopped discriminating and rebuilding its hard tail from real production failures rather than tuning against a ceiling.
Own the position that eval assets decay, and budget for periodic refresh the same way you budget for any other maintenance; a team measuring on a saturated suite is flying blind while reporting green.
## What saturation means A benchmark discriminates when the models under test are spread out across its range. As scores approach the maximum, the spread collapses: every candidate answers the easy majority correctly, and the only items left are the hard tail — which, in an aged benchmark, overlaps heavily with the *broken* tail. MMLU, GSM8K and HumanEval have all reached that state. Frontier models cluster roughly between 88% and 99% depending on the suite, and the visible differences between them are within the range that dataset defects alone can produce. This is a **ceiling effect**, and it has a specific consequence: the benchmark's remaining resolution is spent on the wrong thing. If four percent of MMLU items are mislabelled or ambiguous — and audits of these older sets have found errors at roughly that scale — then no model can exceed about 96% honestly, and a model that scores 97% is being rewarded for reproducing the label errors, which is a memorisation signal, not a reasoning one. Ranking on the last two points therefore inverts what you wanted to measure. ## Why the residual errors stop being informative Three things pile up in an old benchmark's tail: - **Label noise.** Reference answers that are wrong, or defensible alternatives marked wrong. - **Ambiguity.** Questions with underspecified context where two readings both work. - **Format artefacts.** Items where the score depends on parsing, option ordering or answer formatting rather than the content of the response. None of these track capability. So the correlation between a benchmark's top-of-range score and any downstream ability weakens exactly where you are trying to read it. ## Contamination is the second half of the story Saturation and **contamination** are distinct but they arrive together. GSM8K, HumanEval and MMLU have been in public repositories, papers, blog posts and tutorials for years. Whatever filtering a lab does, the items and their solutions have been reproduced, paraphrased and discussed all over the open web. A high score can therefore reflect recall of a seen item rather than the ability the benchmark was designed to probe. The evidence for this is indirect but consistent: when researchers built fresh grade-school-maths items in the same style as GSM8K (the GSM1k work) and when they generated symbolic variants of the same problems (GSM-Symbolic), several models dropped noticeably relative to their GSM8K score — the signature of partial memorisation. ## What replaced them, and why The successor suites were built specifically to restore headroom, and they attack different axes: - **MMLU-Pro** keeps the multiple-choice knowledge format but widens the option set and filters for items that require reasoning rather than recall, pushing scores back down the range. - **GPQA-Diamond** uses graduate-level science questions written so that a non-expert with a search engine still fails them. - **ARC-AGI-2** targets abstract pattern induction that humans find easy and models find very hard, deliberately resisting a knowledge-recall shortcut. - **Humanity's Last Exam** aggregates expert-authored questions across many fields at the top of the difficulty range. - **SWE-bench Verified** (a human-validated subset of SWE-bench), **SWE-bench Pro** and **Terminal-Bench** moved the format from a single completion to multi-step agentic work in a real environment, where the score depends on end-state verification rather than string match. - **GDPval** goes further out, scoring deliverables on real occupational tasks by expert human comparison against human-produced work. Note the trend across that list: away from short recall items, toward tasks that are hard to memorise, verified on an end state, and closer to work someone would pay for. ## The practical lesson Benchmarks have a lifecycle. A new one is informative; a mature one is a summary statistic; a saturated one is a historical artefact. Two habits follow. First, when reading a score, look at where it sits in the range — a benchmark where every candidate is above ninety tells you nothing about which candidate to pick. Second, expect your own internal suite to saturate too. If every configuration you try passes 58 of 60 items, the suite has stopped doing its job and needs harder items drawn from where the system currently fails, not another round of tuning against a ceiling. The interview version of this answer is compact: saturation means no headroom, no headroom means the surviving gap is dataset noise plus memorisation, and that is why the field moved to harder, newer and more agentic suites — and why your own eval set needs the same treatment over time.
- How would you tell saturation apart from contamination when a score looks suspiciously high?Saturation shows up as every candidate model clustering near the maximum, including weak ones — the whole field is compressed. Contamination shows up as a specific model scoring far above its performance on freshly written items of the same difficulty. The discriminating test is a held-out or newly authored variant set: saturation leaves all models high, contamination leaves the affected model dropping while others hold.
- Your own 60-item internal suite is now passed almost perfectly by every candidate. What do you do?Treat it as saturated and rebuild the hard tail. Mine current production failures for items the system actually gets wrong, tighten the pass criterion where a partially correct response was scoring as a pass, and keep the old items as a cheap regression floor rather than the discriminator. The goal is to restore spread on the axis where you still need to make decisions.
- Does a saturated benchmark have any remaining use?Yes, as a floor check rather than a ranking. A frontier model scoring far below the pack on GSM8K signals something broken — a bad prompt template, a parsing bug, an over-aggressive safety filter. It is a smoke test with a known expected value, which is genuinely useful; it just cannot order the models above that floor.
saying these in an interview costs you the question
- Reads a two-point gap on a saturated benchmark as a real capability difference
- Assumes 100% is attainable and all residual errors are model errors
- Thinks saturation and contamination are the same phenomenon
- Believes a benchmark stays discriminating forever once published
- Never expects an internal eval suite to saturate too