skip to content

Why can an LLM's emergent ability vanish when the benchmark metric changes?

level: seniorimportance: should knowfreq 38%

answer

  1. ask how it was measured
  2. the y-axis is doing work
  3. all-or-nothing scoring
  4. exact match hides partial progress
  5. small test sets cannot resolve rare successes

basics

~20 s

Sharp capability jumps are often produced by all-or-nothing scoring. Exact match on a multi-step task requires every step to be right, so steady per-step improvement is squeezed through a threshold and looks discontinuous. Score with partial credit and the curve is smooth.

solid answer

~50 s

An emergent ability is usually reported as a capability that is absent in smaller models and present in larger ones, with a sharp jump between them. The influential critique is that the discontinuity frequently lives in the **metric**, not in the model. Take a four-field extraction task scored by exact match on all four fields: if per-field accuracy moves 0.5 to 0.7 to 0.9 across model sizes, the all-fields score moves 0.06 to 0.24 to 0.66 — a hockey stick manufactured by a nonlinear scoring rule. Swap in per-field partial credit, edit distance, or the log-likelihood of the correct answer, and the same runs plot as a smooth trend. Small evaluation sets compound this, because you cannot resolve a true accuracy of 1% with 100 items. The critique does not prove nothing ever emerges — it means a sharp curve is a claim to interrogate, not evidence on its own.

go deeper

for a junior

Know what an emergent ability is claimed to be — absent in small models, present in large ones — and that the shape of the curve depends on how the task was scored.

for a middle

Explain the arithmetic: exact match on a multi-step task multiplies per-step accuracies together, so linear underlying progress plots as a sharp jump. Name partial credit and log-likelihood as the fixes.

for a senior

Demonstrate how you would audit the claim in practice: re-score with a graded metric, check eval-set size for resolution, confirm prompts were held constant across sizes, and check the jump reproduces.

for a principal

Own the roadmap consequence — do not plan a launch on the assumption that an ability will appear at the next scale, and separate the model's smooth competence curve from the product threshold your workflow actually requires.

## What the emergence claim says The standard framing of an emergent ability is: a capability that is not present in smaller models, is present in larger ones, and cannot be predicted by extrapolating the smaller models' performance. The plots that made this famous show near-chance accuracy across several model sizes and then a steep rise — the visual signature of something appearing rather than improving. The claim matters because it implies you cannot forecast: you must build the bigger model to discover what it can do. ## The metric-choice critique The critique is that many of these curves are artifacts of how the task is scored. Two ingredients produce a fake discontinuity. **Nonlinear scoring.** Suppose a task requires four sub-decisions and is scored by exact match — all four must be right. If per-field accuracy across three model sizes is 0.5, 0.7 and 0.9, the exact-match scores are 0.5^4 = 0.06, 0.7^4 = 0.24 and 0.9^4 = 0.66. The underlying competence rose linearly; the reported metric rose 4x then 3x and looks like an ignition. Nothing appeared. A scoring rule that multiplies probabilities together will always compress low competence towards zero and expand high competence, and the compression is what reads as "absent in smaller models". **Insufficient resolution.** If the true accuracy of the smaller model is 1% and your eval set has 100 items, you will usually record zero. Zero looks like inability. Enlarge the set or use a continuous score — the probability the model assigns to the correct completion, for instance — and the smaller model turns out to be measurably, if weakly, on the right track. The distinction between "cannot" and "can, rarely" is exactly the distinction the emergence claim depends on, and a small test set erases it. Apply either fix — partial credit per field, token edit distance, Brier score, log-likelihood of the target — and a substantial share of published emergent curves flatten into ordinary smooth trends. ## What survives the critique Three things, and a candidate who states only the debunking is answering half the question. 1. **Some abilities are genuinely non-smooth in the ways that matter.** The critique shows that many reported discontinuities are metric-induced; it does not show that none are real, and it does not cover every observed jump. 2. **All-or-nothing is often the metric the product needs.** If your pipeline routes a claim only when all four extracted fields are correct, exact match *is* your business metric. The smooth underlying curve is scientifically truer and operationally irrelevant. The honest statement is: the capability improves smoothly, and the usable threshold is crossed abruptly. 3. **Unpredictability is still partly real.** Knowing that the underlying curve is smooth lets you extrapolate the underlying quantity, but the mapping from that quantity to "does the feature work" still contains a threshold you have to locate empirically. ## How to interrogate an emergence claim When someone shows you a step-change plot, work through this: - **What is the y-axis?** Exact match, all-of-N, or any thresholded score is a warning sign. Ask for a graded score on the same runs. - **How many items?** Near-zero accuracy on a small set carries no information about whether the ability exists at low levels. - **Was the prompt held constant?** Different shot counts, formats or instructions across model sizes confound the comparison entirely. - **Is there a continuous proxy?** The probability assigned to the correct answer is available even when the sampled answer is wrong, and it usually reveals the smooth trend hiding under the step. - **Does the jump reproduce?** A single run at high sampling temperature on a small set produces plenty of accidental steps. ## Why an interviewer asks this They want to hear you read a benchmark number sceptically. The weak answer treats a plot as a fact about the model. The strong answer separates three things: the model's underlying competence, the scoring rule applied to it, and the product threshold that decides whether a feature ships. Confusing any two of those is how teams end up planning a roadmap around a capability that either was already there or never arrived.

  • How would you check an emergence claim before repeating it?
    Re-score the same runs with a graded metric — per-field partial credit, edit distance, or the log-likelihood the model assigns to the correct answer — and check the evaluation set is large enough to resolve small nonzero accuracy. Confirm the prompt, shot count and sampling settings were identical across model sizes. If the step survives all of that, it is worth taking seriously.
  • Does the critique mean no ability is ever genuinely emergent?
    No. It shows that many published discontinuities are manufactured by discontinuous metrics and small test sets, not that smooth underlying curves are universal. The defensible position is that a sharp plot is a claim requiring scrutiny rather than evidence in itself, and that the burden is on the claimant to show the jump persists under graded scoring.
  • If the underlying curve is smooth, why do product teams still see abilities switch on?
    Because products apply their own threshold. A pipeline that needs all four extracted fields correct experiences the model's smooth improvement as a step at the point where the joint success rate crosses what the workflow tolerates. The capability is continuous; the usability of it is not. Say both, and locate the threshold empirically rather than assuming it.

saying these in an interview costs you the question

  • Treats a sharp benchmark jump as proof a new capability appeared
  • Calls smaller models incapable without checking partial-credit scores
  • Claims the mirage critique proves nothing is ever emergent
  • Ignores test-set size when measuring near-zero accuracy
  • Compares model sizes with different prompts or shot counts

context