skip to content

A new CoT prompt lifts GSM8K by 4 points — what would you check before believing it?

level: principalimportance: should knowfreq 33%

answer

  1. one run is not a measurement
  2. paired interval on identical items
  3. diff the harness before the models
  4. perturbed variants expose memorisation
  5. ceiling effects shrink real headroom

basics

~20 s

Check four things before accepting the gain: run-to-run variance with confidence intervals on paired items, harness differences such as shot count and answer extraction, contamination and saturation on a decade-old public set, and whether the gain transfers to the actual product task.

solid answer

~50 s

Treat a 4-point move on a public reasoning suite as a hypothesis. First, variance: a single run at non-zero temperature on a 1,319-item test set has a standard error of roughly one point, so re-run both arms several times, compare on the same items, and report a paired interval — many reported gains disappear here. Second, harness drift: shot count, stop sequences, and the answer-extraction rule differ between arms far more often than teams expect, and a looser extractor alone can buy several points. Third, benchmark hygiene: grade-school maths sets have been in public corpora for years, so re-run on perturbed variants that change names and numbers, and note that these suites are close to saturated for frontier models as of mid-2026, which compresses the headroom a real gain could occupy. Fourth, transfer: your product is not grade-school arithmetic. A gain that does not reproduce on an in-domain suite is not a gain you can ship.

go deeper

for a junior

Know that one benchmark run is noisy and that both arms must be scored the same way on the same items. Asking for a re-run before believing a delta is the right instinct at this level.

for a middle

Explain the mechanics: binomial standard error on a set of about 1,300 items, sampling variance at non-zero temperature, and how answer extraction or shot count can move a score independently of the model.

for a senior

Show you would design the comparison — paired runs, fixed seeds, identical harness, perturbed variants to test memorisation — and interpret where the score sits relative to ceiling before drawing conclusions.

for a principal

Own the evidence standard for the organisation: what a team must show before a prompt or model change ships, which suites are for comparability and which are for decisions, and when to retire a saturated benchmark instead of mining its last points.

## Why this question is a judgment question Nothing about a benchmark number is self-validating. The four-point claim is compatible with a real improvement, with noise, with a harness change, with memorisation, and with a metric that has run out of room. A lead's job is to know which checks separate those, in what order, and how much each is worth spending on. ## Variance before anything else A reasoning suite of about 1,300 items has a binomial standard error near one percentage point at accuracies in the eighty-to-ninety range, and sampling at non-zero temperature adds more on top. A single run per arm therefore cannot distinguish a four-point gain from a two-point gain plus luck with any confidence. The fix is cheap: several runs per arm with different seeds, comparison on identical items, and a paired interval rather than two independent means. Paired comparison matters because item difficulty dominates the variance, and pairing removes it. If the interval crosses zero, stop here. ## Harness equivalence Most surprising benchmark deltas are harness deltas. The things that quietly differ between arms: number and choice of few-shot exemplars, the exact instruction wording, stop sequences that truncate a chain mid-derivation, maximum output tokens, and above all the answer-extraction rule. A grader that accepts a bare trailing number scores higher than one that requires a fixed marker, and that difference alone can be worth several points on a chatty model. Before comparing arms, diff the harness configuration and confirm both arms are graded by the identical extractor. Re-scoring the old arm's saved generations with the new extractor is a fast way to isolate this. ## Contamination Grade-school word problems and competition maths sets have been publicly available for years and are extensively reproduced in tutorials, forums and derivative datasets, so assume some presence in pretraining corpora. Contamination inflates absolute scores and, worse, inflates them unevenly across models, making cross-model comparison unreliable. The practical detection method is perturbation: regenerate the problems from templates with different names, quantities and surface phrasing while preserving the underlying structure, and compare performance on the original items against the perturbed ones. A large drop on the perturbed variants is a memorisation signal. Published efforts along these lines — freshly written held-out sets matching the original distribution, and symbolically regenerated variants of grade-school problems — exist precisely because this drop turned out to be real for a number of models. Applying the same perturbation to both arms of your comparison also answers a narrower question: does the new prompt's advantage survive when recall is not available? ## Saturation and headroom By mid-2026 grade-school arithmetic and much of the standard competition-maths set are close to ceiling for frontier models. Two consequences follow. A four-point gain in the high nineties is a different claim from a four-point gain in the sixties — it is a large fraction of the remaining error, and it should be scrutinised harder, not less. And the remaining errors on a saturated suite are disproportionately label errors, ambiguous items and extraction quirks rather than reasoning failures, so "improvements" there often mean the prompt got better at the harness, not at reasoning. When headroom is gone, move the measurement to a harder, more recent suite rather than mining the last points from an exhausted one. ## Construct validity and transfer Even a real, contamination-free gain answers a narrow question: does this prompt help on short arithmetic word problems? A reasoning suite drawn from multi-step symbolic tasks illustrates the gap — accuracy on a shuffled-object tracking subtask can move without the traces getting any better, because the task rewards bookkeeping rather than argument. If the product is claims adjudication or incident triage, the only evidence that matters is a task suite built from your own items. Public suites are for comparability with other people's numbers; task suites are for shipping decisions. ## What a satisfying answer proposes Re-run both arms n times with fixed seeds and identical harness settings, report a paired confidence interval, re-grade with a single extractor, repeat the comparison on perturbed variants of the items, check where the score sits relative to ceiling, and require replication on an in-domain suite before the prompt goes near production. Then state the residual uncertainty plainly rather than presenting a point estimate as a result.

  • How would you check for contamination when you cannot inspect the training data?
    Perturb the items and compare. Regenerate the problems from templates with different names, numbers and phrasing but the same structure, then measure the drop between original and perturbed versions; a large asymmetric drop is a memorisation signal. Freshly authored held-out items matching the original distribution do the same job. Neither proves contamination, but a stable score across perturbations largely rules it out.
  • Why does a four-point gain near the ceiling deserve more scrutiny than the same gain mid-range?
    Because near ceiling four points is a large share of the remaining error, and the remaining errors on a saturated suite are enriched in label mistakes, ambiguous items and extraction quirks rather than genuine reasoning failures. A prompt that harvests those is optimising the harness. Mid-range, the same delta is far more likely to reflect an actual capability difference.
  • Your prompt reproduces its gain on perturbed items but not on the in-domain suite. What do you conclude?
    That the improvement is real but does not transfer to your task. That is a valid and common outcome: short arithmetic word problems exercise a narrow slice of reasoning. Do not ship on the public number. Either investigate what differs about your items — longer context, domain vocabulary, less crisp answers — or accept the prompt only where the evidence covers it.

saying these in an interview costs you the question

  • Accepts a single run per arm as a measurement
  • Compares arms graded by different answer extractors
  • Assumes a public benchmark gain transfers to the product task
  • Treats an old public suite as uncontaminated by default
  • Reports a point estimate with no interval or seed count

context