skip to content

Your 60-item eval suite shows a 4-point win — is that difference real?

level: seniorimportance: should knowfreq 36%

answer

  1. a suite is a sample, samples wobble
  2. precision grows with square root of n
  3. sixty items, plus or minus ten points
  4. pair the runs, count only flips
  5. report intervals, never bare point estimates

basics

~20 s

Almost certainly not. On 60 pass/fail items scoring around 80 percent, the 95 percent interval is roughly plus or minus 10 points, so a 4-point gap is well inside noise. Detecting it needs pairing, graded scores, or several hundred items.

solid answer

~50 s

A pass rate on 60 items has a sampling error of about 5 points in each direction one standard error, so the 95 percent interval around 80 percent spans roughly 70 to 90. A 4-point difference cannot be separated from that. Three ways out, in order of cost. **Pair the comparison**: run both versions on the same items and count only the items that flipped, which throws away the shared variance and typically cuts the requirement from over a thousand items to a few hundred. **Use a graded score** instead of binary pass/fail, since a 0-to-4 rubric carries more information per item. **Add items**, guided by an actual power calculation rather than instinct. If none of those is affordable, be honest about what the suite is: a smoke net that catches large regressions, not an instrument that can adjudicate small wins. Report intervals, never bare point estimates.

code

python · 9 lines
python
import math


def half_width(p, n, z=1.96):
    return z * math.sqrt(p * (1 - p) / n)


for n in (60, 200, 500, 1500):
    print(n, round(100 * half_width(0.80, n), 1))

go deeper

for a junior

Know that an eval score is an estimate from a sample, that small suites produce noisy numbers, and that a small difference between two runs may mean nothing at all.

for a middle

Be able to reason about the size of the noise: precision improves with the square root of the item count, so a 60-item suite carries roughly a 10-point interval and cannot settle a 4-point question.

for a senior

Show the practical toolkit — pair both versions on the same items and inspect the flips, prefer graded rubrics over binary pass/fail, and separate nondeterminism variance from sampling variance in what you report.

for a principal

Own the budget tradeoff: decide which questions justify a suite large enough to answer them, state plainly what the cheap suite can and cannot certify, and guard against the winner's curse when many variants are screened at once.

## Where the uncertainty comes from An eval suite is a sample. The number you report is an estimate of how the system would score on the population of requests the suite was drawn from, and every estimate from a finite sample carries sampling error. For a binary pass/fail metric at pass rate p over n items, the standard error is the square root of p times one-minus-p over n. At p = 0.8 and n = 60 that is about 0.052, so the 95 percent interval is roughly plus or minus 10 percentage points. Two versions scoring 78 and 82 on such a suite are indistinguishable; re-running the same two versions on a different 60-item draw could easily reverse the ordering. The uncomfortable arithmetic is that precision improves with the square root of n. Halving the interval requires quadrupling the suite. Going from plus-or-minus 10 to plus-or-minus 2.5 points means 60 items becomes 960. ## What it costs to detect a 4-point difference For two independent groups at roughly 80 percent, detecting a 4-point difference with conventional confidence and power needs on the order of 1,500 items **per version** — far beyond what most teams will hand-label. That number is the honest answer to why small suites cannot settle small questions. ## Pairing is the biggest single win You are not, however, forced into the independent-groups design. Both versions can run on the *same* items, which makes the comparison paired. Items that both versions pass, and items both fail, carry no information about which is better; only the discordant items — pass under A and fail under B, or the reverse — do. Because the shared difficulty of the items cancels out, the required sample drops sharply: if roughly 10 percent of items flip in either direction and the net difference is 4 points, a few hundred paired items suffices rather than a few thousand. The matching statistical test for paired binary outcomes is McNemar's test, and the practical version of it is simply: look at the flip counts in both directions and ask whether the imbalance is bigger than chance. Pairing has a second benefit that matters more day to day. The flipped items are a short, concrete list you can read. Twelve items that regressed tell you *what* broke; a 4-point delta tells you nothing. ## Graded metrics carry more information per item Binary pass/fail discards everything about *how* wrong an answer was. A rubric scored 0 to 4, or a continuous similarity score, has lower variance relative to the effect you are trying to detect, so the same number of items resolves smaller differences. The tradeoff is that graded scores are harder to label consistently, and inconsistency reintroduces the noise you just removed — which is why rubric clarity and sample size are the same conversation. ## Noise from the system, not just the sample Sampling error is not the only variance. Model outputs are nondeterministic, so the same item can pass on one run and fail on the next. Running each item several times and averaging reduces this, but repeated runs of the same item are not independent observations of the population — five runs of 60 items does not give you the precision of 300 items. It gives you a better estimate of *those 60 items*, which is a different and smaller improvement. When per-item variance is high, report the per-item pass proportion rather than a single run's pass/fail, and be explicit that the suite measures both quality and stability. ## Multiple comparisons If you evaluate ten prompt variants against a baseline on a small suite, the best-looking one is selected partly on noise, and its measured advantage is biased upward — the winner's curse. The same applies to slice-level reporting: with enough slices, something always looks regressed. The defenses are to pre-register which comparison decides, to confirm the winner on a fresh or held-out portion of the set, and to treat exploratory scans as hypothesis generation rather than as results. ## Living with a suite you cannot afford to grow Most teams cannot label 1,500 items per release. The realistic posture is to say clearly what the suite can and cannot do: it can catch a 15-point collapse, a broken output format, a failure concentrated in one slice — and it cannot certify a 3-point improvement. Pair every comparison, report intervals rather than bare point estimates, keep a small high-signal set for fast iteration and a larger one for decisions that matter, and reserve real statistical claims for the questions that justify the labeling budget. The failure mode this prevents is the most common one in applied LLM work: a team shipping a change on a 60-item suite, seeing the gain evaporate in production, and losing trust in evaluation altogether.

  • Why does running both prompt versions on the same items reduce the sample size you need?
    Because it removes item difficulty from the comparison. Items both versions pass, or both fail, carry no information; only the discordant ones do. That cancels the shared variance and typically cuts the requirement from thousands of items to hundreds. It also produces a readable list of exactly which items flipped, which a delta in percentage points never does.
  • You run each item five times to handle nondeterminism. Does that give you the precision of five times as many items?
    No. Repeated runs of the same items give a better estimate of those items, not a broader sample of the population, so population-level uncertainty barely moves. What they do give you is a per-item pass proportion, which separates genuine quality from instability. Both matter, but only adding distinct items narrows the interval on the overall estimate.
  • You compare ten prompt variants on one small suite and ship the best. What is wrong?
    The winner is selected partly on noise, so its measured advantage is biased upward and will shrink on re-measurement. Pre-register which comparison decides, or confirm the leading variant on a held-out portion of the set before shipping. Treat a broad scan as hypothesis generation, not as a result.

saying these in an interview costs you the question

  • The number went up, therefore the change is better
  • Sixty items is fine, they are carefully chosen
  • Repeated runs of the same items add sample size
  • Just pick whichever variant scored highest across ten trials
  • Confidence intervals are academic, we ship on the mean

context