skip to content

Prompt search gains 8 points on dev but none on held-out test — why?

level: seniorimportance: must knowfreq 58%

answer

  1. best-of-N is biased upward
  2. score has sampling error, selection exploits it
  3. noise does not transfer, quality does
  4. same set reused across thousands of comparisons
  5. touch the test split exactly once

basics

~20 s

Selecting the best of hundreds of candidates on one small dataset captures that dataset's noise along with real quality. The winner's score is the maximum of many noisy estimates, so it is biased upward — and the part that came from noise does not transfer.

solid answer

~50 s

This is a multiple-comparisons problem, the same one that makes best-of-N leaderboard scores optimistic. Each candidate's dev score is an estimate with real sampling error; when you evaluate 400 candidates and take the maximum, you are selecting for high true quality *and* for a lucky draw of items. With a 100-item set the standard error is around 4-5 points, and the maximum over a few hundred roughly independent estimates sits a couple of standard errors above the mean — so a large part of an 8-point gain can be selection bias with no real improvement behind it. Mitigations are structural: use a three-way split and touch the test set exactly once; make the dev set big enough that the noise floor is smaller than the gains you care about; rotate or re-shuffle validation items across rounds so no fixed subset can be memorized; limit how many candidates compete; and re-confirm the top few finalists on fresh data before declaring a winner.

code

python · 9 lines
python
import random

def best_dev_score_of_identical_prompts(n_candidates, n_items, true_acc=0.7, seed=0):
    rng = random.Random(seed)
    scores = [
        sum(rng.random() < true_acc for _ in range(n_items)) / n_items
        for _ in range(n_candidates)
    ]
    return max(scores), sum(scores) / len(scores)

go deeper

for a junior

Know that the score used to pick the winner is optimistic, and that a prompt must be checked on data the search never saw before you trust the number.

for a middle

Explain selection bias concretely: each score has sampling error, taking the maximum over many candidates picks up that error, and the effect grows with candidate count and shrinks with evaluation-set size.

for a senior

Design the loop against it: three-way split touched once, rotating folds, bounded candidate counts, finalist re-confirmation on fresh items, and a reported number that is the held-out score rather than the search's best.

for a principal

Own the reporting standard across teams — what a prompt-optimization result must state (candidates compared, items scored, held-out confirmation) before it justifies a production change or a claim of improvement.

## The mechanism A prompt-optimization loop is a selection procedure. It measures many candidates on one dataset and returns the argmax. Two things drive a candidate's measured score: its true quality on the task distribution, and the luck of which items landed in the evaluation set and how sampling fell during decoding. Selection cannot distinguish them. The winner is therefore, systematically, a candidate that is both decent *and* lucky, and its measured score is an upward-biased estimate of its true quality. The luck does not reproduce on new data; the quality does. That gap is exactly the dev-to-test drop. ## How large the bias is The scale is easy to reason about. For a binary-correct metric at true accuracy 0.7 on n items, the standard error is sqrt(0.7 × 0.3 / n): about 4.6 points at n = 100, about 3.2 at n = 200, about 1.4 at n = 1000. Now take the maximum over N roughly independent candidates. For a few hundred candidates the expected maximum of that many draws sits close to three standard errors above their common mean. At n = 100 and N = 400, a set of prompts with *identical* true quality would still produce a best-observed score more than ten points above the average one. Real candidates are correlated — they descend from the same seed and share most of their text — which shrinks the effect considerably, but it does not remove it. The practical reading is that with a small dev set and a large candidate count, a mid-single-digit "improvement" can be entirely selection bias. ## Why long searches make it worse Overfitting here compounds with rounds, not just with candidates. Each round selects survivors on the same items, so the surviving lineage accumulates whatever quirks fit those specific items — an exemplar that resolves one ambiguous case in the dev set, a constraint that helps a handful of idiosyncratic inputs. Ten generations of beam-4 search with four proposals each has performed thousands of comparisons against one fixed dataset. The dataset has effectively become training data, and the final dev score has the status of a training score. A second, subtler leak: if you keep peeking at the test set to see how a run is doing and change the search based on what you see, the test set has entered the loop. Each peek costs some of its independence, and after enough of them it is a second dev set. ## Countermeasures **Three-way split, discipline on the third.** Train/optimize on one split, select on dev, and evaluate the single chosen prompt on a test split exactly once. If you need to compare several finalists on test, expect a smaller version of the same bias and report it honestly. **Size the dev set to the effect you care about.** If a 2-point improvement matters, a 100-item set cannot see it — the noise floor swamps it. Work backwards from the smallest gain worth acting on to the number of items needed, and remember that paired comparison on identical items buys a lot of effective sample size for free. **Rotate the evaluation data.** Re-sample or re-shuffle the validation items each round, or hold several folds and rotate them. A prompt that only wins on one particular fold gets exposed; a genuinely better prompt wins across folds. This costs nothing extra beyond having more labelled data available. **Bound the number of comparisons.** Fewer, better-motivated candidates overfit less than a huge random sweep. If you cannot reduce candidates, at least budget evaluations so that only a small number of finalists get many items, which keeps the high-variance comparisons out of the final decision. **Re-confirm finalists on fresh items.** Take the top three or five by dev score and re-score them on data none of them were selected on. Rankings frequently reorder — that reordering is the bias becoming visible, and choosing after re-confirmation is materially better than trusting the search's own argmax. **Prefer robustness at equal score.** When two prompts are statistically tied, prefer the shorter, more general one over the one that accumulated many narrow rules; those rules are the most likely to be dev-set artifacts. ## What to report Report the chosen prompt's held-out score as the expected quality, not the search's best dev score, and state how many candidates competed and on how many items. A search summary without the candidate count hides the size of the multiple-comparisons problem it created.

  • Does correlation between candidates make this better or worse than the independent-draws estimate?
    Better. Candidates in a search share a seed and most of their text, so their scores move together and the effective number of independent trials is far below the raw candidate count. That shrinks the expected maximum of the noise. It does not eliminate the bias — the loop is still choosing an argmax on one dataset — so the qualitative conclusion stands even though the ten-point arithmetic is a loose upper bound.
  • How would you size the dev set before starting a search?
    Start from the smallest improvement worth acting on. For a proportion metric, the standard error is roughly sqrt(p(1-p)/n), so detecting a 2-point difference needs on the order of a thousand items unpaired. Paired comparison on identical items cuts that substantially because item difficulty cancels. If you cannot get that much labelled data, accept up front that only large gains are detectable and set the patience threshold accordingly.
  • You have only 150 labelled items in total. How do you run this responsibly?
    Keep a small untouched test slice, rotate cross-validation folds over the rest so no single split drives selection, restrict the candidate count sharply, and treat the result as a hypothesis rather than a validated win. Then spend effort on getting more labelled data, because with that little data the search's resolution — not the search algorithm — is the binding constraint.

Let a hundred people each flip ten coins and crown the one who got nine heads. Nothing about their technique transfers to the next ten flips.

saying these in an interview costs you the question

  • Reports the best dev score as expected production quality
  • Believes more candidates always give a better final prompt
  • Tunes on the test set and calls the result held-out
  • Assumes a large dev gain must reflect a real improvement
  • Adds narrow rules for individual dev items and calls it generalization

context