skip to content

Why hold out a separate test set when iterating on a prompt by hand?

level: middleimportance: must knowfreq 62%

answer

  1. you are the optimizer in this loop
  2. the set you stare at gets fitted
  3. one set to steer, one to report
  4. peeking converts test into dev
  5. few-shot pool must be disjoint

basics

~20 s

Repeated hand-tuning fits a prompt to the examples you keep looking at. A held-out test set, scored rarely and never read during tuning, is the only honest estimate of how the prompt behaves on inputs you have not seen.

solid answer

~50 s

When you edit a prompt, read the failures, and edit again, **you are the optimizer** and the eval set is your loss function. Any optimizer that sees the same data every step fits that data, so the score on the set you iterate against drifts upward faster than real quality does. The fix is the same one supervised learning uses: split the set. Take a labelled corpus — say 120 anonymised mortgage-servicing emails with gold intent labels — and split it into a dev half you may inspect item by item on every revision, and a test half you score only at a release candidate. The dev number steers; the test number is the one you report and gate on. The moment you start tuning against test, it has become a second dev set and you need a fresh holdout.

code

python · 10 lines
python
import hashlib

def bucket(item_id, dev_fraction=0.5):
    digest = hashlib.sha256(item_id.encode()).hexdigest()
    return "dev" if int(digest[:8], 16) % 100 < dev_fraction * 100 else "test"

ids = [f"email-{i}" for i in range(120)]
assigned = [bucket(i) for i in ids]
print(assigned.count("dev"), assigned.count("test"))
print(bucket("email-7") == bucket("email-7"))

go deeper

for a junior

Know the vocabulary and the rule: a dev set is what you look at while iterating, a held-out test set is what you score at the end, and quoting the tuned-on score as your quality number is the classic mistake.

for a middle

Be ready to explain the mechanism — repeated hand-tuning fits the visible examples — and to name concrete hygiene: stable hash-based splits, few-shot examples kept out of both halves, stratification, and never repairing test items to suit the prompt.

for a senior

Show that you read the dev-test gap diagnostically and decide from it: which revisions were patches, whether the split is stratified by the right unit, and when the loop has run out of signal and needs harder items rather than more edits.

for a principal

Own the policy side: who is allowed to touch the holdout, how often it may be scored, when a refresh is justified given that it destroys comparability with historical numbers, and how eval sets are versioned so decisions across teams stay comparable.

## Why a single eval set stops being honest Manual prompt iteration is a search: you write a prompt, score it on some examples, read the wrong outputs, add a clause or an example, and score again. That loop has all the parts of an optimization algorithm — a candidate, an objective, and an update rule — except that the update rule is a human reading failures. Every optimizer that is allowed to see the same data at every step will fit that data, and a human is no exception. After twenty revisions, the score on the set you have been staring at reflects two things blended together: genuine improvement, and adaptation to the particular items in front of you. You cannot separate them from inside that number. ## Dev and test, and the discipline between them The standard remedy is to split the labelled data before you start. - The **dev set** (also called the validation set) is the working surface. You may read every item, inspect individual failures, group them, and score it after every edit. Its number is a steering signal, not a claim about quality. - The **test set** is locked. You score it rarely — at a release candidate, on a schedule, or when you are about to change something structural — and you do not read its items while tuning. Its number is what you report, gate on, and put in a decision document. The discipline is not about the mechanics of splitting; it is about what you are allowed to *look at*. A test set that you score after every prompt edit has been converted into a dev set through peeking, even though you never opened the file. Each look leaks a little information about the test items into your next edit. ## The specific ways hand-tuning overfits Human overfitting does not look like gradient descent, but the mechanisms are recognisable: - **Patching the visible failures.** You see three emails misrouted and add a rule — "if the message mentions both an escrow shortage and a due-date change, classify as escrow" — that is really a description of those three emails rather than a general policy. - **Selection bias over candidates.** If you generate twenty prompt variants and keep the one with the best dev score, that winner's score is optimistically biased: the maximum of twenty noisy estimates tends to land on a lucky one, not just a good one. The more variants you compare, the larger the inflation. - **Contaminating the prompt with eval items.** Pulling few-shot examples out of the same pool you score against guarantees the model has seen the answers. Keep the few-shot pool disjoint from both dev and test. - **Repairing inconvenient items.** Deleting or re-labelling test items because the current prompt gets them wrong is the most direct form of cheating available. Genuinely mislabelled items should be fixed — but adjudicated blind, without looking at what the prompt predicted. ## Split hygiene A few practices keep the split meaningful over months: - **Split on a stable key.** Hash a natural item id to assign the bucket, so adding data later never migrates an existing item from test into dev. - **Split by the unit generalisation is about.** If ten emails come from one servicing account, they are not independent; put the whole account on one side, or the test score measures memorisation of that account's vocabulary. - **Stratify.** Preserve the class balance and, where you can, the rare-but-important cases, so that both halves can actually detect the failure modes you care about. - **Version the sets.** A score is only comparable to another score taken on the same set version; record which version produced which number. ## Reading the two numbers together A large gap between dev and test — dev 89%, test 78% — is the signal the split exists to produce. It usually means the last several revisions were fitting dev specifics, and the honest number is the lower one. The right response is to do error analysis on the *test* failures, then treat those insights as general policy, not to quietly report the dev number. A small gap that persists across revisions means your iteration is finding real improvements. When the dev set is exhausted — the prompt passes essentially everything and you can no longer tell revisions apart — the loop has run out of signal, and you need harder or fresher dev items rather than more edits. Refreshing the test set is a bigger decision: it restores an unspent peeking budget, but it breaks comparability with every historical number, so version it and re-baseline the current prompt on the new set at the same time.

  • How do you tell whether a three-point dev improvement is real?
    Compare the two prompts on the *same* items and look at what flipped, not at two independent accuracy numbers. Count items the new prompt fixed versus items it broke; if it fixed 5 and broke 4, there is no improvement regardless of the headline. Paired comparison on identical items is far more sensitive than comparing totals, because it removes the variance from item difficulty.
  • What do you do when the test score comes in far below the dev score?
    Treat the test number as the truth and diagnose the gap. Read the test failures for patterns: if they are the same failure modes the dev set already showed, your recent edits were dev-specific patches; if they are new modes, the two halves differ in distribution and the split was not stratified properly. Either way, re-derive the fix as a general rule, then re-lock the holdout.
  • How do you avoid contaminating the eval set with your few-shot examples?
    Keep a separate example pool that is never scored. If examples must come from the same corpus, carve them out before the dev/test split and mark them permanently ineligible. Otherwise every few-shot example is a leaked answer, and the eval measures copying rather than generalisation — which shows up later as a large drop the first time you run on genuinely fresh inputs.

saying these in an interview costs you the question

  • Says one large eval set is enough if it is representative
  • Reports the score from the set they tuned against
  • Draws few-shot examples from the same items they score
  • Deletes or re-labels test items the prompt got wrong
  • Treats a two-point move on a small set as a real gain

context