skip to content

Why do iterative prompt-rewrite loops end up with long prompts that generalize worse?

level: seniorimportance: must knowfreq 50%

answer

  1. it only ever adds, never deletes
  2. dev score up, held-out score down
  3. overfitting, in natural language
  4. clauses that name specific dev cases
  5. inert caveats nobody ablated

basics

~20 s

Each round only ever adds: caveats accumulate because nothing is deleted, and the additions encode the specific failures the rewriter was shown rather than the rule behind them. The prompt fits the development set, so held-out performance falls while the loop reports progress.

solid answer

~50 s

Prompt collapse has three compounding mechanisms. **Monotone accretion** — a rewriter asked to fix failures adds clauses and almost never removes any, so an 80-word instruction becomes a 900-word one where later clauses contradict earlier ones. **Overfitting to the shown evidence** — the rewrite encodes the exact dev items ("if the report mentions a crushed pallet, label it DAMAGE") instead of the general rule, which scores beautifully on the set it was tuned against and worse on new data. **Optimizing the scorer** — if the score is a proxy, particularly a judge model, the loop learns the phrasings the proxy rewards rather than the behaviour you wanted. The diagnostic is a widening gap between the tuning score and a held-out score the rewriter never sees, usually alongside a prompt-length curve climbing every round. Guard it by scoring on a rotating held-out split, requiring deletions or a length budget, keeping the best-scoring version rather than the latest, and ablating suspicious clauses to check they change behaviour at all.

go deeper

for a junior

Know that a prompt-rewrite loop tends to make prompts longer every round, and that a longer prompt is not automatically a better one — it has to be tested on examples it was not tuned on.

for a middle

Explain the two core mechanisms — accretion because rewriters add and never delete, and overfitting because the rewrite encodes the specific failures it was shown — and why a held-out split is what exposes both.

for a senior

Show the operating discipline: separate tuning and held-out sets, prompt-length and per-slice tracking, ablation of suspicious clauses, keeping a champion rather than the latest round, and stopping when held-out gains fall inside noise.

for a principal

Own the governance question. Decide which prompts are worth optimizing at all, who reviews and signs off on a model-written instruction, how versions are pinned to model releases and rolled back, and how you keep a metric that a loop is actively optimizing against from quietly becoming the goal.

## What collapse looks like A concrete pattern: a loop starts from an 80-word instruction scoring 0.71 on the development set. Ten rounds later the instruction runs to 900 words and scores 0.89 on dev — and 0.63 on a held-out test set it never saw, below where it started. Read the prompt and it is full of clauses like "be very careful", "double-check your answer before responding", and a list of specific cases that look suspiciously like the failures from rounds three through seven. Every individual round looked like an improvement. The artifact is worse than the thing it started from. ## Mechanism 1 — monotone accretion Ask a model to fix failures and it adds. Deletion requires deciding that an existing sentence is unnecessary, which the rewriter has no evidence for — the failing cases say what is missing, never what is redundant. So the instruction grows monotonically. Two costs follow. Later clauses start contradicting earlier ones, and the executing model has to resolve a conflict you never intended it to see. And long instructions dilute attention: the rules that actually mattered are now one paragraph among twenty, which is why a longer prompt can score worse even on the material it was tuned for. A specific sub-case worth naming: clauses that are *inert*. Successive rounds append "be very careful", "think carefully", "double-check" — each intuitively reasonable, none changing behaviour measurably. They survive because nothing tests whether they do anything. An ablation — delete the clause, re-score — usually shows the score is unchanged, and that clause was pure cost. ## Mechanism 2 — overfitting to the shown failures The meta-prompt shows failing cases so the rewriter can diagnose the general defect. What it often does instead is memorize the cases: the rewrite grows a special rule for each one. This is textbook overfitting, with a natural-language hypothesis space instead of weights, and it has the same signature — training score up, generalization down. It is aggravated by small evaluation sets. Prompt optimization is usually run against tens or low hundreds of items because scoring costs model calls, and at that size a handful of memorized rules can move the dev score several points. It is also aggravated by showing the same failures round after round, which lets successive rewrites pile rule upon rule for the same idiosyncratic items. ## Mechanism 3 — optimizing the scorer Any automated score is a proxy for what you actually want. A loop with enough rounds will find the gap. If the objective rewards long, hedged, structured answers, the loop discovers prompts that produce long, hedged, structured answers regardless of whether they are more correct. The result is a prompt that satisfies the measurement and not the goal — Goodhart's law arriving through a rewrite loop. The tell is a score that keeps improving while human spot-checks of the outputs do not. ## How you detect it - **Two-set discipline.** Score every round on a held-out split the rewriter is never shown, and plot both curves. Divergence is collapse. If the two sets track each other, the gains are real. - **Prompt length over rounds.** A curve that only rises is a warning even when scores look fine. - **Per-slice scores.** Aggregate accuracy hides the case where the prompt got better on the majority class and much worse on a minority slice that the shown failures happened not to include. - **Read the diff.** Someone should read what changed each round. Clauses naming specific dev items, or exhortations with no operational content, are visible in seconds and invisible in a metric. - **Ablation.** Delete a suspicious clause and re-score. If nothing moves, the clause was never doing work. ## How you prevent it **Hold out data the rewriter cannot see, and rotate it.** A fixed held-out set consulted every round eventually leaks through your own decisions about when to stop; rotating splits, or reserving a final set touched only once, keeps the estimate honest. **Budget the prompt.** Impose a word or token ceiling in the edit instruction, or require that each round delete at least as much as it adds. Forcing a deletion decision surfaces which clauses are load-bearing. **Keep a champion.** Track the best held-out score seen and keep that version, not the latest. Loops that always ship the last round guarantee that a bad final round becomes production. **Vary the evidence.** Sample different failures each round, cluster them, and prefer showing three distinct failure modes over ten instances of one — the rewrite then has to find the rule rather than enumerate cases. **Stop early.** Halt when held-out improvement is within noise, when the prompt exceeds its budget, or when the critiques go generic. More rounds are not free and, past the turn, are actively harmful. **Keep the human in the diff.** A model-written production prompt should be reviewed, versioned and rollback-able like code. The reviewer's job is precisely the one the metric cannot do: noticing that the instruction now encodes the dev set.

  • How would you prove a specific added clause is doing nothing?
    Ablate it: remove that clause alone, re-score the prompt on the same held-out split, and compare against the version that contains it. If the difference is inside run-to-run noise — repeat both a few times, since sampling makes single runs unreliable — the clause is inert and should be deleted. Doing this systematically after a run typically strips a large fraction of an inflated prompt.
  • The dev score keeps climbing and the held-out score is flat. Is the loop still worth running?
    No. A flat held-out curve means the extra rounds are buying dev-set fit, not capability, and the prompt is getting longer and more brittle in exchange for nothing. Stop, ship the best held-out version rather than the last, and treat the remaining gap between the two curves as a measure of how much of the apparent gain was illusory.
  • Why does a held-out set stop being honest if you consult it every round?
    Because your stopping and selection decisions leak information from it into the prompt — you are choosing the version that happens to score best on that particular sample, which is a mild form of fitting it. Rotating the split between rounds, or reserving a final set that is scored exactly once before shipping, restores an unbiased estimate.
  • Does a length ceiling risk losing genuinely necessary instructions?
    It can, which is why the ceiling should be a forcing function rather than a hard truncation: the rewriter must decide what to cut to make room. That decision is where redundancy surfaces. If the score drops when the budget binds, that is real evidence the removed content was load-bearing — evidence you never get from a loop that only appends.

Like a policy document that only ever receives amendments: every clause was added for a defensible reason, and the accumulated result is long, self-contradictory and worse than the original.

saying these in an interview costs you the question

  • Assumes a rising development score means the prompt genuinely improved
  • Never holds out data the rewriting model has not seen
  • Ships the final round's prompt rather than the best-scoring one
  • Believes adding 'be careful' and 'double-check' clauses reliably changes behaviour
  • Treats a long, exhaustively caveated prompt as evidence of thorough optimization
  • Never reads the round-by-round diff of a model-written prompt

context