A prompt search shows no gain for six rounds — how do you decide it has converged?
answer
- flat curve can be noise, not convergence
- compare on identical items, paired
- local optimum versus real ceiling
- patience threshold above the noise floor
- read the failures before another round
basics
~20 sSix flat rounds is a plateau only if the flatness survives measurement. Re-score the incumbent and its challengers on the same evaluation items, look at the paired difference and its uncertainty, and only then decide between stopping, restarting from a new seed, or widening the operators.
solid answer
~50 sConvergence in a noisy black-box search is a statistical claim, not an observation that the best score stopped moving. First separate signal from noise: score the incumbent and challengers on **identical** evaluation items so comparisons are paired, and put an interval around the difference — with a couple of hundred items, a two- or three-point gap is often inside sampling noise. Then ask which kind of plateau you have. A **local optimum** means your operators cannot reach anything better, and the answer is a random restart from a different seed prompt, a larger perturbation, or a broader operator set. A **ceiling** means the remaining errors are not fixable by prompting — ambiguous labels, missing context, genuine model capability limits — and you see it by reading the failing cases, not by running more rounds. In practice you run a patience rule (stop after N rounds with no gain above a threshold) plus a hard budget cap, and confirm the winner on data the loop never touched.
go deeper
Know that scores wobble between runs, so a flat round or two does not mean the search has finished, and that a stopping rule and a budget cap are set in advance.
Explain paired evaluation on identical items, why it reduces comparison noise, and how a patience-based early-stopping rule with a threshold above the noise floor is configured.
Diagnose the plateau: distinguish an operator-limited local optimum from a data or metric ceiling by reading failures, and choose between restart, wider operators, and stopping. Confirm the winner on untouched data.
Frame it as spend allocation: decide when further prompt search is worth less than fixing labels, adding retrieval, or changing model, and define the marginal-value rule that ends a run across the team's projects.
## Why "the score stopped moving" is not enough Every score in an automatic prompt-engineering loop is an estimate. It carries variance from decoding randomness, from judge or scorer variability, and above all from the finite evaluation set. With 200 items and accuracy near 0.7, the standard error of a single measurement is roughly 3 percentage points, so two prompts whose true quality is identical will routinely appear 4 or 5 points apart, and a genuinely better prompt can look flat for a round or two. Declaring convergence from a flat best-score curve therefore conflates two very different situations: the search has genuinely run out of reachable improvements, or it is producing improvements too small to see through the noise floor. ## Step one: measure the difference, not the scores The cheapest large win is **paired evaluation**. Score the incumbent and every challenger on exactly the same items, then compare per-item outcomes rather than aggregate percentages. Item difficulty cancels out of the comparison, so the variance you are fighting drops sharply and a much smaller evaluation set gives a reliable ranking. From the paired outcomes you can compute a simple interval — a bootstrap over items, or a sign test over items where the two prompts disagree — and answer the actual question: is the best challenger's advantage distinguishable from zero at this sample size? If the interval straddles zero, you do not have evidence of a plateau; you have evidence that your evaluation resolution is coarser than the improvements on offer. The options are to increase items per comparison for the top candidates only, or to accept that further gains are below the resolution you are willing to pay for and stop deliberately. ## Step two: classify the plateau If the flatness is real, diagnose which kind it is. **Local optimum from the operator set.** The neighbourhood you defined cannot reach anything better. This is the common case for greedy search with conservative operators — every proposal reweighs the same phrasing. The tell is that all recent proposals are minor variants of the incumbent. Remedies: a random restart from a different seed prompt, a deliberately large perturbation of the incumbent, a broader operator set that can restructure sections rather than reword them, or carrying a population instead of a single incumbent so several regions stay alive. **Task or data ceiling.** The remaining failures are not addressable by prompting. Read fifty failing items by hand and bucket them. If a large share are mislabelled gold answers, ambiguous cases where reasonable humans disagree, or items that require context the prompt was never given, no rewording will fix them and the measured ceiling is a property of the dataset, not the prompt. This diagnosis is worth more than another thousand model calls, and it redirects effort to data curation, retrieval, or a different model. **Scorer saturation.** The metric may already be near its maximum on the items the loop sees, so improvements land in a region the score cannot express. That again ends the useful life of this loop rather than justifying more rounds. ## Step three: encode the decision as a stopping rule Do not decide plateau by eye in production runs. Standard practice combines several rules and stops on whichever fires first: - **Patience**: stop after N consecutive rounds without an improvement exceeding a threshold (a threshold larger than the noise floor you measured, not zero). - **Hard budget cap**: a maximum number of rounds, model calls or dollars, set before the run. - **Marginal-value rule**: stop when the improvement per unit of spend falls below what the improvement is worth to you. - **Restart policy**: rather than stopping outright, spend the remaining budget on a fresh restart; several independent short runs frequently beat one long one on a rugged landscape, and they also give you a spread of outcomes rather than a single point. ## Step four: confirm the winner honestly The best dev score at the end of a long search is an optimistic number, because the winner was chosen by that same set. Confirm the chosen prompt on data the loop never scored against, once, and report that number. If the confirmation score is much lower than the search score, the plateau debate was moot: the recent rounds were fitting the evaluation set rather than improving the prompt. ## What a strong answer sounds like Name the noise floor and how you measured it, describe paired comparison as the first move, distinguish local optimum from ceiling using failure inspection, state the concrete stopping rule you would configure, and say what you would do with the remaining budget — restart, widen operators, or stop and reinvest in data.
- How do you set the improvement threshold in a patience rule?Anchor it to measured noise, not to zero. Score one fixed prompt several times on the evaluation set and observe the spread, or bootstrap over items to get a standard error. Set the threshold at roughly that scale — improvements smaller than your measurement resolution are not something the loop can reliably chase. Then pick patience so that the expected cost of the extra rounds stays inside your budget.
- Why might several short restarts beat one long run at the same total budget?On a rugged, noisy landscape a single greedy run commits early to whichever lineage got lucky and then spends the rest of the budget polishing inside one basin. Independent restarts sample several basins, so the best-of-restarts outcome is usually higher and you also get a variance estimate across runs. The cost is losing the deep refinement that a long run can achieve inside a genuinely good region.
- What tells you the ceiling is in the data rather than the prompt?Manual inspection of the residual failures. If a large fraction are mislabelled references, genuinely ambiguous items, or cases needing information the prompt never receives, then no wording change can score them correctly and the metric's apparent ceiling belongs to the dataset. That finding redirects work to label adjudication or to supplying the missing context, which is usually a bigger win than more search.
saying these in an interview costs you the question
- Declares convergence because the best score stopped moving
- Compares candidates on different random samples of items
- Keeps running rounds without inspecting any failing cases
- Sets the improvement threshold at zero, so noise counts as progress
- Reports the final search score as the expected production quality