skip to content

An automated prompt search stalls immediately — how do you fix a low-diversity candidate pool?

level: seniorimportance: should knowfreq 36%

answer

  1. look at the spread, not the maximum
  2. thirty rewrites of one seed
  3. cluster count, not candidate count
  4. dedup before spending scoring calls
  5. widen the source, not the sample size

basics

~20 s

First confirm the pool is the problem: if candidate scores cluster within noise, the search had nothing to choose between. Then dedup near-identical candidates, generate from multiple seeds, demonstration subsets and proposers, and spend the freed scoring budget on genuinely distinct prompts.

solid answer

~50 s

A search that plateaus in round one usually has a generation problem, not a search problem. Diagnose it by looking at the *spread* of candidate scores: if every candidate lands within the dataset's own noise band, the pool is semantically one prompt written thirty ways, and no search strategy can rescue that. The classic cause is paraphrase-only generation — thirty rewrites of one seed instruction for a wine-recommendation classifier collapse to a handful of distinct ideas once you normalise wording. Fixes, roughly in order of payoff: dedup before scoring (exact match after normalisation, then a similarity threshold over n-grams or embeddings, keeping one per cluster); generate from several independent sources rather than one seed — induction from disjoint demonstration subsets, structured slot mutation, and paraphrase together; vary the proposer model or its temperature; and reinvest the scoring calls you saved on duplicates into the distinct candidates that remain.

code

python · 20 lines
python
def tokens(prompt):
    return set(prompt.lower().split())

def jaccard(a, b):
    ta, tb = tokens(a), tokens(b)
    return len(ta & tb) / len(ta | tb)

def dedup(candidates, threshold=0.8):
    kept = []
    for candidate in candidates:
        if all(jaccard(candidate, k) < threshold for k in kept):
            kept.append(candidate)
    return kept

pool = [
    "Classify the wine review as positive or negative.",
    "Classify this wine review as positive or negative.",
    "Decide whether the taster enjoyed the wine; answer yes or no.",
]
print(len(dedup(pool)))

go deeper

for a junior

Know that a prompt search needs candidates that actually differ from each other, and that many rewrites of one instruction are mostly the same prompt in different words.

for a middle

Be able to explain the mechanics: dedup by normalisation then similarity, generate from several sources rather than one seed, and describe why scoring duplicates wastes the budget that limits the whole run.

for a senior

Demonstrate the diagnosis. Compare score spread against measured evaluation noise, count clusters rather than candidates, and separate a generation failure from a search failure or a task that simply is not prompt-limited.

for a principal

Own the budget argument: distinct clusters per scoring call is the quantity that matters, and more samples from a collapsed generator is a spend with no expected return. Decide when to stop optimizing and invest in data, retrieval or evaluation instead.

## The symptom You run an automated prompt optimization loop. Round one produces thirty candidates, the best scores 0.71 against a seed baseline of 0.70, round two produces thirty more and the best is still 0.71. Nothing is improving. The instinct is to blame the search strategy — switch from hill-climbing to something fancier — but the far more common cause is that the pool never contained a meaningfully different prompt. ## Diagnosing pool collapse rather than search failure The cheap diagnostic is the **spread of candidate scores within a round**. Compute the standard deviation of scores across the pool, and compare it to the noise of your evaluation set (re-score the same candidate twice, or bootstrap the dev set, to estimate that noise). Two readings: - **Spread is at or below noise.** The candidates are effectively the same prompt. This is a generation problem. A better search algorithm changes nothing, because there is no signal to hill-climb on. - **Spread is well above noise but the maximum never moves across rounds.** Now it is plausibly the search or the metric — the pool has variety, and the loop is failing to exploit it. A second, more direct diagnostic: cluster the pool by similarity and count clusters. Thirty candidates that reduce to four clusters tell you exactly what happened. ## Why the pool collapsed The usual culprit is **single-seed paraphrase generation**. Asking one model for N rewrites of one instruction samples a tight semantic neighbourhood: the outputs differ in word order, politeness and length, not in strategy. A second culprit is **greedy or low-temperature proposing**, which returns the model's modal phrasing repeatedly. A third is **narrow demonstrations** — inducing every candidate from the same five examples produces instructions that all encode the same narrow reading of the task. ## Dedup before you spend a single scoring call Scoring is the expensive half of the loop: each candidate costs one model call per dev example. Scoring a duplicate buys the same answer twice. A practical three-stage dedup: 1. **Exact match after normalisation** — lowercase, collapse whitespace, strip trailing punctuation. This alone removes a surprising share. 2. **Surface similarity** — token or character n-gram Jaccard above a threshold (0.8 is a common starting point) means "same prompt, reworded"; keep one representative. 3. **Semantic similarity** — cluster by embedding cosine similarity and keep one or two per cluster. This catches candidates that share no vocabulary but say the same thing. Keep the *shortest* representative of each cluster by default: it costs less at inference for the rest of the prompt's life. ## Restoring real diversity Dedup shrinks a bad pool; it does not make a good one. To widen the pool at the source: - **Multiple generation modes.** Mix induction from labelled pairs, structured slot mutation, and paraphrase in one round rather than relying on any single mode. They fail in different directions, so their union is broader than any one of them scaled up. - **Disjoint demonstration subsets.** Induce candidate one from examples 1–5, candidate two from 6–10, and so on. Different subsets emphasise different aspects of the task and induce substantively different instructions. - **Multiple seeds.** If you must paraphrase, paraphrase three different seed instructions, not one thirty times. - **Proposer variation.** Different model families phrase instructions differently; two proposers give you two distributions instead of one. Raising temperature helps, but past a point it buys incoherence rather than diversity. - **Explicit diversity pressure.** Show the proposer the candidates already generated and ask for one that takes a *different approach* — a different persona, a different decomposition, a different output contract. This is more effective than asking for "more variety" in the abstract. ## Pool size is not the lever people think it is Doubling a collapsed pool doubles cost and adds nothing. The useful quantity is *distinct clusters per unit of scoring budget*. A pool of twelve genuinely distinct candidates will beat sixty near-duplicates, and it costs a fifth as much to evaluate. When you do have budget to grow the pool, spend it on new generation modes before spending it on more samples from the mode you already have. ## Know when the pool is not the problem If the pool is diverse, scores spread widely, and the winner still fails to beat the baseline on a held-out split, the failure is elsewhere: the task may not be prompt-limited at all, the evaluation set may be too small or too noisy to resolve the differences you care about, or the gains you saw are overfitting to the tuning split. Diversity fixes a real and common failure, but it is not a universal explanation — say so, rather than adding candidates forever.

  • How do you tell noise from a real improvement when the candidates are all close?
    Estimate evaluation noise first — re-score one candidate several times, or bootstrap-resample the dev set — and treat any gap smaller than that band as no result. Then confirm the winner on a held-out split it was never scored against. If the ranking does not survive the held-out set, the round found nothing, however confident the tuning-split numbers look.
  • Is there a downside to aggressive deduplication?
    Yes. Surface-similarity thresholds can collapse candidates that differ in one load-bearing word — a negation, a format constraint, a tie-break rule — and those small edits sometimes carry the whole gain. Keep the threshold conservative, prefer semantic clustering over raw string overlap for the final pass, and log what you dropped so a suspicious result can be traced back.
  • Would you ever deliberately keep a low-diversity pool?
    In the final rounds, yes. Once a strong prompt exists, you are polishing rather than exploring, and a tight neighbourhood of near-paraphrases is exactly the right pool for a small hill-climbing step. The mistake is using that regime at the start, when you have no reason to believe the seed is anywhere near the best available prompt.

A search algorithm is a shopper and the candidate pool is the shelf. If the shelf holds thirty jars of the same jam under different labels, no amount of careful shopping produces a different breakfast.

saying these in an interview costs you the question

  • Blaming the search algorithm when candidate scores show no spread
  • Generating thirty paraphrases of a single seed and calling it a pool
  • Scoring duplicate candidates instead of deduping first
  • Treating a larger pool as automatically a more diverse one
  • Reading tuning-split gains as real without a held-out check

context