In automatic prompt optimization, where do candidate prompts come from?
answer
- generation sets the ceiling, search only selects
- three sources, different diversity profiles
- infer the instruction from labelled pairs
- rewrite a seed versus vary template slots
- cold start, broaden, then polish
basics
~10 sThree sources dominate: inducing an instruction from labelled input/output pairs, resampling paraphrases of a seed instruction, and mutating slots in a fixed template. They trade diversity against staying on-task, and most systems combine them.
solid answer
~50 sCandidate generation is the proposal half of prompt optimization, and the search can only find what generation proposes. **Instruction induction** shows a proposer model a handful of labelled pairs — say five raw date strings and their ISO-8601 normalisations — and asks what instruction would have produced those outputs; sampling it repeatedly, and from different demonstration subsets, yields a varied pool with no seed prompt required. **Paraphrase resampling** takes an instruction that already works and asks for N rewrites; it is cheap and stays on task, but explores only a tight neighbourhood. **Template mutation** splits the prompt into slots — persona, task statement, constraints, output-format clause — and varies them combinatorially, which is reproducible and tells you which slot mattered, though the grid only contains variants you thought of. In practice you cold-start with induction, broaden with mutation, and polish with paraphrase.
go deeper
Know that automatic prompt optimization needs a pool of candidate prompts before anything can be scored, and be able to name where they come from: inferred from examples, paraphrased from a seed, or built by varying parts of a template.
Be ready to explain each generation mode mechanically — what the proposer is shown, what it returns, and what diversity each mode actually produces — and to say why paraphrases of one seed explore far less than induction from different demonstration subsets.
Show that you treat generation as the ceiling on the whole run: hold generation data out from scoring data, dedup before spending scoring calls, and track candidate length and format validity alongside score so you don't ship a prompt that is expensive or unparseable.
Own the decision of whether to automate at all, and the split of budget between proposing and scoring. Argue when a hybrid pipeline is worth its complexity versus a few hours of human prompt work, and what evidence would tell you the search is finding real signal rather than noise.
## Why generation is the half that decides the ceiling Automatic prompt engineering is a search loop: propose candidate prompts, score them against a dataset, keep the winners, repeat. The scoring and the search strategy get most of the attention, but they are both *selectors* — they can only return the best member of the pool that generation handed them. If every candidate says roughly the same thing in slightly different words, a perfect scorer and a perfect search still return roughly the seed prompt. Candidate generation sets the ceiling; everything downstream just finds it. ## Source 1 — instruction induction from demonstrations Here you have labelled data but no prompt. You show a proposer model a small set of input/output pairs and ask it to infer the instruction that would produce those outputs: give it five raw date strings (`2026/03/04`, `4 March 2026`, `03-04-26`) alongside their normalised `2026-03-04` forms and ask what instruction a system was following. Strengths: it needs no hand-written seed, and models routinely produce phrasings a human would not write ("rewrite each date in ISO 8601, resolving two-digit years to the 2000s") that outscore the obvious wording. Weakness: the induced instruction is only as general as the demonstrations. Five slash-separated dates induce an instruction about slashes, which then fails on the prose dates. Two mitigations matter: draw several *disjoint* demonstration subsets and induce separately from each — different subsets induce genuinely different instructions, which is diversity for free — and sample the proposer at non-zero temperature rather than greedily. ## Source 2 — paraphrase resampling You already have an instruction that works acceptably and you ask a model for thirty rewrites of it: "restate this instruction thirty different ways, preserving meaning." For a wine-recommendation classifier, that turns one seed sentence into a local neighbourhood of variants. This is the cheapest source and the safest — every candidate is still about the right task, so few are wasted on the scorer. It is also the most prone to collapse: paraphrases of one seed cluster tightly in meaning, and after dedup a pool of thirty may hold five genuinely distinct ideas. Paraphrase resampling is a polishing move, not a discovery move; treat it as generating the local neighbourhood for a hill-climbing step rather than as the whole pool. ## Source 3 — structured template mutation Decompose the prompt into named slots — persona clause, task statement, constraint list, reasoning cue, output-format clause — and generate candidates by filling those slots combinatorially, or by mutating one slot at a time while holding the rest fixed. Two real advantages. First, **attribution**: because candidates differ along known axes, the scores tell you *which slot* moved the metric, and that knowledge survives into the next task. Second, **reproducibility**: a grid is deterministic and does not cost proposer tokens. The cost is combinatorial blow-up — five slots with four options each is 1024 candidates, far more than most scoring budgets allow — so you sample the grid rather than enumerate it, or mutate greedily one slot at a time. The other limit is that a grid can only contain variants you enumerated; it will never surprise you the way induction does. ## Choosing and mixing - **No working prompt, some labelled data** → induction, several samples from several demonstration subsets. - **A decent prompt already** → template mutation for breadth, paraphrase for the final polish. - **A production loop** → seed with induction, broaden with mutation, and reserve paraphrase for the last rounds where you are shaving points. Most mature setups are hybrids, and they generate more than instruction text: the few-shot exemplar set bundled with the instruction is itself a candidate dimension. ## Failure modes to name in an interview - **Pool collapse.** Dozens of candidates, near-identical after normalising whitespace and casing. Dedup before scoring; scoring duplicates burns the budget twice for one answer. - **Length drift.** Proposer models tend to produce longer, more hedged instructions each round. Longer prompts cost more at inference forever after, so track candidate length as well as score. - **Leakage.** If the demonstrations you induce from are drawn from the same split you score on, the induced instruction can encode answers and the reported gain is fictional. Hold out the scoring split from the generation split. - **Malformed candidates.** A candidate that breaks the output contract (drops the format clause, adds commentary) scores zero for a parsing reason, not a quality reason. Validate candidates structurally before spending scoring calls on them.
- Why sample the proposer several times instead of taking one high-quality proposal?Because you are building a pool for a search, not writing a prompt. A single greedy proposal gives the search nothing to compare, and the first phrasing a model produces is rarely the best-scoring one. Sampling at non-zero temperature — and, better, inducing from different demonstration subsets — produces candidates that differ in substance rather than wording, which is what the scoring step needs to discriminate.
- How would you keep template mutation from exploding combinatorially?Don't enumerate the grid. Either sample it — draw a fixed number of random slot combinations and score those — or mutate greedily: fix all slots, vary one at a time, keep the best value for that slot, then move to the next. Greedy mutation is linear in slots rather than exponential, and it gives you a per-slot attribution report as a side effect, at the cost of missing slot interactions.
- When is candidate generation the wrong tool entirely?When the task is failing for a reason no wording fixes: missing context, a retrieval step returning the wrong documents, a model that lacks the capability, or an evaluation set that disagrees with itself. Prompt search will happily report a two-point gain that is noise. Check that a careful human prompt and a bad human prompt actually score differently before you automate the search.
saying these in an interview costs you the question
- Assuming the search can find phrasings generation never proposed
- Generating only paraphrases and calling the pool diverse
- Inducing instructions from the same split used for scoring
- Treating a bigger candidate pool as automatically a better one
- Ignoring that generated prompts steadily get longer each round