In an evolutionary prompt search, what do crossover and mutation actually do?
answer
- one parent versus two parents
- mutation perturbs, crossover recombines
- prompts have natural section cut points
- selection needs elitism and diversity
- clone population by generation three
basics
~20 sMutation edits one parent prompt — rewording an instruction, swapping an exemplar, changing the output format. Crossover splices sections from two parent prompts into a child. Selection then keeps the highest scorers as the next generation's parents.
solid answer
~50 sEvolutionary search maintains a population of prompts rather than a single incumbent. **Mutation** perturbs one parent: rewrite a constraint, drop or replace an example, tighten the output spec. **Crossover** builds a child from two parents by taking some sections from each — prompts have natural cut points (role, task instruction, constraints, exemplars, output format), which is what makes recombination meaningful rather than random text splicing. Selection then keeps the top scorers, usually with elitism so the best prompt is never lost. The reason to reach for this over hill-climbing is modularity: if one parent got the recall wording right and another got the output format right, crossover can combine both wins in one step, which a single-incumbent climber would have to rediscover sequentially. The characteristic failure is premature convergence — by generation three the population is a set of near-identical prompts, and further generations buy nothing.
code
python · 12 linesimport random
SECTIONS = ("role", "instruction", "constraints", "exemplars", "format")
def crossover(parent_a, parent_b, rng=random.Random(0)):
return {s: (parent_a if rng.random() < 0.5 else parent_b)[s] for s in SECTIONS}
def mutate(prompt, rewrite, rng=random.Random(1)):
child = dict(prompt)
target = rng.choice(SECTIONS)
child[target] = rewrite(prompt[target])
return childgo deeper
Know the vocabulary: a population of prompts, mutation as an edit to one prompt, crossover as combining two, and selection as keeping the best scorers for the next round.
Explain why prompts are represented as sections so operators produce coherent children, and describe the mutation rate and selection pressure as the knobs that trade exploration against exploitation.
Show judgment about when recombination pays — separable prompt components, rugged landscape, parallel budget — and how you detect and correct premature convergence with diversity metrics, elitism and restarts.
Argue the cost case: a population method multiplies evaluation spend, so justify it against a cheaper climber and decide what improvement per thousand model calls would make you keep funding it.
## Population instead of an incumbent Hill-climbing carries one prompt and asks "can I do better than this?". Evolutionary search carries a population — say twenty prompts — and asks "which of these deserve to be parents?". The extra state is the point: several distinct lineages explore different regions of prompt space at once, and good ideas from different lineages can be merged rather than competing until one wins. ## Representing a prompt so operators are meaningful Operators only work if the representation has structure. A prompt is rarely one undifferentiated blob; it usually decomposes into a role or persona line, the task instruction, constraints and edge-case rules, few-shot exemplars, and an output-format spec. Treating those as fields gives crossover natural cut points and gives mutation a target to pick. Splicing two prompts at a random character offset produces garbage; splicing at section boundaries produces a coherent child. ## Mutation operators Mutation is a unary operator: one parent in, one variant out. Typical operators are rewording an instruction while preserving intent, tightening or loosening a constraint, adding an explicit negative rule, replacing one exemplar, reordering sections, or changing the requested output format. Two knobs matter. **Rate** — how many sections are perturbed per child — controls step size; too low and the population stagnates, too high and children stop resembling their fit parents. **Strength** — small rewording versus wholesale replacement of a section — controls whether you are exploiting a good region or jumping to a new one. Mixed operator sets, with mostly small edits and occasional large jumps, tend to behave better than any single setting. ## Crossover operators Crossover is binary: two parents in, one or two children out. Uniform crossover picks each section independently from one parent or the other; single-point crossover takes a prefix of sections from one parent and the suffix from the other. The bet crossover makes is that fitness is at least partly **separable** — that the contribution of a good constraints block is largely independent of which exemplars sit beneath it. When that bet holds, recombination is a genuine shortcut. When the prompt's parts interact strongly (an output format that only works because a specific exemplar demonstrates it), crossover produces incoherent children and mutation carries most of the progress. ## A worked setting Consider optimizing the prompt for a fraud-alert triage assistant that must label each alert as escalate, monitor or dismiss, with a one-line justification. After a few generations the population holds one parent whose constraint block is unusually good at catching the rare escalate cases, and another whose output-format block reliably produces the exact three-field structure downstream systems parse. Neither is the best overall prompt. A crossover child that takes the first parent's constraints and the second parent's format block can beat both immediately; a hill-climber holding only the better-scoring parent would have to rediscover the other's format improvement by chance. ## Selection and diversity Selection decides who reproduces. Truncation selection keeps the top-N by score; tournament selection samples a few candidates and takes the best of that sample, which gives weaker prompts a survival chance and preserves diversity. Elitism copies the single best prompt into the next generation unchanged so a lucky mutation cannot lose it. The dominant failure is **premature convergence**: strong selection pressure makes one lineage dominate within two or three generations, the population becomes near-clones, crossover between clones produces the parent again, and only mutation still generates novelty — an expensive way to run hill-climbing. Countermeasures are milder selection pressure, higher mutation rate, explicitly injecting fresh random candidates each generation, or measuring population diversity (pairwise dissimilarity of prompts) and reacting when it collapses. ## Cost Population methods are expensive: population size times evaluation items times generations, all in model calls. They are also naturally parallel — a whole generation can be scored concurrently — which makes them fit rate-limited APIs better than a strictly sequential chain of dependent steps. Against noisy scores, selection amplifies luck: a candidate that scored high on one lucky evaluation subset becomes a parent and spreads. Scoring all candidates in a generation on the *same* evaluation items makes the comparison paired and much less noisy. ## When to prefer this family Reach for evolutionary search when the prompt has several separable components you expect to improve independently, when the landscape looks rugged enough that a single climber keeps stalling, and when you have parallel budget to spend. Prefer a simpler beam or hill-climb when the prompt is short, the budget is small, or improvements are obviously sequential refinements of one idea.
- When would crossover actively hurt, and what would you do instead?Crossover assumes the prompt's sections contribute roughly independently. When they are entangled — an output-format block that only works because a particular exemplar demonstrates it — recombining parents produces children that lose the coupling and score worse than both. Symptoms are children consistently below their parents. The fix is to fall back to mutation-only search, or to define crossover over larger coherent chunks so coupled parts move together.
- How would you detect that the population has prematurely converged?Track diversity, not just the best score. Measure pairwise dissimilarity across the population — edit distance or embedding distance over prompt text — and watch it per generation. A collapsing spread with a flat best score means crossover is now producing near-copies and only mutation adds novelty. React by raising the mutation rate, softening selection pressure, or injecting fresh random candidates.
- Why score every candidate in a generation on the same evaluation items?Because it makes comparisons paired. If each candidate is scored on a different sample, part of the observed score gap is just which items it drew, and selection promotes lucky sampling rather than better prompts. Shared items cancel item difficulty from the comparison, so a smaller subset gives a reliable ranking, which matters because evaluation is the dominant cost in the loop.
Crossover is like combining two draft memos — one has the sharper argument, the other the cleaner formatting — into a single version that keeps the best half of each.
saying these in an interview costs you the question
- Describes crossover as splicing prompts at random character offsets
- Treats mutation and crossover as interchangeable operators
- Assumes a converged population means the global optimum was found
- Runs generations without elitism and loses the best prompt
- Scores each candidate on a different random sample of items