skip to content

Automatic Prompt Engineering

Letting a search loop write the prompt: generate candidates, score them against a dataset, keep the winners, repeat — up to compiled prompt programs like DSPy. Interviewers ask about it to see whether you would optimize prompts systematically instead of hand-tuning them forever.

on this pageshow

explore

questions

22

In prompt optimization, what does a meta-prompt that rewrites another prompt contain?

level: juniorimportance: must knowfreq 45%

answer

  1. a prompt about a prompt
  2. four slots, not free-form prose
  3. current text plus task spec
  4. failure evidence is load-bearing
  5. delimited output, no commentary

basics

~20 s

A meta-prompt is a prompt whose subject is another prompt. It carries four things: a specification of the task, the current prompt verbatim, concrete evidence of how that prompt failed, and an instruction to emit a revised prompt in a fixed format.

solid answer

~50 s

A meta-prompt asks a model to do prompt engineering rather than the task itself, so everything it needs must be in the text — it cannot see your eval harness. Four slots do the work: **the task specification** (what the prompt is supposed to achieve, and what a correct output looks like); **the current prompt**, quoted verbatim inside delimiters so the model does not confuse it with the surrounding instructions; **failure evidence** — a handful of concrete cases showing input, the output the current prompt produced, and the expected output; and **an edit instruction plus an output contract**, e.g. "rewrite the instruction so these cases come out right without breaking the ones that pass; output only the new instruction between `<prompt>` tags". A typical shape is to feed the five worst-scoring outputs from the last evaluation run. Without the failure evidence the model has nothing to diagnose and will simply pad the prompt with generic caution.

code

markdown · 21 lines
markdown
# Task
Classify a shipping-incident report as DELAY, DAMAGE, LOST or OTHER.
Output exactly one label, uppercase, nothing else.

# Current instruction
<current>
Read the report and say what kind of incident it is.
</current>

# Failing cases (5 worst-scoring from the last run)
1. input: "Pallet arrived crushed on one corner, contents wet."
   got: "It looks like the shipment was damaged in transit."
   expected: "DAMAGE"
2. input: "Scanned in Rotterdam on the 3rd, no events since."
   got: "DELAY"
   expected: "LOST"

# Your job
Rewrite the instruction so these cases come out right without breaking
cases that already pass. Prefer removing ambiguity over adding caveats.
Output ONLY the new instruction between <prompt> and </prompt>.

go deeper

for a junior

Be able to say plainly that a meta-prompt is a prompt whose input is another prompt, and name the four things it must carry: the task, the current prompt, real failures, and the edit instruction.

for a middle

Explain why the failure evidence is what makes the rewrite targeted rather than decorative, and why the output format has to be constrained so a loop can extract the candidate automatically.

for a senior

Show judgment about which failures to sample and how many, why passing cases are worth including, and that a rewrite is a candidate to be scored with the production executor, never an accepted improvement.

for a principal

Own the question of who reviews model-written prompts before they ship: how candidates are versioned, diffed and rolled back, and where the boundary sits between an automated rewrite and a human-owned instruction.

## What a meta-prompt is A meta-prompt is a prompt whose subject matter is another prompt. The model reading it is not performing the task; it is performing prompt engineering *on* the task prompt and returning a new candidate. This is the atom of automatic prompt engineering: whatever search or scoring machinery sits around it, the actual writing of a new prompt happens inside one meta-prompt call. The critical property is that the rewriting model is blind. It cannot see your evaluation harness, your dataset, your score, or the production traffic the prompt failed on. Anything it should reason about has to be serialized into the meta-prompt text. Most weak meta-prompts fail for exactly this reason: they say "improve this prompt" and give the model no basis on which to decide what "improve" means, so it falls back on generic writing advice — more structure, more caveats, more politeness — none of which is grounded in the task's actual failures. ## The four slots **1. Task specification.** State what the prompt is for, what the inputs look like, what a correct output looks like, and any hard constraints (output format, forbidden content, length). Without this the rewriter infers the task from the prompt it is being asked to fix — which means it inherits that prompt's misunderstanding. **2. The current prompt, verbatim and delimited.** Quote it inside an unambiguous fence so the model can tell the material-under-edit apart from the instructions addressed to it. Meta-prompts confuse models when the two run together; the model then "helpfully" answers the task prompt instead of rewriting it. **3. Failure evidence.** Concrete cases, each showing the input, what the current prompt actually produced, and what was expected. This is the load-bearing slot. A rewrite grounded in "on these three refund questions the model answered from memory instead of quoting the policy text" produces a targeted edit; a rewrite grounded in nothing produces prose. **4. Edit instruction and output contract.** Say what kind of edit you want ("fix these failures without regressing the passing cases", "prefer deleting an ambiguous sentence over adding a new one", "keep it under 200 words") and how to return it ("output only the new instruction between `<prompt>` and `</prompt>`, no commentary"). The contract exists so the surrounding loop can extract the candidate mechanically instead of parsing an essay. ## Choosing the evidence Show a small, diverse sample — commonly the worst-scoring handful from the last evaluation round, deduplicated so five near-identical items do not dominate. Showing everything is both expensive and counterproductive: the more specific cases you paste, the more the rewrite tends to encode those exact items rather than the rule behind them. Some practitioners also include one or two *passing* cases, labelled as such, so the rewrite has something to avoid breaking. ## Failure modes of the meta-prompt itself - **"Make it better" with no evidence.** The model appends adjectives. Nothing measurable changes. - **No output contract.** The model returns a rewrite wrapped in explanation, and the loop stores the explanation as the prompt. - **Task spec and current prompt merged.** The model answers the task instead of editing. - **Asking for a rewrite and an explanation in one turn without separating them.** Fine if you parse the delimiters; a bug source if you do not. - **Accepting the rewrite unscored.** A meta-prompt produces a *candidate*, never a verified improvement. It must be run against the same evaluation as the incumbent before it replaces anything. ## The rewriter is not necessarily the executor The model that rewrites and the model that runs the resulting prompt need not be the same. A stronger rewriter can write instructions a weaker executor cannot reliably follow, so scoring must always use the executor that will run in production. Conversely, a rewriter that shares the executor's blind spots may keep proposing edits that read well and change nothing. ## What sits around it One meta-prompt call is a single edit. Turning it into an optimization method requires deciding how candidates are proposed at scale, how they are scored, and how the search over them proceeds — separate concerns with their own machinery. The meta-prompt is where the field's actual leverage lives, though: the difference between a loop that improves a prompt and one that inflates it is usually visible in the four slots above.

  • How many failing cases would you paste into one meta-prompt, and why not all of them?
    A handful — typically five to ten, sampled across distinct failure clusters and deduplicated. Beyond that you pay context cost for little diagnostic gain, and the rewrite starts encoding the specific items shown rather than the rule behind them. Diversity matters more than volume: five different ways the prompt breaks is far more useful than fifty instances of the same break.
  • Why insist that the rewritten prompt come back inside delimiters with no commentary?
    So the loop can extract the candidate mechanically. Without a contract the model returns "Here's an improved version, I've clarified the tone…" and a naive extractor stores that preamble as part of the prompt. Delimiters also let you detect a malformed response and retry rather than silently corrupting the candidate pool.
  • Should the model doing the rewriting be the same one that will run the prompt?
    Not necessarily, but the scoring must use the executor. A stronger rewriter often writes instructions a weaker executor cannot follow, so a candidate that looks excellent is only real if it scores well when the production model runs it. Using the same model for both is simpler and avoids that mismatch, at the cost of inheriting its blind spots.

saying these in an interview costs you the question

  • Thinks a meta-prompt is just asking the model to make the prompt better
  • Omits failing examples and expects the model to guess what broke
  • Pastes the entire evaluation set into the meta-prompt
  • Accepts the rewritten prompt without re-scoring it against the incumbent
  • Confuses a meta-prompt with a system prompt for the task itself

context

open as a page

In automatic prompt optimization, where do candidate prompts come from?

level: middleimportance: must knowfreq 48%

basics

~10 s

Three sources dominate: inducing an instruction from labelled input/output pairs, resampling paraphrases of a seed instruction, and mutating slots in a fixed template. They trade diversity against staying on-task, and most systems combine them.

open as a page

Why is automatic prompt optimization run as a discrete, gradient-free search?

level: middleimportance: must knowfreq 62%

basics

~20 s

A prompt is discrete text, not a continuous parameter vector, so no gradient points toward a better prompt. Optimization instead runs as black-box search: propose edited candidates, score each on a dataset, keep the winners, repeat under a fixed budget.

open as a page

In DSPy, what does a signature declare, and why isn't it just a prompt string?

level: middleimportance: must knowfreq 46%

basics

~20 s

A DSPy signature declares one step's named inputs and outputs plus a short description of the task — for example context, question -> answer. The wording, formatting rules and examples that make up the actual prompt are generated by the framework, not written by you.

open as a page

In automatic prompt optimization, how do you pick the scoring metric for a task?

level: middleimportance: must knowfreq 70%

basics

~20 s

Match the scorer to the output shape: exact match or F1 when one label is right, execution-based checks when the output can be run and verified, overlap scores only as a rough proxy, and a judge model only for genuinely open-ended text.

open as a page

Why do iterative prompt-rewrite loops end up with long prompts that generalize worse?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Each round only ever adds: caveats accumulate because nothing is deleted, and the additions encode the specific failures the rewriter was shown rather than the rule behind them. The prompt fits the development set, so held-out performance falls while the loop reports progress.

open as a page

Prompt search gains 8 points on dev but none on held-out test — why?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Selecting the best of hundreds of candidates on one small dataset captures that dataset's noise along with real quality. The winner's score is the maximum of many noisy estimates, so it is biased upward — and the part that came from noise does not transfer.

open as a page

A prompt search shows no gain for six rounds — how do you decide it has converged?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Six flat rounds is a plateau only if the flatness survives measurement. Re-score the incumbent and its challengers on the same evaluation items, look at the paired difference and its uncertainty, and only then decide between stopping, restarting from a new seed, or widening the operators.

open as a page

In DSPy, what does compiling a program with BootstrapFewShot actually produce?

level: seniorimportance: must knowfreq 40%

basics

~20 s

Compiling returns a new copy of the program whose predictors now carry selected demonstrations (and, with instruction optimizers, rewritten instruction text). No model weights change. That prompt state is the artifact, and it can be saved to JSON and loaded later.

open as a page

When can a small proxy eval set replace the full gold set for scoring prompt candidates?

level: seniorimportance: must knowfreq 54%

basics

~20 s

Only for screening. A cheap subsample can rank candidates roughly and cull the obviously bad ones, but promotion decisions and reported numbers must come from the full gold set, because a moderately correlated proxy reorders candidates exactly where they are closest.

open as a page

Why can accuracy mislead when scoring prompt candidates on an imbalanced defect set?

level: juniorimportance: should knowfreq 50%

basics

~20 s

Accuracy is dominated by the majority class. On a manufacturing set that is 2% defective, a candidate that never flags a defect still scores 98%, so the optimizer ranks it top while it fails completely on the class you built the system for.

open as a page

How do you generate few-shot exemplar sets as prompt candidates automatically?

level: middleimportance: should knowfreq 40%

basics

~20 s

Run the current prompt over training inputs, keep only the traces whose final answer matched the label, and bundle sampled subsets of those verified traces as candidate exemplar sets. The exemplar set is then searched alongside the instruction text.

open as a page

How does the APE method (Zhou et al., 2023) discover an instruction from examples?

level: middleimportance: should knowfreq 42%

basics

~20 s

APE treats writing an instruction as search. An LLM infers candidate instructions from input-output examples of the task, each candidate is scored by actually running it on held-out data, and the best scorers are resampled into semantically similar variants until scores stop improving.

open as a page

Why does a prompt-rewrite loop critique the prompt before editing it?

level: middleimportance: should knowfreq 35%

basics

~20 s

Splitting diagnosis from editing forces the model to state, in writing, why the prompt failed. That written reason constrains the next edit to the observed defect, makes the change reviewable by a human, and gives a stopping signal when critiques go generic.

open as a page

In an evolutionary prompt search, what do crossover and mutation actually do?

level: middleimportance: should knowfreq 42%

basics

~20 s

Mutation edits one parent prompt — rewording an instruction, swapping an exemplar, changing the output format. Crossover splices sections from two parent prompts into a child. Selection then keeps the highest scorers as the next generation's parents.

open as a page

In DSPy, what changes when you swap Predict for ChainOfThought on a step?

level: middleimportance: should knowfreq 34%

basics

~20 s

Swapping the module keeps the same signature but changes how that step is executed: ChainOfThought extends the declared outputs with a reasoning field the model fills before answering. Calling code, training data and metric are untouched — you change one line and recompile.

open as a page

An automated prompt search stalls immediately — how do you fix a low-diversity candidate pool?

level: seniorimportance: should knowfreq 36%

basics

~20 s

First confirm the pool is the problem: if candidate scores cluster within noise, the search had nothing to choose between. Then dedup near-identical candidates, generate from multiple seeds, demonstration subsets and proposers, and spend the freed scoring budget on genuinely distinct prompts.

open as a page

How do you detect prompt candidates that game the scorer instead of solving the task?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Compare the optimized score against a held-out set the search never touched, and read the top candidates' actual outputs. A score that climbs on the search set while flat or falling on held-out data is the signature of a candidate exploiting the scorer.

open as a page

How do you spend a fixed evaluation budget across candidates in a prompt search?

level: principalimportance: should knowfreq 40%

basics

~20 s

Spend unevenly. Score every candidate cheaply on a small shared subset, eliminate the clearly weak ones, and reallocate the saved evaluations to survivors — successive halving. Budget in model calls and tokens, not candidate counts, and reserve capacity to confirm finalists.

open as a page

How do you handle a base-model upgrade for a compiled DSPy program in production?

level: principalimportance: should knowfreq 28%

basics

~20 s

Treat the compiled artifact as a build output pinned to the model it was compiled against. On an upgrade, first evaluate the existing artifact on the new model against a held-out set, then recompile and compare all three options — old artifact, new artifact, and uncompiled — before deciding what to ship.

open as a page

Your prompt search scores every candidate with a judge model — how do you keep that affordable and trustworthy?

level: principalimportance: should knowfreq 38%

basics

~20 s

Treat the judge as expensive infrastructure: cascade cheap programmatic checks in front of it so it only ranks survivors, pin its model version and rubric for the whole campaign so scores stay comparable across generations, and re-validate it against human labels on a fixed anchor set.

open as a page

In automated prompt search, how do you choose the proposer model and candidate budget?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

Generation is usually the cheap half: proposing N prompts costs N calls, while scoring them costs N times the evaluation-set size. That asymmetry argues for a strong proposer and a small, well-chosen pool, with the optimized prompt always validated on the model that will actually serve it.

open as a page