skip to content

How does the APE method (Zhou et al., 2023) discover an instruction from examples?

level: middleimportance: should knowfreq 42%

answer

  1. search over instruction space
  2. propose, score, resample
  3. the model infers the instruction
  4. examples in, instruction out
  5. a discovered step-by-step trigger beat the handwritten one

basics

~20 s

APE treats writing an instruction as search. An LLM infers candidate instructions from input-output examples of the task, each candidate is scored by actually running it on held-out data, and the best scorers are resampled into semantically similar variants until scores stop improving.

solid answer

~60 s

APE — Automatic Prompt Engineer — is the reference formulation for letting a model write its own instructions, and it has three moves. **Induction**: show an LLM a set of input-output pairs from the target task and ask it to state the instruction that would produce those outputs; a common template is "I gave a friend an instruction and several inputs… the instruction was ___". **Selection**: run each candidate instruction on a scoring set with the executing model and rank by an objective such as execution accuracy or the log-probability of the desired answer. **Resampling**: take the top candidates and ask the model for semantically equivalent variations of them, score those too, and iterate — a Monte Carlo search around the current best. The paper's headline result was that APE discovered *"Let's work this out in a step by step way to be sure we have the right answer"*, which beat the hand-written *"Let's think step by step"* on zero-shot chain-of-thought benchmarks. The lesson that survives: instruction quality is empirical, and a model plus a scorer finds phrasings humans do not.

code

markdown · 16 lines
markdown
# Induction (propose an instruction from demonstrations)
I gave a friend an instruction and several inputs. The friend read the
instruction and wrote an output for every one of the inputs.
Here are the input-output pairs:

  input: "2 apples, ate 1"      output: "1"
  input: "5 chairs, added 3"    output: "8"

The instruction was:

# Resampling (vary a surviving candidate)
Generate a variation of the following instruction while keeping the
semantic meaning.

Input: Solve the word problem and give only the final number.
Output:

go deeper

for a junior

Recall the shape: a model reads examples of the task, guesses the instruction that produced them, and the guesses are tested against data. Know that the testing, not the writing, is what makes it work.

for a middle

Explain all three moves — induction from demonstrations, selection by running candidates with the executing model, resampling of the winners — and why execution accuracy or answer log-probability is the ranking signal.

for a senior

Show that you would treat a tuned instruction as a versioned artifact pinned to a model version, question whether the scoring set is representative, and check that the instruction is actually the bottleneck before spending budget on the search.

for a principal

Own when the method is worth its data and compute at all: what it costs to build and maintain the labelled set, which prompts in a portfolio deserve optimization, and how re-tuning is scheduled against model upgrades.

## The problem APE poses Hand-tuning a prompt is a person guessing at phrasings and eyeballing a few outputs. APE (Zhou et al., 2023 — "Large Language Models Are Human-Level Prompt Engineers") reframes that as a search problem with three components: a *space* of natural-language instructions, a *proposal mechanism* that samples from it, and a *scoring function* over a dataset. The claim that made it notable was not that the search was clever — it is deliberately simple — but that an LLM is a good enough proposer that the simple search works. ## Step 1 — instruction induction The proposal step is inverse inference: given examples of the task's inputs and outputs, ask a model what instruction would have produced them. The published templates are conversational rather than technical — a forward-mode template of the shape "I gave a friend an instruction and several inputs. The friend read the instruction and wrote an output for every one of the inputs. Here are the input-output pairs: … The instruction was", and a reverse-mode variant that puts the blank in the middle for models trained to infill. This is the part that makes APE a *meta-prompting* method: the model is writing prompts, and the prompt that makes it do so is itself the object of design. Note what the induction step consumes: demonstrations, not a description. You need a labelled sample of the task, which is the same asset any honest evaluation needs. If you cannot produce twenty correct input-output pairs, you cannot run APE — and arguably you do not yet know what the task is. ## Step 2 — selection Each candidate instruction is executed on a scoring subset and ranked by an objective. The two APE reports are execution accuracy (run the instruction, compare the output against the gold label with a task-appropriate match) and log-probability of the target answer under the executing model, which gives a smoother signal on tasks where exact match is brittle. To keep cost sane the method filters aggressively: score all candidates on a small slice, discard the clearly bad ones, and only spend a large scoring budget on survivors. The non-negotiable detail is that the score comes from *running* the instruction with the model that will run it in production. An instruction that reads beautifully and scores badly is a bad instruction. This is the discipline that separates APE from prompt-writing advice. ## Step 3 — resampling The iterative variant takes the surviving high scorers and asks the model for variations that preserve meaning — a paraphrase-style meta-prompt of the shape "generate a variation of the following instruction while keeping the semantic meaning". Those variants are scored, the pool is re-ranked, and the cycle repeats. It is a local search around the current champion rather than a fresh draw from the whole space, which is why it converges quickly and why it can also converge on a local optimum. ## The result that made it famous APE's discovered zero-shot chain-of-thought trigger — "Let's work this out in a step by step way to be sure we have the right answer" — outperformed the human-written "Let's think step by step" on the arithmetic and reasoning benchmarks the paper evaluated. The finding is worth repeating in an interview not because that exact string is still the best trigger (it is model-specific and the models have all changed), but because of what it demonstrated: the mapping from phrasing to behaviour is not something human intuition reads accurately, and a scored search finds phrasings a person would not write. ## What APE assumes - **A labelled dataset exists** and is representative. Every downstream claim rests on the scoring set. - **The metric is trustworthy.** If the score is a proxy, the search will optimize the proxy; this is where most real APE deployments go wrong. - **The instruction is the bottleneck.** If the failure is missing knowledge, a bad retrieval step, or a model that cannot do the task at all, no phrasing recovers it. - **The executor is stable.** Prompts tuned against one model version can lose their edge when the model changes, so a tuned instruction is a versioned artifact with an expiry, not a permanent asset. ## Why it still matters As of mid-2026 nobody runs the 2023 code, but the formulation is the skeleton of everything that came after: propose, score against data, keep winners, iterate. When an interviewer asks about automatic prompt engineering, describing that loop and the assumptions it rests on is a stronger answer than naming a current tool, because the tools change and the loop does not.

  • What does APE need that a hand-tuned prompt does not?
    A labelled dataset. Induction needs input-output pairs to infer an instruction from, and selection needs held-out items to score candidates on. That requirement is the real cost of the method and also its main benefit — it forces you to define what correct output means before you start optimizing phrasing.
  • Why score with log-probability of the target answer rather than exact-match accuracy?
    Exact match is a step function: many candidates score identically zero, which gives the search no gradient to follow. Log-probability of the desired answer under the executing model is smooth, so a candidate that is nearly right ranks above one that is wildly wrong. The tradeoff is that it needs token-level scores and can reward confident-but-wrong phrasings.
  • A tuned instruction stops outperforming the baseline after a model upgrade. Why?
    Prompt optimization fits the phrasing to one executor's idiosyncrasies. A new model has different sensitivities, so idiosyncratic wins evaporate while the genuinely clear parts of the instruction survive. Treat a tuned prompt as a versioned artifact pinned to a model version, and re-run the search — or at least re-score the champion against the baseline — on every upgrade.

saying these in an interview costs you the question

  • Thinks APE improves a prompt without any labelled data
  • Believes the discovered instruction is universally optimal across models
  • Says the candidates are judged by reading them rather than running them
  • Confuses instruction induction with fine-tuning the model
  • Assumes better phrasing can fix a task the model cannot do at all

context