How do you generate few-shot exemplar sets as prompt candidates automatically?
answer
- demonstrations are a candidate dimension too
- run the prompt, keep the traces that were right
- labels used as a filter, not as output
- the model imitates its own format best
- right answer can hide wrong reasoning
basics
~20 sRun the current prompt over training inputs, keep only the traces whose final answer matched the label, and bundle sampled subsets of those verified traces as candidate exemplar sets. The exemplar set is then searched alongside the instruction text.
solid answer
~50 sA prompt candidate is not just instruction text — the demonstrations bundled with it are a candidate dimension too, and often the higher-leverage one. The standard automatic recipe is **bootstrapping**: run the current (or a stronger teacher) prompt across labelled training inputs, capture each full trace including any intermediate reasoning, discard every trace whose final answer disagrees with the label, and treat the survivors as an exemplar pool. Candidates are then subsets sampled from that pool — different sizes, different orderings, different coverage of classes or input shapes. Self-generated demonstrations tend to beat hand-written ones because they are already in the model's own formatting and reasoning idiom, so the model imitates them cleanly. The main hazards are traces that reached the right answer by wrong reasoning, and drawing exemplars from the same split you later score on, which leaks labels and inflates the reported gain.
code
python · 15 linesdef run_prompt(prompt, raw_date):
return raw_date.replace("/", "-")
def bootstrap_exemplars(prompt, trainset, max_keep=4):
kept = []
for raw, gold in trainset:
prediction = run_prompt(prompt, raw)
if prediction == gold:
kept.append((raw, prediction))
if len(kept) == max_keep:
break
return kept
trainset = [("2026/03/04", "2026-03-04"), ("4 March 2026", "2026-03-04")]
print(bootstrap_exemplars("Normalise the date to ISO 8601.", trainset))go deeper
Know that a prompt includes its examples, and that those examples can be produced automatically by running the prompt over labelled data and keeping only the cases it got right.
Explain the bootstrap loop mechanically — run, capture the full trace, filter against the label, sample subsets as candidates — and say why self-generated demonstrations imitate better than hand-written ones.
Show the production discipline: strict train/dev/test separation so exemplars never leak labels, programmatic verification of intermediate steps where possible, and stratified sampling so classes the model already fails on are not squeezed out of the pool.
Own the tradeoff between demonstration count and permanent inference cost, decide when a teacher-model bootstrap is worth its expense versus fine-tuning, and set the policy for regenerating verified pools as input distributions drift.
## The exemplar set is a candidate, not a constant Most discussions of automatic prompt engineering picture the search moving over instruction wording. In practice a prompt shipped to production is instruction text **plus** a set of demonstrations, and on many tasks swapping the demonstrations moves the metric more than rewording the instruction does. Treating the exemplar set as a fixed input to the search — hand-written once, never varied — leaves that gain on the table. So generation has two outputs: candidate instructions and candidate exemplar sets, and a candidate prompt is a pairing of the two. ## Bootstrapping: generate demonstrations, then filter by correctness The workhorse technique is bootstrapping from your own labelled data: 1. Take labelled training inputs (inputs with gold outputs). 2. Run a prompt over them — either the current best candidate, or a stronger, more expensive "teacher" model whose traces you intend the cheaper production model to imitate. 3. Capture the **whole trace**: the input, any intermediate reasoning the model produced, and the final answer. 4. Keep only traces whose final answer matches the gold label. Discard the rest. 5. Sample subsets of the survivors as candidate exemplar sets. Step 4 is what makes this work. You are not trusting the model's output; you are using the labels as a filter, so every demonstration that reaches the prompt is verified correct end to end. For a date-normalisation task, a trace that turns `4 March 2026` into `2026-03-04` is kept; one that produces `2026-04-03` is thrown away, along with whatever reasoning led there. ## Why self-generated demonstrations often beat hand-written ones - **Format fidelity.** A generated trace already uses the model's own spacing, ordering and phrasing conventions, so the model imitates it without friction. Hand-written exemplars frequently differ in some small formatting respect that the model then reproduces inconsistently. - **Reasoning idiom.** If the prompt asks for intermediate reasoning, a bootstrapped trace shows reasoning in the shape that model actually produces, rather than a human's abbreviated version. - **Scale.** You can bootstrap hundreds of verified traces from a labelled set in one pass; nobody hand-writes hundreds. - **Teacher distillation.** Running the filter over a stronger model's traces lets a cheaper model imitate reasoning it could not reliably produce unaided — the gain shows up as demonstrations, not weights. ## What varies across exemplar-set candidates Once you have a verified pool, the candidates differ along several axes, and a search can move over all of them: - **Which traces** are included — the largest source of variation. - **How many** — more demonstrations cost tokens on every future request, and the curve flattens; the search should be allowed to pick a smaller set that ties. - **Ordering** — models are sensitive to demonstration order, particularly the last one before the query. - **Coverage** — deliberately spanning classes, input formats or difficulty bands rather than sampling uniformly, so the set does not accidentally show one label five times. ## The failure modes that matter **Right answer, wrong reasoning.** Correctness filtering checks the final answer only. A trace can arrive at the right label through a lucky guess or a bogus intermediate step, and once it is a demonstration the model imitates the bogus step. On tasks with intermediate structure, verify the reasoning too where you can do so programmatically — run the code, check the arithmetic, validate the schema — rather than trusting final-answer agreement alone. **Label leakage.** If exemplars are drawn from the same split you score candidates on, the prompt is being shown the answers to its own exam. Keep a strict split: bootstrap from train, score on dev, report on a held-out test set the search never touched. **Bias propagation.** Bootstrapping keeps what the current prompt already gets right, so it systematically over-represents the inputs the model already handles. If your model fails on one class, that class produces few surviving traces and appears rarely in candidate sets, which entrenches the failure. Stratify the pool by class or failure mode rather than sampling uniformly. **Cost.** Bootstrapping costs one model call per training input per round, on top of scoring. Cache the traces: a verified pool from an earlier round is reusable across later rounds as long as the task and the teacher have not changed. **Distribution shift.** Exemplars bootstrapped from last quarter's inputs teach last quarter's patterns. When inputs drift, the verified pool should be regenerated, not just the instruction re-tuned.
- Why bootstrap from a stronger teacher model rather than the model you will deploy?Because the demonstrations only need to be produced once, while the deployed model reads them forever. A stronger model produces correct traces on harder inputs, and after correctness filtering the cheaper model imitates reasoning it could not reliably generate itself. The caveat is idiom mismatch — if the teacher's format differs noticeably from the deployed model's own style, some of the imitation advantage is lost.
- Your bootstrapped pool has almost no traces from one class. What does that tell you?That the current prompt fails on that class, since only correct traces survive the filter. Left alone it compounds: the class stays under-represented in every candidate set, so the prompt never learns it. Stratify the pool so each class contributes exemplars, hand-write demonstrations for the starved class to seed it, or accept that the gap is a data or capability problem rather than a prompting one.
- How many demonstrations should a candidate exemplar set contain?Let the search decide rather than fixing it, but score cost as well as accuracy. Every demonstration is tokens paid on every production request, and the accuracy curve typically flattens well before the context limit. If a smaller set ties within evaluation noise, prefer it — the cheaper prompt wins on total cost of ownership even at identical quality.
saying these in an interview costs you the question
- Treating the exemplar set as fixed while only the instruction is searched
- Keeping generated demonstrations without checking them against labels
- Drawing exemplars from the same split used to score candidates
- Assuming a correct final answer means the shown reasoning is sound
- Bootstrapping uniformly and entrenching classes the prompt already fails