skip to content

Exemplar Selection

Which examples actually earn their tokens: coverage of edge cases, class balance, difficulty mix, and how many shots before returns diminish. Interviewers probe this because badly chosen demonstrations quietly teach the model the wrong rule and the failure looks like a model problem.

part ofPrompt engineeringoverview, primer and where to startread it →
on this pageshow

questions

5

When picking few-shot exemplars, why must every label class appear?

level: juniorimportance: must knowfreq 58%

answer

  1. examples define the answer set
  2. instruction alone is weak evidence
  3. missing class, near-zero recall
  4. abstain class needs a demo too
  5. per-class recall, not overall accuracy

basics

~20 s

Demonstrations define the label space the model treats as live. A class with no example is predicted far less often than it should be, even when the instruction names it, so coverage of every label comes before adding more examples.

solid answer

~50 s

Few-shot demonstrations do more than show formatting — they signal which outcomes are actually in play. A class that never appears in the examples gets predicted much less often than its true rate, even when the instruction lists it, because the examples are the strongest evidence in the prompt about what the answer set looks like. So the first rule of selection is **coverage**: at least one demonstration per label, including rare classes and any fallback outcome such as "insufficient information". The second rule is **balance**: if ten of twelve examples carry one label, ambiguous inputs drift toward that label. For a radiology-report severity tagger with five severity levels plus an "insufficient information" outcome, twelve exemplars at roughly two per outcome cover the space without letting one label dominate. Validate with per-class recall on a held-out set, never overall accuracy alone.

code

json · 13 lines
json
{
  "pool": "radiology-severity-v3",
  "labels": ["routine", "minor", "moderate", "severe", "critical", "insufficient_information"],
  "exemplars_per_label": {
    "routine": 2,
    "minor": 2,
    "moderate": 2,
    "severe": 2,
    "critical": 2,
    "insufficient_information": 2
  },
  "held_out_check": "per-class recall"
}

go deeper

for a junior

Be ready to say plainly that every label needs at least one example, including the rare ones and the fallback class, and that examples carry more weight than the instruction's label list.

for a middle

Explain the mechanism: demonstrations signal which outcomes are live, so an undemonstrated class is under-predicted, and a skewed set pulls ambiguous inputs toward the majority label. Show how per-class recall exposes it.

for a senior

Demonstrate that you validate with a per-class confusion matrix on held-out data, put a floor under rare-but-costly classes, and catch drift in production by comparing predicted label rates against audited ones.

for a principal

Own the tradeoff between mirroring production frequencies and keeping rare classes reachable, and decide when a label space is too large for exemplars at all and should be split across staged prompts instead.

## The problem in one sentence A few-shot prompt is a set of worked examples placed before the real input. Those examples do two jobs at once: they show the shape of a correct answer, and they implicitly tell the model *which answers exist*. The second job is the one people forget, and it is why the composition of the exemplar set — which labels appear and how often — matters as much as the count. ## Coverage: every label needs a demonstration When a task has a fixed set of outcomes (a classifier, a router, a severity tagger), the instruction usually lists them in prose. The demonstrations then either confirm that list or quietly contradict it. If the instruction names six labels but the examples only ever produce four, the model has conflicting evidence, and the concrete evidence tends to win: the two undemonstrated labels are emitted at a rate far below their true frequency, and the inputs that deserve them get pushed into whichever demonstrated label is closest. The classes most often left out are exactly the ones you cannot afford to lose: - **Rare-but-costly classes.** The severe cases, the fraud cases, the safety escalations. They are rare in any sample you draw, so a set assembled by grabbing typical records misses them. - **The fallback or abstain class.** Labels like "insufficient information", "out of scope" or "needs human review" exist precisely so the model has somewhere to put inputs that do not fit. If no example ever abstains, the model learns that abstaining is not an option and forces a confident wrong answer instead. - **Newly added classes.** A label added to the instruction last week but never added to the exemplar set behaves as if it does not exist. ## Balance: how skewed can the set be? Coverage is binary; balance is a dial. A set that is technically covered but heavily skewed — say ten of twelve examples on one label — pulls ambiguous inputs toward the majority label. Roughly even counts per class are the safe default for a small set. Two refinements are common: - Weight slightly toward classes the model gets wrong, rather than toward classes that are frequent in traffic. The examples are teaching material, not a sample survey; the model does not need help with the case it already handles. - Put a floor under rare classes. Even if a class is 0.5% of traffic, give it at least one, usually two, demonstrations — one clean instance and one borderline one. There is a real tension here: matching production frequencies makes the prompt's priors realistic, while equalizing counts makes rare classes reachable. In practice, reachability wins for small sets, because a class the model never produces has a recall of zero regardless of how well-calibrated its prior is. ## A worked example A radiology-report severity tagger has five severity levels plus an "insufficient information" outcome, six outcomes in total. A twelve-exemplar set at two per outcome gives full coverage and even balance in about the token budget of a page of text. The second example per outcome is deliberately not a twin of the first: one is a clear-cut case, the other sits near the boundary with an adjacent severity. That way the set teaches both the centre and the edge of each class. After assembling it, the check is per-class: run a held-out set and read recall for each of the six outcomes separately. Overall accuracy hides the failure, because the rare severities contribute so few rows that a class with zero recall barely moves the aggregate number. ## When you have far more labels than token budget With forty or a hundred labels, one example per class is not affordable. The usual moves are: group labels into families and demonstrate the families, then resolve within a family in a second call; keep the full label list and its definitions in the instruction while spending the exemplars only on the distinctions the model actually confuses; or split the task so each prompt sees a smaller label space. What you should not do is silently drop classes from the exemplar set and hope the instruction carries them — that is the exact failure this whole discussion is about. ## What coverage does not fix Good label coverage does not make the prompt robust to every problem. It does not fix an ambiguous label definition, it does not fix a task the model genuinely cannot do, and it does not remove the need to check that each demonstration's label is actually correct. A mislabeled exemplar teaches a wrong rule with full confidence, and it is much harder to spot than a missing class.

  • If a class is genuinely rare in production, should the exemplar set mirror that rarity?
    Usually not. Mirroring a 0.5% class means it gets zero demonstrations in a twelve-example set, and a class with no demonstration is barely predicted at all. Give rare classes a floor of one or two examples so they stay reachable, and accept that the prompt's implied prior is flatter than reality. If over-prediction of the rare class then shows up on a held-out set, tighten its definition in the instruction rather than deleting its exemplars.
  • You have forty labels and a tight token budget. How do you keep coverage?
    Stop trying to demonstrate every label. Group the forty into a handful of families, demonstrate the families, and resolve inside the chosen family with a second prompt that sees only that family's labels. Alternatively keep full definitions in the instruction and spend the exemplars only on the pairs the model actually confuses, identified from a held-out confusion matrix. Both keep every label reachable without paying forty examples of tokens per request.
  • How would you notice a missing-class problem in production rather than in evaluation?
    Track the predicted label distribution over time and compare it with the distribution of labels that human reviewers assign on audited samples. A class whose predicted rate sits far below its audited rate — especially one at or near zero — is the signature. It is worth alerting on, because the same signal catches a class that was added to the instruction but never added to the exemplar set.

A menu tells you what the kitchen can cook; the dishes actually on display tell you what it will cook. Diners order what they can see.

saying these in an interview costs you the question

  • Assumes the instruction's label list is enough on its own
  • Thinks adding more examples of common classes helps rare ones
  • Duplicates one rare example many times to force balance
  • Reports overall accuracy while a class has zero recall
  • Leaves out the abstain or fallback class because it looks trivial

context

open as a page

How many few-shot exemplars are enough, and how do you find that point?

level: middleimportance: must knowfreq 56%

basics

~20 s

Find it empirically: sweep the shot count on a held-out set and ship the first point where added examples stop paying. Gains typically flatten after a handful, while every extra example keeps costing tokens and latency on every request.

open as a page

Should a fixed few-shot exemplar set favor diversity or typical cases?

level: middleimportance: should knowfreq 47%

basics

~20 s

Both, in a specific order: cover the distinct modes of the traffic you actually receive, then weight roughly toward the common ones. Cloning one typical case wastes tokens; filling the set with exotic cases teaches a distribution your users do not send.

open as a page

Why add negative exemplars showing what a classifier must not flag?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Positive-only demonstrations teach where the rule applies but never where it stops, so the model over-generalizes and flags look-alikes. Near-miss negatives — cases that resemble a hit but are not one — draw the boundary the positives leave undefined.

open as a page

Who owns a few-shot exemplar pool, and when must it be refreshed?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

An exemplar pool is policy encoded as data, so it needs a named owner in the domain team, provenance and PII review on every entry, versioning alongside the prompt, and a refresh triggered by policy changes, label-set changes and drift — not by the calendar alone.

open as a page