When picking few-shot exemplars, why must every label class appear?
answer
- examples define the answer set
- instruction alone is weak evidence
- missing class, near-zero recall
- abstain class needs a demo too
- per-class recall, not overall accuracy
basics
~20 sDemonstrations define the label space the model treats as live. A class with no example is predicted far less often than it should be, even when the instruction names it, so coverage of every label comes before adding more examples.
solid answer
~50 sFew-shot demonstrations do more than show formatting — they signal which outcomes are actually in play. A class that never appears in the examples gets predicted much less often than its true rate, even when the instruction lists it, because the examples are the strongest evidence in the prompt about what the answer set looks like. So the first rule of selection is **coverage**: at least one demonstration per label, including rare classes and any fallback outcome such as "insufficient information". The second rule is **balance**: if ten of twelve examples carry one label, ambiguous inputs drift toward that label. For a radiology-report severity tagger with five severity levels plus an "insufficient information" outcome, twelve exemplars at roughly two per outcome cover the space without letting one label dominate. Validate with per-class recall on a held-out set, never overall accuracy alone.
code
json · 13 lines{
"pool": "radiology-severity-v3",
"labels": ["routine", "minor", "moderate", "severe", "critical", "insufficient_information"],
"exemplars_per_label": {
"routine": 2,
"minor": 2,
"moderate": 2,
"severe": 2,
"critical": 2,
"insufficient_information": 2
},
"held_out_check": "per-class recall"
}go deeper
Be ready to say plainly that every label needs at least one example, including the rare ones and the fallback class, and that examples carry more weight than the instruction's label list.
Explain the mechanism: demonstrations signal which outcomes are live, so an undemonstrated class is under-predicted, and a skewed set pulls ambiguous inputs toward the majority label. Show how per-class recall exposes it.
Demonstrate that you validate with a per-class confusion matrix on held-out data, put a floor under rare-but-costly classes, and catch drift in production by comparing predicted label rates against audited ones.
Own the tradeoff between mirroring production frequencies and keeping rare classes reachable, and decide when a label space is too large for exemplars at all and should be split across staged prompts instead.
## The problem in one sentence A few-shot prompt is a set of worked examples placed before the real input. Those examples do two jobs at once: they show the shape of a correct answer, and they implicitly tell the model *which answers exist*. The second job is the one people forget, and it is why the composition of the exemplar set — which labels appear and how often — matters as much as the count. ## Coverage: every label needs a demonstration When a task has a fixed set of outcomes (a classifier, a router, a severity tagger), the instruction usually lists them in prose. The demonstrations then either confirm that list or quietly contradict it. If the instruction names six labels but the examples only ever produce four, the model has conflicting evidence, and the concrete evidence tends to win: the two undemonstrated labels are emitted at a rate far below their true frequency, and the inputs that deserve them get pushed into whichever demonstrated label is closest. The classes most often left out are exactly the ones you cannot afford to lose: - **Rare-but-costly classes.** The severe cases, the fraud cases, the safety escalations. They are rare in any sample you draw, so a set assembled by grabbing typical records misses them. - **The fallback or abstain class.** Labels like "insufficient information", "out of scope" or "needs human review" exist precisely so the model has somewhere to put inputs that do not fit. If no example ever abstains, the model learns that abstaining is not an option and forces a confident wrong answer instead. - **Newly added classes.** A label added to the instruction last week but never added to the exemplar set behaves as if it does not exist. ## Balance: how skewed can the set be? Coverage is binary; balance is a dial. A set that is technically covered but heavily skewed — say ten of twelve examples on one label — pulls ambiguous inputs toward the majority label. Roughly even counts per class are the safe default for a small set. Two refinements are common: - Weight slightly toward classes the model gets wrong, rather than toward classes that are frequent in traffic. The examples are teaching material, not a sample survey; the model does not need help with the case it already handles. - Put a floor under rare classes. Even if a class is 0.5% of traffic, give it at least one, usually two, demonstrations — one clean instance and one borderline one. There is a real tension here: matching production frequencies makes the prompt's priors realistic, while equalizing counts makes rare classes reachable. In practice, reachability wins for small sets, because a class the model never produces has a recall of zero regardless of how well-calibrated its prior is. ## A worked example A radiology-report severity tagger has five severity levels plus an "insufficient information" outcome, six outcomes in total. A twelve-exemplar set at two per outcome gives full coverage and even balance in about the token budget of a page of text. The second example per outcome is deliberately not a twin of the first: one is a clear-cut case, the other sits near the boundary with an adjacent severity. That way the set teaches both the centre and the edge of each class. After assembling it, the check is per-class: run a held-out set and read recall for each of the six outcomes separately. Overall accuracy hides the failure, because the rare severities contribute so few rows that a class with zero recall barely moves the aggregate number. ## When you have far more labels than token budget With forty or a hundred labels, one example per class is not affordable. The usual moves are: group labels into families and demonstrate the families, then resolve within a family in a second call; keep the full label list and its definitions in the instruction while spending the exemplars only on the distinctions the model actually confuses; or split the task so each prompt sees a smaller label space. What you should not do is silently drop classes from the exemplar set and hope the instruction carries them — that is the exact failure this whole discussion is about. ## What coverage does not fix Good label coverage does not make the prompt robust to every problem. It does not fix an ambiguous label definition, it does not fix a task the model genuinely cannot do, and it does not remove the need to check that each demonstration's label is actually correct. A mislabeled exemplar teaches a wrong rule with full confidence, and it is much harder to spot than a missing class.
- If a class is genuinely rare in production, should the exemplar set mirror that rarity?Usually not. Mirroring a 0.5% class means it gets zero demonstrations in a twelve-example set, and a class with no demonstration is barely predicted at all. Give rare classes a floor of one or two examples so they stay reachable, and accept that the prompt's implied prior is flatter than reality. If over-prediction of the rare class then shows up on a held-out set, tighten its definition in the instruction rather than deleting its exemplars.
- You have forty labels and a tight token budget. How do you keep coverage?Stop trying to demonstrate every label. Group the forty into a handful of families, demonstrate the families, and resolve inside the chosen family with a second prompt that sees only that family's labels. Alternatively keep full definitions in the instruction and spend the exemplars only on the pairs the model actually confuses, identified from a held-out confusion matrix. Both keep every label reachable without paying forty examples of tokens per request.
- How would you notice a missing-class problem in production rather than in evaluation?Track the predicted label distribution over time and compare it with the distribution of labels that human reviewers assign on audited samples. A class whose predicted rate sits far below its audited rate — especially one at or near zero — is the signature. It is worth alerting on, because the same signal catches a class that was added to the instruction but never added to the exemplar set.
A menu tells you what the kitchen can cook; the dishes actually on display tell you what it will cook. Diners order what they can see.
saying these in an interview costs you the question
- Assumes the instruction's label list is enough on its own
- Thinks adding more examples of common classes helps rare ones
- Duplicates one rare example many times to force balance
- Reports overall accuracy while a class has zero recall
- Leaves out the abstain or fallback class because it looks trivial