How do few-shot demonstrations signal the label space, and what happens to unseen classes?
answer
- the demonstrated strings are the output vocabulary
- instruction lists it, examples never show it
- a coverage matrix over the schema
- exact spelling, exact casing
- meaningful words beat opaque codes
basics
~20 sThe label strings appearing in the demonstrations act as the effective output vocabulary. A class named only in the instruction but never demonstrated is emitted rarely or never, so every class you want back must appear at least once, spelled exactly as your parser expects.
solid answer
~50 sDemonstrations advertise three things about the output: **which labels exist**, **exactly how they are spelled**, and roughly **how often each occurs**. The first is the one teams get wrong in production. A logistics exception classifier can list eight classes in its instruction, demonstrate six of them, and then essentially never emit `customs_hold` — the class is grammatically available but has no evidence in the pattern, so recall on the rarest and often most important class silently sits near zero. The second matters for parsing: `Customs Hold`, `customs_hold` and `CUSTOMS-HOLD` are different strings, and casing drift across exemplars produces casing drift in the output. The third is a design lever — the demonstrated mix is read as a prior over classes, so a block that is nine parts `routine` will push borderline cases toward `routine`. Audit the exemplar block as a coverage matrix against the schema, not as a pile of good examples.
go deeper
Know that the model mostly answers with labels it has actually seen in the examples, so every class you want back needs at least one example, spelled exactly the same way each time.
Explain the three signals the label position carries — membership, exact spelling, rough frequency — and why a class named only in the instruction comes back near never.
Turn it into an audit: a coverage matrix of schema classes against exemplars, re-run whenever the schema changes, plus the silent-recall failure you would look for in production metrics on rare classes.
Own the asymmetry. Decide which classes are expensive to miss, set the deliberate over-representation policy for them, and make schema changes require an exemplar update so coverage cannot rot as the taxonomy grows.
## Label space is a format property, not a content one It is tempting to think of the label as the *answer* to each demonstration and therefore a matter of correctness. Operationally it behaves more like part of the format: the set of strings that occupy the label position defines the vocabulary the model draws from when it fills that position for the live query. Ablations of in-context learning make the same point from the other side — replacing labels with strings from outside the task's label set damages accuracy far more than shuffling correct labels within the set. The set is doing work that the individual assignments are not. ## The unseen-class failure Consider a shipment-exception classifier with eight outcomes in its schema: `damaged`, `misrouted`, `delayed_weather`, `delayed_carrier`, `address_invalid`, `refused`, `lost`, `customs_hold`. The instruction lists all eight. The exemplar block, assembled from whatever real tickets were handy, contains six — the two rarest never made it in. What you observe in production is not an error. It is silence: `customs_hold` cases come back as `delayed_carrier` or `misrouted`, plausibly and confidently, and nobody notices until a downstream team asks why the customs queue is empty. The class is available in the instruction, but the demonstrations supply no evidence that it is ever the answer, and the nearest demonstrated neighbour absorbs it. This is a coverage problem with a coverage fix. Treat the exemplar block as a matrix: schema classes down one axis, exemplars across the other, and require at least one cell filled per class. That check is cheap, mechanical, and catches the defect at authoring time rather than in a quarterly audit. It also has to be re-run whenever the schema grows — adding a ninth class to the instruction without adding an exemplar for it produces exactly the same silence. ## Exact strings and the verbalizer The label position teaches spelling as well as membership. Three practical consequences: - **Casing and separators must be consistent.** If two exemplars say `customs_hold` and one says `Customs Hold`, expect both forms in the output, and expect whatever consumes the response to be the thing that discovers it. - **The words themselves matter.** The natural-language surface of a label — the *verbalizer* — carries meaning the model can use. `refused` and `address_invalid` are semantically informative; `class_4` and `class_7` are not. Opaque codes throw away the model's prior knowledge and force the demonstrations to teach the entire mapping from scratch, which is precisely the situation where exemplar labels stop being robust to noise. If your downstream system needs numeric codes, have the model emit the meaningful string and map it to the code in your own code. - **Near-collisions are a real failure source.** `delayed_weather` and `delayed_carrier` share a prefix and a concept. If the demonstrations do not make the discriminating feature visible in the input text, the model has nothing to separate them on and the split will be close to arbitrary. ## The demonstrated distribution is a prior Beyond membership, the *proportions* in the exemplar block are read as information about how often each label occurs. A block with nine `routine` examples and one `escalate` signals that escalation is rare, and borderline cases will drift toward `routine`. This is a lever, and it points in a direction that depends on your costs: if a missed `customs_hold` is expensive and a false one is cheap, an exemplar block that mirrors true production frequency — where the class is 0.5% of traffic — is the wrong design. Over-representing the rare, expensive classes relative to the wild distribution is a legitimate and common choice. The honest caveat is that the exemplar mix is a crude control. It interacts with everything else in the prompt and is not a calibrated dial, so the mix should be validated on a held-out set with the class distribution you actually care about, not reasoned about from first principles. ## Open-ended outputs Not every task has a finite label set. When the output is free text — a summary, a rewritten paragraph, an extracted list — the demonstrations still define an implicit output space: the length, the register, the vocabulary, the level of hedging, whether units are included, whether the model says "unknown" or guesses. Everything in the label-space argument transfers. The most useful special case is teaching abstention: if none of your demonstrations shows the model declining to answer, it will not decline, and you have effectively removed "I don't know" from the output vocabulary. Where a wrong confident answer is costly, at least one exemplar should demonstrate the abstention, in the exact form you want it.
- Should the exemplar mix mirror the true class frequencies in production?Not automatically. The demonstrated mix reads as a prior, so mirroring a distribution where a class is 0.5% of traffic will bury it. If a miss on that class is expensive and a false positive is cheap, over-represent it relative to the wild distribution. Treat it as a crude lever and validate on a held-out set with the distribution you actually care about.
- Does the label-space argument apply when the output is free text rather than a class?Yes. The demonstrations define an implicit output space: length, register, vocabulary, hedging, whether units appear. The sharpest case is abstention — if no exemplar shows the model declining to answer, it effectively cannot, because "I don't know" is not in the demonstrated vocabulary. Where a confident wrong answer is costly, demonstrate the abstention in the exact form you want.
- What is wrong with using opaque codes like class_4 as labels?They discard the model's prior knowledge. A meaningful verbalizer such as address_invalid lets the model connect the label to what it already knows about the concept; class_4 forces the demonstrations to teach the whole mapping from scratch, which is exactly where exemplar labels stop being robust and where you need far more of them. Emit the meaningful string and map it to your code internally.
saying these in an interview costs you the question
- Assumes listing a class in the instruction is enough for it to be emitted
- Lets label casing and spelling drift across exemplars
- Uses opaque numeric codes as labels to match a downstream schema
- Copies production class frequencies into the exemplar block without thought
- Never re-audits exemplar coverage after adding a class to the schema