skip to content

Should a fixed few-shot exemplar set favor diversity or typical cases?

level: middleimportance: should knowfreq 47%

answer

  1. think distribution, not pretty examples
  2. span the modes first
  3. then weight toward real volume
  4. near-duplicates cost tokens, teach nothing
  5. score per mode, not aggregate

basics

~20 s

Both, in a specific order: cover the distinct modes of the traffic you actually receive, then weight roughly toward the common ones. Cloning one typical case wastes tokens; filling the set with exotic cases teaches a distribution your users do not send.

solid answer

~50 s

When you commit to one fixed set of demonstrations, the set has to stand in for the whole input distribution, so the goal is *spanning*, not *averaging*. Start by segmenting real traffic into distinct modes — different phrasings, lengths, formats, sources — and give each mode a demonstration; that is what buys generalization to inputs you have not seen. Then weight the remaining slots toward the modes that dominate volume, because a set built entirely from exotic cases teaches a distribution your users never send and burns tokens on it. Near-duplicate examples are the clearest waste: two demonstrations that differ only in surface wording contribute one example's worth of information at two examples' cost. Verify the choice empirically on a held-out slice stratified by mode, so a gain on the common mode cannot mask a collapse on a rare one.

go deeper

for a junior

Know that examples should look like the inputs the system really receives, and that copies of one example do not help. Be able to say you would look at real traffic before choosing.

for a middle

Explain the allocation: cover distinct input modes first, then weight the rest toward volume, and identify near-duplicates as wasted slots. Mention scoring per mode on a held-out set.

for a senior

Show that you stratify evaluation by mode so a rare mode cannot collapse unnoticed, that you ablate to find redundant examples, and that you refresh the set when the traffic distribution shifts.

for a principal

Own the decision to stay with a static set at all — weighing its reviewability, versioning and reproducibility against dynamic alternatives — and set the policy for how mode coverage is defined and re-derived across teams.

## Why the question exists A fixed few-shot set is a lossy summary of your entire input distribution, squeezed into a handful of examples. Two instincts pull in opposite directions. One says: show the model the most typical case, repeatedly, so the pattern is unmistakable. The other says: show the model the widest possible variety, so nothing surprises it. Neither instinct is right on its own, and interviewers ask about it because the reasoning reveals whether you think about the input distribution at all. ## Spanning beats averaging The useful reframing is coverage of *modes*. Real traffic is rarely one smooth blob; it is a handful of clusters. A contract-clause tagger, for example, sees clauses that arrive as clean single sentences, as multi-paragraph blocks with nested sub-clauses, as OCR output with broken line breaks, and as fragments quoted mid-sentence in an email. Those are four modes, and a model that has only seen the first will handle the other three worse — not because the task changed, but because the demonstrations never showed the task being done under those conditions. So the first pass over the exemplar budget is: one demonstration per distinct mode. This is where diversity earns its keep. The generalization you get is not magical breadth; it is specifically the breadth you demonstrated. ## Then weight toward what actually arrives Once every mode has a representative, spend remaining slots proportional to volume — with a floor for modes that are rare but expensive to get wrong. This is where similarity to the expected query distribution matters. An exemplar set stuffed with unusual, ambiguous, adversarial cases has two costs: - **Token cost with no return.** Exotic cases tend to be long, and they occupy the budget that a common case would have used. - **A distorted implied prior.** Demonstrations are evidence about what kinds of input the model is about to see. If they are all edge cases, the model treats ordinary input as if it were subtle, and starts over-thinking or over-hedging on the 90% of traffic that is straightforward. ## Near-duplicates are the real waste The most common defect in hand-built exemplar sets is not too much diversity or too little — it is redundancy. Someone assembles eight examples by picking eight rows that all looked good, and five of them are structurally identical. Those five cost five examples' worth of tokens and latency and contribute roughly one example's worth of information. A cheap discipline: before adding an example, ask what it teaches that the set does not already teach. If the honest answer is "nothing, it's just another one of those", drop it and spend the slot on an uncovered mode. ## How to actually decide The selection is an empirical question, so make it one: 1. Pull a sample of real inputs and label the modes — by source, format, length, or whatever axis actually varies. 2. Build a held-out evaluation slice that is *stratified* by mode, with enough rows per mode to read a number. 3. Try candidate exemplar sets and compare per-mode scores, not just the aggregate. The aggregate is dominated by the common mode and will happily hide a rare mode collapsing to zero. 4. Keep the set that wins on the worst mode without losing meaningfully on the common one. ## The static-set constraint is doing real work Everything above assumes you must commit to one set that is used for every request. That constraint is what makes the diversity-versus-similarity trade a trade at all: if you could choose demonstrations per request, the question would be posed differently and would come with its own operational costs, which is a separate technique. There are good reasons to stay static — a fixed set is simpler to review, version, and reason about, and it keeps the prompt identical across requests, which matters for reproducibility and for anything downstream that assumes a stable prompt. So "just pick per query" is not automatically the better answer, and an interviewer will expect you to justify the choice rather than assume it. ## What the answer sounds like when it is wrong A weak answer says "more diverse is always better" with no reference to what the traffic looks like, or picks examples by which ones read nicely rather than by what they cover. A strong answer starts from the observed input distribution, treats the exemplar budget as slots to allocate, and ends with a measurement on a stratified held-out set.

  • How would you detect that your fixed set is redundant rather than diverse?
    Ablate. Remove one example at a time and re-score on a stratified held-out set; an example whose removal changes nothing is contributing nothing and can be replaced. Structurally, sort candidate examples by the axes you know vary — source, length, format, label — and look for slots where several examples share every value. Those clusters are the redundancy, and each one is a free slot for an uncovered mode.
  • When is a deliberately narrow, homogeneous exemplar set the right call?
    When the input distribution really is narrow — a single upstream system emitting one format — or when the task is a strict transformation where the only thing the examples need to convey is the mapping rule. Extra variety there adds tokens and can introduce inconsistencies between examples. The test is empirical: if per-mode scores are already flat and high, breadth has nothing left to buy.
  • Your traffic distribution shifts after a product launch. What changes about the exemplar set?
    New modes appear — different phrasing, new document types, new intents — and the old set no longer spans them. Re-sample recent traffic, re-derive the modes, and check per-mode scores on the new slices before deciding. Often the fix is small: swap out one or two now-unrepresentative examples rather than rebuilding. The important part is having the stratified evaluation set refreshed too, or you will be measuring the old world.

saying these in an interview costs you the question

  • Says more diverse examples are always better, with no reference to traffic
  • Picks exemplars by how good they read rather than what they cover
  • Fills the set with near-identical rows and calls it consistency
  • Judges candidate sets on aggregate accuracy only
  • Assumes edge cases should dominate because they are hard

context