skip to content

When can few-shot CoT exemplars hurt accuracy on unusual inputs?

level: seniorimportance: should knowfreq 42%

answer

  1. examples teach shape, not just method
  2. out-of-distribution inputs suffer most
  3. forced into the nearest demonstrated case
  4. reasoning length mirrors the exemplars
  5. slice the eval by exemplar similarity

basics

~20 s

Worked examples teach format as well as method. When a new input does not resemble them, the model tends to force it into the demonstrated shape — mapping an unusual case onto the nearest example rather than reasoning about what is actually in front of it.

solid answer

~60 s

Exemplars act as a strong prior over the shape of the response, and that prior is only helpful where the input distribution matches. Take a veterinary intake triage prompt with four worked cases encoding the clinic's escalation ladder: each shows a symptom set, three or four lines of reasoning, and a tier. An animal arrives with a combination none of the four covers, and the model does what pattern continuation does — it picks the nearest demonstrated case, reproduces its reasoning length and structure, and outputs a tier from the demonstrated set instead of reasoning that this case is off-ladder. The reasoning looks confident and well-formed, which makes the failure hard to spot in spot checks. The tells are structural: outputs on odd inputs mirror exemplar length and section order almost exactly, and accuracy on the tail is far below accuracy on cases that resemble the examples. Fixes are an explicit escape hatch in the instruction, routing tail inputs to a zero-shot path, or an eval sliced by exemplar similarity so the tail is visible at all.

go deeper

for a junior

Know that worked examples pull the model toward the shape of those examples, so an input unlike them can get squeezed into a demonstrated answer even when it does not belong there.

for a middle

Explain the mechanism — pattern continuation carries reasoning length, structure and the observed answer set — and give a concrete symptom such as tail cases receiving exemplar-length reasoning.

for a senior

Demonstrate the diagnosis: slice evals by similarity to the exemplars, A/B a zero-shot variant on the tail, and act on the result by adding an escape hatch or routing the tail away from the few-shot prompt.

for a principal

Own the design consequence — that a few-shot prompt encodes an assumption about the input distribution, that assumption drifts, and someone must own the review cadence, the tail metric, and the human fallback that keeps the failure cheap.

## Exemplars are a prior, not an illustration It is tempting to read few-shot examples as an explanation the model reads and then sets aside. Mechanically it is closer to the opposite: the model is continuing a pattern, and everything consistent across the examples — reasoning length, section order, vocabulary, the set of answer values that appear, the confidence register — becomes part of the pattern being continued. That is precisely why few-shot CoT gives such good format consistency. It is also why it degrades in a specific, recognizable way when the input falls outside the region the examples cover. ## The failure, concretely A veterinary clinic triages intake with a prompt carrying four worked examples. Each demonstrates a symptom set, roughly four lines of reasoning, and a final tier drawn from the clinic's escalation ladder: routine, same-day, urgent, emergency. The examples were written from the four most common presentations. An animal arrives with a symptom combination that none of the four resembles — say, a chronic condition interacting with an acute complaint in a way the exemplars never demonstrate. Three things tend to happen at once: **Nearest-example capture.** The model latches onto whichever exemplar shares surface features and reuses its reasoning path, weighing the factors that mattered *there* rather than the ones that matter here. **Length and structure lock.** The reasoning comes out at roughly exemplar length. A case that genuinely needs more deliberation gets four lines, because four lines is what the pattern says an answer looks like. **Answer-space collapse.** The tier chosen is one that appears in the examples. If the right response was "this does not fit the ladder, escalate to a human", nothing in the prompt demonstrates that this is even an option. The output is fluent, correctly formatted, and confidently wrong — the worst combination for review, because format correctness is what a reviewer checks first. ## How you would know Aggregate accuracy hides this completely, because the tail is by definition a minority of traffic. The diagnostic move is to **slice the eval by similarity to the exemplars**: score cases that resemble the demonstrated presentations separately from those that do not. A large gap between the two slices is the signature. In production the proxies are cheap: output-length variance that is suspiciously low, an answer distribution that is a near-perfect match for the exemplar distribution, and a heavier concentration of human overrides on inputs that look unlike any example. A useful A/B is to run the same tail cases through a zero-shot variant. If zero-shot beats few-shot on the tail while losing on the head, you have confirmed lock-in rather than general task difficulty. ## What to do about it **Add an explicit escape hatch.** State in the instruction that the examples illustrate common cases only, that unusual combinations should be reasoned about from first principles, and that a specific out-of-ladder outcome exists. Demonstrating that outcome once is worth more than saying it, but even the instruction alone usually helps. Measure it — this is exactly the kind of change that feels effective and sometimes is not. **Route by input.** Send inputs that match known presentations to the few-shot prompt for its consistency, and send the residue to a zero-shot path that is free to reason and free to say "unclear". This concedes format consistency exactly where it was doing harm. **Cap the blast radius.** For a triage-shaped task, the cheapest robust design is that anything the model cannot place confidently goes to a person. Lock-in becomes a routing inefficiency instead of a clinical error. **Reconsider whether you need full traces.** Sometimes the value the exemplars were providing was format, not method. An instruction with a strict output template and no worked reasoning gives the format without the pull toward demonstrated reasoning paths. ## The judgment to show in an interview The strong answer does not say "few-shot is bad". It says the examples encode an assumption about the input distribution, names the symptoms that show the assumption has broken, and describes an eval that would surface it before a user does. Say plainly that head accuracy and tail accuracy are different numbers and that only one of them is usually being measured.

  • How would you detect this in production before users complain?
    Slice quality metrics by how closely each input resembles the exemplars and watch the two slices separately — a widening gap is the signal. Cheap proxies help too: unusually low variance in output length, an answer distribution that mirrors the exemplar distribution, and human overrides concentrated on inputs unlike any example. Sample those overridden cases into the eval set so the tail keeps being measured.
  • Would an instruction telling the model the examples are illustrative actually help?
    Often yes, and it is the cheapest thing to try — an explicit escape clause plus a named out-of-band outcome gives the model somewhere to go when nothing fits. But treat it as a hypothesis, not a fix: demonstrated behaviour outperforms stated behaviour, so verify on the tail slice rather than assuming. If it does not move the number, routing the tail to a separate path is the next step.
  • Does this argue for keeping the exemplar count small?
    Not directly — the problem is coverage of the input space, not count as such, and a small unrepresentative set can lock in just as hard as a large one. The relevant question is which region of the distribution your examples describe and how much traffic lives outside it. Treat exemplar composition as something driven by measured tail performance rather than by a number chosen up front.

saying these in an interview costs you the question

  • Claims more examples always improve accuracy on unusual inputs
  • Judges few-shot output quality by format correctness alone
  • Reports one aggregate accuracy number with no tail slice
  • Assumes exemplars affect content but never reasoning length
  • Thinks a well-formed confident trace means the case was understood

context