In few-shot prompting, what do the demonstrations actually teach the model?
answer
- look past the correct answers
- not just the input-output mapping
- the allowed label vocabulary and shape
- shuffled labels barely hurt accuracy
- removing demonstrations hurts far more
basics
~20 sDemonstrations mainly transmit the task framing: which outputs are legal, what the output format is, and what kind of input to expect. A well-known result is that corrupting the demonstration labels hurts accuracy far less than removing the demonstrations altogether.
solid answer
~50 sThe intuitive story is that the model infers the input-to-output function from the pairs. The evidence says that is only part of it. In experiments where demonstration labels are deliberately shuffled to the wrong classes, accuracy drops only modestly, while removing the demonstrations entirely drops it a lot — so most of the signal is carried by the **label space** (which strings are legal answers), the **output format** and the **input distribution** that tells the model what kind of task this is. Practically, that means examples are excellent at pinning an unusual vocabulary or a machine-parseable shape, and weak at teaching a genuinely novel mapping: if a rule is new, state it as an instruction rather than hoping four examples imply it. The finding is not absolute — with many demonstrations, stronger models do start to override their priors and follow a flipped mapping.
code
python · 13 linesprompt = """Classify the claim. Labels: total_loss, repairable, fraud_review.
Claim: Rear bumper cracked, airbags intact.
Label: repairable
Claim: Vehicle submerged to roofline, frame twisted.
Label: total_loss
Claim: Third collision filed this month, no police report.
Label: fraud_review
Claim: Windshield chipped by road debris.
Label:"""go deeper
Know that examples show the model what a good answer looks like — the allowed labels and the exact output shape — and that more examples is not automatically better.
Explain the shuffled-label finding: corrupting demonstration labels hurts far less than removing demonstrations, so format, label space and task framing carry most of the signal.
Show how this drives curation: cover every label and the boundary cases, keep the format byte-identical to what you parse, state new rules as instructions, and ablate the example block against a fixed eval set.
Be ready to say where the evidence is contested — enough demonstrations let stronger models override their priors — and to set a team policy of measuring example sets rather than propagating prompt folklore.
## The intuition being tested When you paste four labelled examples into a prompt, the natural assumption is that the model reads off the function: input A goes to label X, input B goes to label Y, therefore this new input goes to whichever label the pattern implies. This is the assumption an interviewer is probing when they ask what demonstrations do, because the experimental evidence complicates it. ## The shuffled-label result The well-known experiment runs three conditions on a classification task: no demonstrations at all, demonstrations with correct labels, and demonstrations whose labels have been randomly reassigned to the wrong (but still valid) classes. If demonstrations taught the mapping, the third condition should be catastrophic — you are showing the model wrong answers. What is observed instead is that random labels retain most of the benefit of correct labels, while removing demonstrations altogether costs far more. The gap between "correct labels" and "wrong labels" is small compared with the gap between "some demonstrations" and "none". ## What the demonstrations therefore carry Four channels explain the effect: - **Label space.** The set of strings that count as an answer. Show three claims labelled `total_loss`, `repairable` and `fraud_review` and the model now emits one of those three, instead of inventing "write-off" or a paragraph of prose. - **Output format.** The exact shape you will parse: one word on the line after `Label:`, a JSON object with two keys, a fixed field order. This is often the largest practical win, because a downstream parser cares about shape more than nuance. - **Input distribution.** What the incoming text looks like — terse adjuster notes rather than customer emails — which helps the model recognise which of its many learned behaviours is being asked for. - **Task framing.** The overall shape of the job, which cues a capability the model already has rather than installing a new one. The unifying view is that demonstrations **locate** a task the model can already perform, rather than teaching it one. That view also explains the earlier result: locating the task does not require the labels to be right, only for them to be present, plausible and drawn from the right vocabulary. ## Where the claim breaks down Be honest about the limits, because a good interviewer will push. - **Scale and shot count matter.** With enough demonstrations, stronger models do begin to follow a deliberately flipped mapping, overriding their own priors about what the label should be. "Labels don't matter" is an overstatement of a finding about which channel dominates at a handful of shots. - **Truly novel mappings are different.** If your task is `route to team 7 when the estimate exceeds the deductible by more than 40%`, no realistic number of examples reliably conveys the threshold. A rule that can be stated should be stated. - **Edge cases still need correct labels.** The marginal examples — the ones near a decision boundary — are exactly where the mapping does carry information, and where a wrong label teaches a wrong boundary. - **It is a permission to be deliberate, not sloppy.** Nobody should ship deliberately wrong labels. The finding tells you where to spend curation effort, not that correctness is optional. ## How to use this when curating examples 1. **Cover the label vocabulary.** Every legal output should appear at least once, so the model never has to guess whether a class exists. 2. **Make the format identical to what you parse.** Whatever the examples show is what you will get, including stray whitespace, quoting style and capitalisation. 3. **Choose examples for coverage, not for being typical.** Two boring cases plus two boundary cases beat four easy ones. 4. **Put genuinely new rules in the instructions.** Examples illustrate a rule; they are a poor channel for defining one. 5. **Keep the set small and measure.** Gains from additional shots usually saturate early on a well-specified formatting task — a sweep of zero, two, four and sixteen examples typically flattens after about four, and the extra tokens then buy nothing but cost. ## How to test it on your own task Run the ablation rather than trusting folklore: score your eval set with demonstrations removed, with labels shuffled, and with the format shown but the content replaced by placeholders. If shuffling barely moves the number, your examples are doing formatting work and you can trim them aggressively. If shuffling collapses accuracy, the mapping genuinely is being learned in context and your example choice deserves real curation effort. Either answer is useful; the mistake is not measuring.
- If labels barely matter, why not just use random ones?Because the finding is about which channel dominates, not about correctness being optional. Wrong labels still cost accuracy at the margin, they teach wrong decision boundaries on the cases nearest the boundary, and with many demonstrations stronger models do follow the mapping you show them. Use correct labels; spend the saved effort on covering the label space, the format and the edge cases.
- How would you tell whether your examples are contributing anything at all?Ablate them against a fixed eval set. Score with the examples, without them, with labels shuffled, and with format-only placeholders. The deltas tell you whether the examples are buying format compliance, mapping accuracy or nothing. If removing four examples costs less than their token price at your volume, delete them.
- What would you do if the task needs a rule the model has clearly never seen?State the rule explicitly in the instructions, define each label in words, and then use examples to illustrate the rule rather than to imply it. Demonstrations are a weak channel for a novel function, especially a numeric threshold or a policy exception. Keep one example per tricky branch so the model sees the rule applied, not just declared.
saying these in an interview costs you the question
- Assumes demonstrations work purely by teaching the input-output mapping
- Concludes that demonstration labels can be random in production
- Says more examples always improve accuracy
- Uses examples to define a brand-new rule instead of stating it
- Picks only typical examples and ignores boundary cases