skip to content

In few-shot prompting, do wrong labels in the demonstrations hurt accuracy?

level: middleimportance: must knowfreq 55%

answer

  1. a well-known and counterintuitive ablation result
  2. shape carries more than truth
  3. random labels from the correct set
  4. unpairing input and label hurts far more
  5. large models can still learn flipped mappings

basics

~20 s

Wrong labels hurt far less than most people expect. On classification tasks, swapping gold labels for random ones drawn from the same label set often costs only a few points, while destroying the input-label format costs much more.

solid answer

~50 s

The influential 2022 "Rethinking the Role of Demonstrations" ablation replaced gold labels in few-shot prompts with random labels from the correct label set and found accuracy degraded only marginally across many models and classification datasets. What the demonstrations mainly supply is the **format** (an input paired with a label, in a fixed shape), the **label space** (which strings count as answers) and the **input distribution** (what kind of text the task operates on). Break any of those — unpair inputs from labels, use labels outside the task's set, show unrelated inputs — and accuracy falls sharply. The nuance matters in interviews: this is a claim about tasks the model already knows how to do, where demonstrations mostly *locate* the task rather than teach it. Later work showed sufficiently large models can learn from flipped labels and override their priors, so "labels never matter" is an overstatement.

go deeper

for a junior

Know that few-shot examples teach by pattern, and that the shape of the examples matters at least as much as their individual correctness. Be able to state the surprising result plainly without overclaiming it.

for a middle

Explain the three components demonstrations supply: format, label space and input distribution. Say what happens when each is broken, and why unpairing inputs from labels is worse than noisy labels.

for a senior

Turn it into a debugging order: check output shape first, then label-set membership, then individual labels. Be ready to say where you would spend an annotation budget on a real prompt and why.

for a principal

Own the boundary conditions. Decide when a task is novel enough that exemplar labels become the specification, set the policy on exemplar review, and resist teams that either over-invest in re-annotation or cite the finding to skip quality control entirely.

## The claim, stated precisely A few-shot prompt shows the model several worked items — an input and its label — before the live query. The intuitive story is that the model reads those pairs, infers the input-to-label mapping, and applies it. If that story were the whole truth, corrupting the labels should wreck performance. A 2022 study titled "Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?" tested exactly that. It took few-shot prompts across a dozen models and roughly sixteen classification and multiple-choice datasets, then replaced each gold label with a label sampled at random from the same task's label set. Accuracy dropped only marginally — typically a small single-digit gap, and far less than the drop from removing the demonstrations entirely. The conclusion was not that demonstrations are useless, but that the *ground-truth pairing* is a smaller part of their value than the field assumed. ## What the demonstrations are actually supplying The same work isolated the components that do carry the signal: - **The format.** The fact that each item appears as an input followed by a label, in a fixed shape, with fixed field names and separators. Strip the pairing — list all inputs, then all labels — and performance collapses. The shape is what tells the model "produce one more of these". - **The label space.** The set of strings that appear in the label position. Draw labels from a foreign vocabulary (English words unrelated to the task, or a different task's classes) and accuracy falls, because the model no longer knows which output tokens are legal. - **The input distribution.** Inputs drawn from the real task distribution beat random unrelated text, because they anchor what kind of thing is being classified. Put differently: the demonstrations are mostly a **task locator**, not a training set. A frontier model has already seen sentiment classification, entity extraction and triage a million times during pretraining. The prompt's job is to pick out which of those latent behaviours to run, and in what output shape. Getting the labels individually correct is a smaller lever than making the request unambiguous. ## Where label correctness does start to matter Do not overgeneralise the result. Three boundaries come up in interviews: 1. **Model scale and prior override.** Follow-up work in 2023 on "larger models do in-context learning differently" showed that with systematically flipped labels — not random noise, but a consistent inversion — small models keep following their semantic prior while sufficiently large models will actually learn the flipped mapping from context. That is direct evidence that big models *can* read the labels; the random-label result says only that a little noise does not throw them off. 2. **Novel or arbitrary mappings.** If the task is genuinely new to the model — mapping internal product codes to routing queues, applying a house style guide that contradicts common usage — the demonstrations are the only place the mapping exists. There is no prior to fall back on, so wrong labels teach a wrong rule. 3. **Generation and reasoning outputs.** The ablation covered classification and multiple choice, where the output is one token from a small set. When the demonstrated output is a paragraph, a derivation or a structured document, the exemplar output *is* the procedure being demonstrated, and errors in it propagate directly into the answer. ## What to do with this in practice The practical reading is a reallocation of effort, not a licence to ship garbage: - **Spend first on format discipline.** Identical field names, identical separators, identical casing and spacing across every exemplar and the live query. This is the cheapest and highest-leverage fix, and it is what breaks most often when prompts are edited by several people over months. - **Spend next on the label space.** Make sure every class you want the model to be able to emit appears in the demonstration block, spelled exactly as your parser expects. - **Still verify the labels**, because a mislabeled exemplar is also a documentation defect: the next engineer reads it as the specification of the task, and the error propagates into evals, into the style guide, and into the next revision of the prompt. Reviewers and downstream humans do not enjoy the model's robustness to noise. ## Diagnosing which one is broken When a few-shot prompt underperforms, the diagnosis order follows the finding. First check whether the outputs are the *right shape* — if the model is emitting prose where you expect a label, or wrapping the answer in commentary, the format is broken and no amount of label curation will help. Then check whether the outputs are *in the label set* — off-vocabulary answers point at label-space gaps. Only after those two should you audit whether the individual demonstration labels are correct. Doing it in the reverse order is how teams spend a week re-annotating exemplars for a prompt whose real defect was an inconsistent separator.

  • So can we stop bothering to verify the labels on our exemplars?
    No. The result says a model tolerates some noise on tasks it already knows, not that labels are inert. Novel or arbitrary mappings have no prior to fall back on, generation-style outputs propagate their errors directly, and a wrong exemplar is read by every future engineer as the specification of the task. Treat clean labels as documentation hygiene even where the model shrugs them off.
  • What breaks worse than random labels, and why?
    Removing the input-label pairing. If you list all inputs and then all labels separately, the model no longer sees the atomic unit it is meant to reproduce, and accuracy falls much further than under label noise. Labels drawn from outside the task's label set are the second worst, because the model loses the vocabulary of legal answers.
  • Does this finding hold for tasks with novel label mappings the model has never seen?
    No, and that is the main boundary. The ablation covered familiar classification tasks where pretraining already supplies the mapping, so demonstrations mainly locate the task. If you are mapping internal ticket codes to routing queues, the demonstrations are the only source of that mapping. There is no prior to recover from noise, so wrong labels teach a wrong rule.

saying these in an interview costs you the question

  • Says demonstrations work like gradient training on a mini dataset
  • Claims label correctness never matters for any few-shot task
  • Confuses random labels with labels from outside the task's label set
  • Generalizes a classification ablation to long generated or reasoned outputs
  • Assumes the finding licenses shipping unaudited exemplar labels

context