How many few-shot exemplars are enough, and how do you find that point?
answer
- no universal number of shots
- sweep it, don't guess it
- ship the knee, not the peak
- cost is per request, forever
- re-sweep after a model change
basics
~20 sFind it empirically: sweep the shot count on a held-out set and ship the first point where added examples stop paying. Gains typically flatten after a handful, while every extra example keeps costing tokens and latency on every request.
solid answer
~50 sThere is no universal k. Treat it as a measured parameter: build a held-out set, evaluate the same prompt at several shot counts — 2, 4, 8, 16, 32 — and plot quality against k. Almost always the curve rises steeply at first and then flattens; ship the knee, not the maximum. In one severity-tagging sweep, 2-shot to 8-shot bought a large gain and 8 to 32 bought almost nothing, so eight was the answer even though thirty-two scored a hair higher. The cost side is what makes the knee the right choice: every exemplar is tokens on *every* request, plus latency and, for long sets, more prompt for the model to keep straight. Tasks with a large label space or complex output structure plateau later; simple binary tasks often plateau at two or three. Re-run the sweep when the task, the label set, or the model changes.
code
python · 9 linesdef first_plateau(scores, min_gain=0.01):
ks = sorted(scores)
for prev, cur in zip(ks, ks[1:]):
if scores[cur] - scores[prev] < min_gain:
return prev
return ks[-1]
print(first_plateau({2: 0.71, 4: 0.79, 8: 0.86, 16: 0.865, 32: 0.868}))go deeper
Know that more examples help only up to a point and that every example costs tokens on each request. Be able to say you would test a few counts rather than pick one by feel.
Walk through the sweep: fix everything but k, grow the set additively, use a doubling ladder, score on held-out data, and ship the knee with a stated noise margin.
Show that you stratify the evaluation so a rare class emerging at higher k is visible, quantify the per-request token delta between knee and peak, and re-run the sweep after model or label-set changes.
Own it as a cost-quality frontier decision across many prompts: standardize the sweep and the noise margin, and treat inherited high-shot prompts as a recurring spend to re-audit on every model migration.
## The shape of the curve Adding demonstrations to a prompt has diminishing returns, and the shape is consistent enough to plan around. The first example or two do most of the work: they establish what the output looks like and which answers are in play. The next few fill in coverage of classes and input modes. Beyond that, additional examples are mostly repeating information the set already carries, and the quality curve flattens. Occasionally it bends downward, when the extra examples add noise, contradictions or mislabeled rows. That means "how many shots?" is not a question with a memorized answer. It is a question with a measurement procedure. ## The measurement procedure 1. **Fix everything except k.** Same instruction, same template, same evaluation set. If you change the examples' wording while changing their count, the sweep tells you nothing. 2. **Choose a ladder, not a range.** Doubling — 2, 4, 8, 16, 32 — covers a wide span in few runs and makes the plateau visible. Sweeping every integer wastes evaluation budget. 3. **Grow the set additively.** The 8-shot set should contain the 4-shot set plus four new examples, so you are measuring the effect of *adding*, not the effect of a different set. 4. **Score on held-out data, stratified by class and input mode.** Aggregate accuracy hides a rare class that only becomes reachable at higher k. 5. **Read the knee.** Pick the smallest k whose score is within noise of the best score. Define "within noise" up front — a fixed margin, or a confidence interval from repeated runs — so the choice is not eyeballed. A concrete result from this procedure: a sweep at 2, 8 and 32 shots that scores roughly 0.71, 0.86 and 0.87. The right ship decision is eight. Thirty-two is four times the example tokens on every single request for a difference you cannot distinguish from noise. ## The cost side The reason to stop at the knee rather than the peak is that examples are not free, and their cost is per request, forever: - **Tokens.** Demonstrations are part of the input on every call. A 32-shot set of realistic examples can easily dwarf the actual user input. - **Latency.** More input tokens means more time before the first output token. - **Attention budget.** Very long prompts give the model more material to reconcile, and inconsistencies between distant examples become likelier as the set grows. A static exemplar set does at least sit in a stable prefix of the prompt, which is friendly to prompt-level caching where a provider offers it — but caching reduces the cost of a long set, it does not make a useless example useful. ## What moves the plateau The knee is task-dependent, and it is worth being able to predict roughly where it will land: - **Later plateau:** many labels with fine distinctions; structured or multi-field outputs; tasks where the rule is easier to show than to state; heterogeneous input formats that each need a demonstration. - **Earlier plateau:** binary or three-way classification with clear definitions; tasks the model already performs well zero-shot, where examples mainly pin down format; short, uniform inputs. If your sweep shows no gain at all from 2 to 32, the honest conclusion is usually that examples are not the bottleneck — the instruction is ambiguous, the labels are inconsistent, or the task is beyond the model. Adding shots will not fix any of those. ## Re-running the sweep The knee is not permanent. It moves when the label set changes, when the input distribution shifts, and — most sharply — when you change models. A newer or larger model often needs fewer shots for the same quality, which means a prompt inherited from an older model may be carrying twenty examples that no longer earn their tokens. Re-running the sweep after a model swap is one of the cheapest cost savings available, and it is a good answer to give when an interviewer asks how you would reduce spend on an existing prompt. ## Reporting the result When you present a k choice, present the curve, the noise margin, and the per-request token delta between the knee and the peak. "Eight shots, within 0.5 points of thirty-two, at a quarter of the example tokens" is an engineering decision. "We use eight because that felt right" is not.
- Your sweep shows no improvement from 2 shots to 32. What do you conclude?That demonstrations are not the bottleneck. The usual causes are an ambiguous instruction, inconsistent labeling in the examples or the evaluation set, or a task the model cannot do at this size. Fix the label definitions and re-check inter-rater agreement on the evaluation data first; if the task is genuinely beyond the model, more shots will never close the gap and you need a different model, a decomposition, or fine-tuning.
- Does the sweep change when you move to a newer model?Yes, usually toward fewer shots. Stronger models extract more from each demonstration and often reach the same quality at a fraction of the examples, so a prompt carrying twenty shots from an older model may be paying for fifteen it no longer needs. Re-running the sweep after any model change is cheap and frequently the largest single token saving available on an established prompt.
- How do you decide the margin that counts as 'within noise' when picking the knee?Estimate run-to-run variance first: evaluate the same configuration several times and look at the spread. Set the margin at least as wide as that spread, and be explicit about it in the write-up. With a small evaluation set, confidence intervals will be wide enough that several k values are indistinguishable — in that case pick the smallest, since the tie-break is cost.
saying these in an interview costs you the question
- Quotes a fixed number of shots as universally correct
- Ships the highest-scoring k regardless of token cost
- Changes the examples and the count in the same experiment
- Reads only aggregate accuracy across the sweep
- Never re-measures after switching models