skip to content

Few-Shot Prompting

Teaching a task by showing examples instead of describing it: zero-shot versus one-shot versus few-shot, and which examples to pick, how to order them, and how to format their labels. Interviewers probe it because badly chosen demonstrations quietly teach the model the wrong rule.

part ofPrompt engineeringoverview, primer and where to startread it →
on this pageshow

questions

17

Why must few-shot examples use the same field names and delimiters as the live query?

level: juniorimportance: must knowfreq 65%

answer

  1. the prompt is one continuous sequence
  2. the pattern is the instruction
  3. separator drift between examples and query
  4. stop the prompt mid-record
  5. render both sides from one template

basics

~20 s

Demonstrations teach by pattern continuation. If the examples use one separator and field set but the live query arrives in a different shape, the query stops reading as the next item in the series, and the model's output format drifts or it keeps writing examples.

solid answer

~50 s

A few-shot prompt is one long sequence, and the model's job is to continue it. The examples establish a template — field names, their order, the separator between items, the casing and spacing — and the live query is supposed to be an incomplete instance of that same template, stopped right where the answer should begin. Consistency is what makes that work. If your wildlife-sighting exemplars are separated by `###` with `Report:` / `Species:` / `Confidence:` fields, but the live report is wrapped in `---` and labelled `Observation:`, the model sees a *different* pattern starting and may invent its own field names, answer in prose, or continue generating more demonstrations instead of answering. The fix is mechanical: build the prompt from one template function so exemplars and the live query cannot drift apart, and end the query mid-pattern at the field you want filled.

code

markdown · 14 lines
markdown
Report: Two grey wolves crossing the ridge at dusk, clear view
Species: grey wolf
Confidence: high
###
Report: Something large moved in the brush, never got a look
Species: unknown
Confidence: low
###
Report: Small deer-like animal at the salt lick, seen briefly
Species: roe deer
Confidence: medium
###
Report: {{live report text}}
Species:

go deeper

for a junior

Be able to say that the model continues a pattern, so examples and the live query must look like items in the same series: same field names, same separator, same spacing.

for a middle

Explain the specific failure modes — invented field names, dropped fields, the model continuing the example series — and the trick of ending the prompt mid-record at the field you want filled.

for a senior

Show how you keep the format from drifting in a real codebase: exemplars stored as structured records, one renderer for exemplars and live query, and a check that catches divergence before it ships.

for a principal

Be clear that demonstration formatting is a probabilistic nudge, not a guarantee, and own the decision about where the real parsing contract lives and what happens on the outputs that still come back malformed.

## The prompt is one sequence There is no structural boundary in a prompt between "your examples" and "the real question". The model receives a single token stream and predicts what comes next. Everything few-shot prompting does rests on that: the examples set up a visible, repeating pattern, and the live query is arranged so that the single most natural continuation is the answer you want. This is why formatting consistency is not cosmetic. The pattern *is* the instruction. Break it and you have quietly asked a different question. ## What has to stay identical Four things drift in practice, and each has a characteristic failure: - **Field names.** `Report:` in the examples and `Observation:` in the query. The model no longer recognises the query as an instance of the demonstrated task, and often re-labels the output with field names it invented. - **Field order and set.** Examples with three fields, query with two, or the same fields in a different order. The model tends to fill in what it saw, producing fields you did not ask for or omitting the one you did. - **The separator between items.** Exemplars split by `###` while the live query is wrapped in `---`. This is the classic case: `---` reads as the start of a *new* section rather than the next item in the list, and the model frequently continues the example series — inventing another fake report and labelling it — instead of answering yours. - **Whitespace and casing.** `Species: grey wolf` in one exemplar, `species:grey wolf` in another. Inconsistency here weakens the pattern and shows up as unstable output casing, which then breaks whatever parses the response. ## Stop mid-pattern The second half of the technique is where the prompt ends. If each exemplar is `Report: … / Species: … / Confidence: …`, the live query should end with `Report: <text>` and then a bare `Species:` — with nothing after it. The model is now completing a partially written record, which is the easiest possible continuation and leaves almost no room to add a preamble like "Sure, here's the classification". Ending the prompt after the input text without the trailing field name invites exactly that preamble, because there is no half-finished line to complete. ## Why it drifts in real systems Almost nobody writes an inconsistent prompt on purpose. It happens because the exemplars live in one place — a fixture file, a spreadsheet, a constant at the top of a module — and the live query is assembled somewhere else, in application code, by a different person, months later. Someone adds a field to the query builder for a new use case and does not update the fixtures. Someone reformats the examples file and a trailing newline changes. A migration rewrites the separator. The durable fix is structural rather than a matter of care: render both the exemplars and the live query through **one template function**, so a change to the shape is applied to every instance by construction. If exemplars are stored as structured records rather than as pre-rendered strings, they cannot fall out of sync with the query at all — the renderer is the single source of the format. ## The distinction worth naming in an interview A sharp answer separates two concerns that often get conflated. Formatting the demonstrations is about **teaching a shape by example** — making the model's most likely continuation be the answer you want, in the layout you want. It is a soft, probabilistic mechanism. It is not the same as *enforcing* a shape at generation time, which is a decode-side concern with different tools and different guarantees. Good demonstration format raises the odds that a parse succeeds and makes the output stable enough to read; it does not make malformed output impossible. ## Prose versus fields One more question follows naturally: should the examples be strictly fielded at all, or is natural prose fine? Fielded exemplars — explicit `Field:` labels on separate lines — give the model an unambiguous slot to fill and give your parser an unambiguous thing to read. Prose exemplars ("the report describes two grey wolves, so this is a confident grey wolf sighting") are more natural and can suit open-ended tasks, but the boundary between input and answer becomes fuzzy, the model's stopping point becomes unpredictable, and downstream extraction turns into regex archaeology. For anything whose output is consumed by code rather than a human, fielded and consistent wins.

  • How should the prompt end so the model does not add a preamble before the answer?
    End mid-record, at the field name you want filled. If exemplars are Report / Species / Confidence, the live query should end with the report text and a bare `Species:` on its own line. The model is then completing a half-written record, which is the easiest continuation and leaves no natural place for "Sure, here's the classification".
  • Is it better to write demonstrations as labelled fields or as natural prose?
    Fielded, whenever code consumes the output. Explicit field names give the model an unambiguous slot to fill and give your parser an unambiguous thing to read, and they make the stopping point predictable. Prose exemplars suit open-ended tasks read by humans, but they blur the input-answer boundary and turn extraction into fragile pattern matching.
  • Why does this consistency decay in production prompts, and how do you prevent it?
    Because exemplars usually live in a fixture or constant while the live query is assembled in application code, by a different person, later. Someone adds a field or reformats one side. Prevent it structurally: store exemplars as structured records and render both them and the live query through a single template function, so the shape cannot diverge by construction.

saying these in an interview costs you the question

  • Treats separators and field names as cosmetic polish
  • Believes the model knows where the examples end and the query starts
  • Ends the prompt after the input without the trailing field name
  • Hand-writes the query format separately from the exemplar format
  • Confuses teaching a shape by example with enforcing it at generation time

context

open as a page

What is dynamic exemplar retrieval in few-shot prompting, and how does it differ from a fixed example block?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Dynamic exemplar retrieval picks the few-shot demonstrations per request, pulling the most similar labelled examples from a pool by vector similarity. A fixed block hardcodes the same demonstrations into the prompt for every request, whatever the input looks like.

open as a page

When picking few-shot exemplars, why must every label class appear?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Demonstrations define the label space the model treats as live. A class with no example is predicted far less often than it should be, even when the instruction names it, so coverage of every label comes before adding more examples.

open as a page

In few-shot prompting, do wrong labels in the demonstrations hurt accuracy?

level: middleimportance: must knowfreq 55%

basics

~20 s

Wrong labels hurt far less than most people expect. On classification tasks, swapping gold labels for random ones drawn from the same label set often costs only a few points, while destroying the input-label format costs much more.

open as a page

Why can permuting the same four few-shot examples swing accuracy by double digits?

level: middleimportance: must knowfreq 62%

basics

~20 s

Order is not cosmetic: the model reads the demonstration sequence as evidence about which label is likely, and examples nearest the end pull hardest. The same four exemplars in different orders can move a sentiment task from the fifties into the high eighties.

open as a page

How many few-shot exemplars are enough, and how do you find that point?

level: middleimportance: must knowfreq 56%

basics

~20 s

Find it empirically: sweep the shot count on a held-out set and ship the first point where added examples stop paying. Gains typically flatten after a handful, while every extra example keeps costing tokens and latency on every request.

open as a page

Why does a per-request retrieved exemplar block destroy prompt-cache hits?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Prompt caching matches an exact token prefix from the start of the request. Retrieved exemplars differ on every call, so everything from that block onward is uncached — placed near the top, the entire prompt is reprocessed every single request.

open as a page

Your few-shot examples show a bare label but the task asks for a label plus a reason. What breaks?

level: middleimportance: should knowfreq 45%

basics

~20 s

The demonstrated shape usually wins over the written instruction. The model emits bare labels and drops the reason, or produces an unstable mix across calls. Fix the exemplars so every one carries both fields, in the order you want them produced.

open as a page

How can nearest-neighbour exemplar retrieval hurt accuracy on a borderline query?

level: middleimportance: should knowfreq 45%

basics

~20 s

Nearest neighbours are similar to the query, not representative of the task. For a borderline case they often all carry one label, or duplicate each other, so the prompt quietly argues for that one answer instead of showing the model where the decision boundary sits.

open as a page

How does contextual calibration use a content-free input like "N/A" to correct label bias?

level: middleimportance: should knowfreq 30%

basics

~20 s

Feed the prompt an input carrying no task evidence, such as "N/A" or an empty string, and read the label probabilities it returns. Those probabilities are the prompt's built-in prior; dividing real predictions by them and renormalising removes most of the bias.

open as a page

Should a fixed few-shot exemplar set favor diversity or typical cases?

level: middleimportance: should knowfreq 47%

basics

~20 s

Both, in a specific order: cover the distinct modes of the traffic you actually receive, then weight roughly toward the common ones. Cloning one typical case wastes tokens; filling the set with exotic cases teaches a distribution your users do not send.

open as a page

How do few-shot demonstrations signal the label space, and what happens to unseen classes?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The label strings appearing in the demonstrations act as the effective output vocabulary. A class named only in the instruction but never demonstrated is emitted rarely or never, so every class you want back must appear at least once, spelled exactly as your parser expects.

open as a page

A retrieved-exemplar store keeps teaching a product name retired six months ago — how do you fix it?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Treat the exemplar pool as a versioned production dataset, not a folder of examples. Give every entry a timestamp and a validity window, filter expired entries out at retrieval, verify anything written back before it becomes teachable, and audit the pool against the current taxonomy on a schedule.

open as a page

Why does a three-approve, one-deny shot set push a few-shot classifier toward approve?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Majority-label bias: the model reads the label mix in the demonstrations as evidence about how often each outcome occurs, so a 3:1 approve-heavy shot set inflates the predicted approve rate on borderline building-permit applications, independent of what each application actually says.

open as a page

Why add negative exemplars showing what a classifier must not flag?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Positive-only demonstrations teach where the rule applies but never where it stops, so the model over-generalizes and flags look-alikes. Near-miss negatives — cases that resemble a hit but are not one — draw the boundary the positives leave undefined.

open as a page

Who owns a few-shot exemplar pool, and when must it be refreshed?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

An exemplar pool is policy encoded as data, so it needs a named owner in the domain team, provenance and PII review on every entry, versioning alongside the prompt, and a refresh triggered by policy changes, label-set changes and drift — not by the calendar alone.

open as a page