skip to content

Your few-shot examples show a bare label but the task asks for a label plus a reason. What breaks?

level: middleimportance: should knowfreq 45%

answer

  1. two descriptions of the task disagree
  2. the nearer, concrete pattern wins
  3. symptom is instability, not uniform failure
  4. every field in every exemplar
  5. the shape you show is the shape you buy

basics

~20 s

The demonstrated shape usually wins over the written instruction. The model emits bare labels and drops the reason, or produces an unstable mix across calls. Fix the exemplars so every one carries both fields, in the order you want them produced.

solid answer

~50 s

When the instruction and the demonstrations disagree about the answer's shape, the demonstrations tend to win — they are the concrete, immediately preceding pattern the model is continuing, while the instruction is a general statement several hundred tokens earlier. So exemplars ending at `Decision: fast_track` will produce bare decisions even though the instruction asked for a justification, and the results are often *inconsistent* rather than uniformly wrong: some calls comply with the instruction, some follow the pattern, and anything parsing the output has to handle both. The rule is to **shape-match the demonstration to the expected answer**: every exemplar shows every field you want back, in the exact order and layout you want, with no field appearing in only some of them. The mismatch runs the other way too — exemplars carrying long rationales when you only need a label will produce long rationales, and you pay for them in tokens and latency on every call.

code

markdown · 14 lines
markdown
Claim: Windscreen chipped by road debris, no injury reported
Decision: fast_track
Reason: low value, single peril, no injury
---
Claim: Kitchen fire, three rooms damaged, tenant hospitalised
Decision: adjuster_visit
Reason: high value with injury, needs on-site assessment
---
Claim: Bicycle stolen from a locked shed, police report attached
Decision: fast_track
Reason: documented theft, within standard limit
---
Claim: {{live claim text}}
Decision:

go deeper

for a junior

Remember that the model copies the shape of the examples. If you want two fields back, show two fields in every example, in the order you want them written.

for a middle

Explain why the demonstrations outweigh the instruction and why the symptom is intermittent rather than uniform, and name the reverse mistake of demonstrating richer output than the consumer needs.

for a senior

Show a diagnosis order for unstable output shape, and treat instruction, demonstrations and parser as three descriptions of one contract that must be kept in agreement in the codebase.

for a principal

Own the design call about what the response should carry at all: whether a rationale is generated on every request or only for the flagged subset, and what the cost and latency of that choice is at production volume.

## Instruction versus demonstration A few-shot prompt usually contains two descriptions of the task: a written instruction near the top, and a set of worked examples below it. When they agree, nothing interesting happens. When they disagree about the *shape* of the answer, the examples generally dominate. The reason is mechanical. The model is continuing a sequence, and the strongest local evidence about what comes next is the pattern in the immediately preceding tokens — three records that each ended after a single label. The instruction is a general claim sitting further back in the context. It is not ignored, but it is competing against a concrete, repeated, adjacent template, and the template usually wins. ## The characteristic symptom is inconsistency What makes this bug annoying is that it rarely fails cleanly. On some inputs the instruction wins and you get `fast_track — low value, single peril`. On others the pattern wins and you get `fast_track`. On a few you get a bare label followed by an unrequested paragraph. A parser written against the first sample breaks on the second, and because the split depends on the input, it survives a quick manual test and fails in production. If you see a prompt whose output format is *usually* right, the first thing to check is whether the exemplars actually demonstrate the shape the instruction asks for. ## Shape-matching, stated as a rule Make the demonstration a literal specimen of the response you want: 1. **Every field you expect back appears in every exemplar.** Not most of them — every one. A field present in three of five exemplars teaches the model that it is optional, and you will get it about that often. 2. **The fields appear in the order you want them generated.** The model produces the record top to bottom; if you want the reason after the decision, show it after the decision in every example. 3. **The value style matches too.** If you want a one-line reason, do not demonstrate three-sentence reasons. The exemplars set the expected length, register and vocabulary just as firmly as they set the field list. 4. **The live query stops at the first field to be produced**, so the model completes the record rather than starting a fresh answer. ## The mismatch runs both ways The expensive direction is often the reverse of the one in the question. Teams build a prompt with richly explained exemplars because it reads well during development, then ship it in a path that only consumes the label. Every call now generates a rationale nobody reads: more output tokens, higher cost, and — because output tokens are generated one at a time — meaningfully more latency on a per-request path. If the consumer wants a label, demonstrate a label. There is also a middle case worth naming: you want a justification, but for humans reviewing flagged cases only. Rather than have every response carry one, the usual designs are a second, cheaper call on the subset that needs it, or two clearly separated fields so the consumer can drop one — but the demonstrations must reflect whichever you choose. The format you show is the format you buy. ## Diagnosing it Given a prompt with unstable output shape, work in this order. First, read the exemplars as if you were the model: what record shape do they demonstrate, field by field? Second, compare that shape to what the instruction asks for and to what your parser expects — three descriptions that must agree. Third, check whether any field is present in some exemplars and absent in others; that is the single most common cause of intermittently missing fields. Only after those should you reach for instruction rewording, because rewording an instruction to fight the demonstrations is arguing with the stronger signal. ## What this is not Shape-matching demonstrations makes the desired output *likely* and stable. It does not make the wrong shape impossible — the model can still emit a malformed record on an unusual input, and any consumer parsing the response needs to handle that. Guaranteeing structure is a separate, decode-time concern with its own tooling. Demonstration format and generation-time enforcement solve overlapping problems from different ends, and a strong answer keeps them distinct: format the examples so the model *wants* to produce the right shape, and validate on the way out because it sometimes will not.

  • What if a field appears in only some of your exemplars?
    You have taught the model that it is optional, and you will get it intermittently — which is worse than never getting it, because a parser written against a good sample passes review and then fails on live traffic. Every field you expect back belongs in every exemplar, in the same position, or it should not be in the format at all.
  • What is the cost of the opposite mistake — exemplars richer than the answer you need?
    You pay for it on every call. Exemplars carrying long rationales produce long rationales, and output tokens are generated sequentially, so it costs both money and latency on a per-request path even when the consumer only reads the label. If downstream code wants a label, demonstrate a label and get the explanation from a separate call on the subset that needs it.
  • If the instruction and the examples conflict, why not just write a firmer instruction?
    Because you are arguing with the stronger signal. The examples are a concrete, repeated pattern immediately before the generation point; the instruction is a general statement further back. Firmer wording sometimes helps and always leaves the conflict in place for the next reader. Fix the exemplars so all three descriptions — instruction, demonstrations, parser — agree.

saying these in an interview costs you the question

  • Assumes an explicit instruction overrides what the examples demonstrate
  • Adds a field to only some exemplars and expects it consistently
  • Ignores the token and latency cost of over-rich demonstrated outputs
  • Fixes shape drift by rewording the instruction instead of the examples
  • Treats demonstration format as a guarantee that output will parse

context