skip to content

How do LangChain few-shot prompt templates select which examples to include?

level: seniorimportance: should knowfreq 44%

answer

  1. static list or a selector, never both
  2. example_prompt renders one example
  3. semantic selector embeds the incoming input
  4. dynamic examples break a cached prefix
  5. chat variant emits human/ai turn pairs

basics

~20 s

A few-shot template either holds a fixed examples list, rendered in full every call, or an example_selector that picks a subset per request — for instance SemanticSimilarityExampleSelector, which embeds the incoming input and returns the k nearest stored examples.

solid answer

~50 s

LangChain gives you `FewShotPromptTemplate` for string prompts and `FewShotChatMessagePromptTemplate` for chat prompts. Both take an `example_prompt` — a template describing how one example renders — plus **either** a static `examples` list **or** an `example_selector`. Static examples are deterministic and cacheable, which is their main virtue: the rendered block is byte-identical every call, so a provider's prompt-prefix cache keeps hitting. A selector such as `SemanticSimilarityExampleSelector` (built via `from_examples(examples, embeddings, vectorstore_cls, k=...)`) or `LengthBasedExampleSelector` chooses per request, which raises accuracy on heterogeneous inputs but costs an embedding call, adds a retrieval dependency in the hot path, and makes the prompt prefix vary — killing that cache and making failures harder to reproduce. In the chat variant each example expands to a human/ai message pair, so the model sees the demonstrations as real turns rather than as quoted text.

code

python · 21 lines
python
from langchain_core.prompts import ChatPromptTemplate, FewShotChatMessagePromptTemplate

example_prompt = ChatPromptTemplate.from_messages([
    ("human", "{input}"),
    ("ai", "{output}"),
])

few_shot = FewShotChatMessagePromptTemplate(
    example_prompt=example_prompt,
    examples=[
        {"input": "2+2", "output": "4"},
        {"input": "2+3", "output": "5"},
    ],
)

final = ChatPromptTemplate.from_messages([
    ("system", "You are a calculator. Reply with the number only."),
    few_shot,
    ("human", "{input}"),
])
print(final.format_messages(input="3+7"))

go deeper

for a junior

Know that a few-shot template pairs an example_prompt describing one example with a list of example dicts, and that it renders the demonstrations before the real input.

for a middle

Explain the two example sources — a static list versus an example_selector — and how the chat variant expands each example into a human/ai message pair inside a larger prompt.

for a senior

Weigh the operational cost of dynamic selection: an embedding plus vector query per request, a destroyed prompt-prefix cache, and answers you cannot reproduce unless you log which examples were chosen.

for a principal

Own the example pool as versioned, evaluated data with a change gate, decide the caching-aware prompt layout, and be able to argue when demonstrations should be replaced by a structured-output constraint instead.

## The two moving parts A few-shot template is composed of a per-example template and a source of examples. `example_prompt` says how one example renders. For the string case that is a `PromptTemplate` such as `"Q: {input}\nA: {output}"`. For the chat case it is a small `ChatPromptTemplate.from_messages([("human", "{input}"), ("ai", "{output}")])`, so every example becomes two real messages. The example source is either `examples` (a list of dicts whose keys match the example prompt's variables) or `example_selector` (an object asked, per request, which examples to use). Supplying both is a configuration error; supplying neither leaves nothing to render. ## Static examples ``` few_shot = FewShotChatMessagePromptTemplate( example_prompt=example_prompt, examples=[{"input": "2+2", "output": "4"}, {"input": "2+3", "output": "5"}], ) final = ChatPromptTemplate.from_messages([ ("system", "You are a calculator."), few_shot, ("human", "{input}"), ]) ``` Note how the few-shot block nests inside a larger chat prompt: it is just another element that expands to multiple messages. In the string variant you instead get `prefix`, the rendered examples joined by `example_separator`, and `suffix` — the suffix being where the live input goes. Static examples have properties production likes. The rendered block is identical every call, so it sits inside the stable prompt prefix that providers cache and that observability tools fingerprint. It costs nothing at request time. It is trivially reviewable in a diff. And when the model misbehaves you can reproduce the exact prompt. Their weakness is coverage. If your inputs span several distinct shapes, a fixed handful of demonstrations either bloats the prompt to cover all of them or biases the model toward whichever shapes made the list. ## Dynamic selection An example selector implements a small interface: given the input variables of this call, return the examples to use. LangChain ships several, of which two matter most. `SemanticSimilarityExampleSelector` embeds each stored example and, at request time, embeds the incoming input and returns the k nearest. Construct it directly from a vector store, or with `from_examples(examples, embeddings, vectorstore_cls, k=2)`, which builds the store for you. This is the selector people mean when they say "dynamic few-shot": the model sees demonstrations that look like the question actually asked. `LengthBasedExampleSelector` takes as many examples as fit a length budget, dropping the rest. It is about token safety rather than relevance — useful when example sizes vary wildly. When a selector is used, the few-shot template's own `input_variables` must include the fields the selector needs, because the selector is called with the formatting inputs. ## The tradeoff an interviewer is listening for Dynamic selection is not free, and the costs are operational rather than conceptual: 1. **Latency and dependency.** Every request now embeds the input and queries a vector store before the model call. That is an extra network hop and an extra thing that can be slow or down, sitting in front of every generation. 2. **Cache destruction.** Provider prompt caching keys on a stable prefix. If the examples change per request and sit near the top of the prompt, nothing before the live input is ever reused. On a high-volume endpoint that can dominate the accuracy win in cost terms. Mitigation: put dynamic examples *after* the stable system content, or keep a static core set and add only one or two dynamic ones. 3. **Reproducibility.** "Why did it answer that?" now requires knowing which examples were selected, so the selected example ids belong in your trace alongside the prompt. 4. **Silent drift.** Adding examples to the pool changes behaviour for inputs you never tested, because a new neighbour can displace an old one. A fixed list changes only when someone edits it in a reviewed diff. ## Practical guidance Start static. A handful of well-chosen, diverse demonstrations is often within noise of a retrieval-based selector, and it is dramatically easier to operate. Move to a selector when you can show, on an evaluation set, that a fixed set cannot cover the input distribution — and then log which examples were chosen, cap k, and place the block where it does the least damage to your cached prefix. Also remember what few-shot examples are *for* in a chat model: with the chat variant they arrive as prior turns, which strongly shapes output format. If the goal is purely a rigid output shape, a structured-output constraint is usually a cheaper and more reliable lever than demonstrations.

  • What does an example selector cost you on a high-QPS endpoint?
    An embedding call and a vector-store query in front of every generation, so extra latency and a new dependency that can fail. Worse, the examples now vary per request, so the provider's cached prompt prefix stops matching and you pay full input-token price on every call. Measure the accuracy gain against that before adopting one.
  • How do you keep dynamic examples from destroying prompt-prefix caching entirely?
    Keep the stable content — system instructions, a fixed core example set — at the top, and place the request-varying examples immediately before the live input. The cache then still covers the invariant prefix. Capping k low also helps, because fewer varying tokens means less of the prompt falls outside the cacheable region.
  • Why do chat few-shot examples influence the model more strongly than the same text quoted in a system message?
    Because the chat variant renders each example as an actual human message followed by an actual ai message, so the model sees demonstrations in the exact position and role its own answer will occupy. Quoted inside a system message they are just described behaviour; as turns they are modelled behaviour, and the format tends to be copied far more faithfully.
  • What breaks when you add new examples to a selector's pool?
    Behaviour changes for inputs nobody re-tested: a new example can become a nearest neighbour and displace one that was previously chosen, silently altering answers. Unlike editing a static list, this happens without a reviewed prompt diff. Treat the example pool as versioned data with an evaluation run gating changes to it.

saying these in an interview costs you the question

  • Passing both examples and an example_selector together
  • Assuming more examples monotonically improve accuracy
  • Ignoring the embedding call a selector adds per request
  • Believing dynamic examples still hit the provider prompt cache
  • Using few-shot demonstrations where structured output would do

context