How many Thought/Action/Observation exemplars should a ReAct prompt carry?
answer
- a handful, not a library
- they teach format, not the task
- re-sent on every step of the loop
- copied surface details are the failure
- complete trajectories, ending in an answer
basics
~20 sA handful — typically two to six complete trajectories — is enough to teach the block grammar and the handoff rhythm. Add more only when a measured failure demands it, because exemplars sit in the prefix and are re-sent on every step of the loop.
solid answer
~50 sExemplars in a ReAct prompt exist to teach a **format**, not a domain. Two to six complete trajectories — each running from the first thought through to a final answer — is the usual range; the original multi-hop QA work used a handful of hand-written traces. Each one must be complete, because a truncated exemplar teaches the model to truncate. Include at least one trajectory that is not a straight-line success, so the model has seen what recovering inside the format looks like. The cost is structural: exemplars live in the stable prefix and are re-sent on every single step, so a ten-step run pays for them ten times. The other cost is imitation — over-specific exemplars get copied, and the model will invoke a tool that exists only in your examples. With a strongly instruction-tuned model, one or two exemplars plus a clear format description usually beats six.
code
markdown · 12 linesQuestion: Which magazine was started first, Arthur's Magazine or First for Women?
Thought: I need the founding year of each magazine, so I will look up the first one.
Action: Search[Arthur's Magazine]
Observation: Arthur's Magazine was an American literary periodical published in the 19th century.
Thought: The founding year is not stated, so I will look for it inside the article.
Action: Lookup[founded]
Observation: It was founded in 1844 in Philadelphia.
Thought: Now I need the founding year of the other magazine.
Action: Search[First for Women]
Observation: First for Women is a woman's magazine launched in 1989.
Thought: 1844 is earlier than 1989, so the first magazine was started first.
Action: Finish[Arthur's Magazine]go deeper
Know that a ReAct prompt usually shows a few complete worked examples of the Thought, Action, Observation pattern, and that each one should run all the way to a final answer.
Explain the mechanics: two to six complete trajectories, teaching format rather than content, with the cost that the whole block sits in the prefix and is re-sent on every step of the loop.
Show you tune this empirically. Talk about counting parse failures and alternation errors on an eval set, recognising imitation tells such as invented tools or copied answers, and deleting exemplars as readily as adding them.
Own the budget argument. Exemplar tokens compete with the growing transcript for the same context window on long runs, so decide deliberately what format guidance stays in the prompt, what moves into a structured interface, and what belongs in weights.
## What exemplars are for A ReAct exemplar is a **complete worked trajectory** pasted into the prompt: a sequence of `Thought:` / `Action:` / `Observation:` blocks ending in a final answer. Its job is to teach three things at once — the literal block labels, the alternation rhythm (never two actions without an observation between them), and the shape of a terminating step. It is a specification written in the form of an example. Crucially it is *not* there to teach the task. If your exemplars are carrying the domain knowledge, you have confused few-shot formatting with few-shot learning, and you will pay for that confusion in imitation errors. ## How many The honest answer is: as few as reliably produce parseable, well-alternating output on your eval set. In practice that lands between two and six complete trajectories. The original multi-hop question-answering demonstrations of ReAct used a handful of hand-written traces over a small action vocabulary. Modern instruction-tuned models need fewer than early base models did — one or two exemplars alongside a prose description of the format is frequently enough, and some models follow the format from the description alone. The method is measurement, not folklore. Build a small set of tasks, run the loop, and count two things: the parse-failure rate on action lines, and the rate of malformed alternation. Add an exemplar only when it moves one of those numbers. Removing exemplars is equally legitimate work — every one you delete is tokens saved on every step of every run. ## What each exemplar must contain - **A full trajectory.** Start at the first thought and end at a final answer. A trajectory that stops mid-loop teaches stopping mid-loop. - **Realistic length.** If your real tasks take five steps and every exemplar takes two, the model will try to finish in two and answer prematurely. - **At least one non-straight-line run.** A trajectory where a first attempt does not yield what was needed and the model adjusts shows the format for continuing, rather than the format for giving up. - **Actions drawn from the real catalogue.** If an exemplar calls a tool you do not actually expose, you have taught the model that tool exists. ## Task-specific or generic? This is the real tradeoff. Exemplars close to your domain improve tool selection and the register of the reasoning; exemplars far from your domain teach the grammar without contaminating the content. Both directions have a failure mode: - **Too domain-specific** and the model imitates surface details — reusing an exemplar's entity names, copying its answer shape, calling its tools, and in the worst case reproducing an exemplar's final answer for a different question. - **Too generic** and format compliance holds but tool choice drifts, because nothing has demonstrated which of your tools suits which kind of subgoal. A workable middle: keep the trajectories in your domain but make each one demonstrably different from the real tasks — different entities, different tool ordering, different length — so there is nothing useful to copy. ## The cost you keep paying Exemplars sit in the prompt prefix. In a ReAct loop the entire transcript is re-sent on every step, so a 1,500-token exemplar block on a ten-step run costs 15,000 input tokens. That is a latency cost as well as a money cost, though a stable prefix is at least friendly to prompt caching, which blunts the money side considerably when the block never changes between turns. What caching cannot recover is context budget: exemplars compete for the same window as the growing transcript of real thoughts and observations, and on long runs that competition is the binding constraint. ## Diagnosing exemplar problems Some specific tells and what they mean: - Model invents a tool → the tool appears in exemplars but not in the catalogue, or the catalogue is weaker than the examples. - Model answers after two steps on a task that needs six → exemplars are too short. - Model echoes an exemplar's entity or answer → exemplars are too close to the live task. - Action lines fail to parse → the format is under-demonstrated; add one exemplar or tighten the format description rather than adding four. - Model emits two actions with no observation between them → the alternation rhythm is not being taught; check that no exemplar shows that pattern. ## Where exemplars go away entirely When the loop protocol moves into a provider's structured tool-calling interface, the format is enforced by the API rather than demonstrated in prose, and the few-shot block shrinks to nothing or to a couple of examples of *judgement* — when to use which tool — rather than examples of syntax. The same happens when the loop is fine-tuned into the weights. In both cases what remains is the interesting part: showing the model good decisions, not showing it where the colons go.
- How would you decide whether your exemplars should come from the agent's real domain?Weigh tool-selection accuracy against imitation risk. Domain-matched trajectories show which tool fits which subgoal, which generic ones cannot. But the closer they sit to live tasks, the more the model copies entities and answers. The usual compromise is same domain, deliberately different specifics — different entities, different tool ordering, different step count — so there is nothing worth copying.
- What tells you the exemplar block is too short rather than too long?Premature termination and parse failures. If the model answers in two steps on tasks that genuinely need five, your exemplars are shorter than reality. If action lines fail to parse or the model emits two actions without an observation between them, the alternation grammar is under-demonstrated. Both are fixed by one better exemplar far more often than by four more.
- Why does an exemplar that stops mid-loop cause trouble?Because the exemplar is the specification. A trajectory that ends after an observation, with no final answer, demonstrates a run that never terminates cleanly, and the model imitates that — trailing off, re-asking itself, or emitting a final answer in a shape you never defined. Every exemplar should run end to end, including the terminating step.
saying these in an interview costs you the question
- Adds more exemplars whenever quality dips, without measuring
- Treats exemplars as the place to put domain knowledge
- Forgets exemplars are re-sent on every step of the loop
- Uses truncated trajectories with no final answer
- Includes tools in exemplars that the agent cannot actually call