skip to content

In a Tree of Thought prompt, what should a few-shot exemplar of a thought demonstrate?

level: juniorimportance: should knowfreq 38%

answer

  1. show form, not the answer
  2. it must stop mid-solution
  3. three visibly different options
  4. borrow from a different instance
  5. demonstrations beat instructions

basics

~20 s

An exemplar should show the shape of one step, not a finished solution: a partial state, several candidate continuations of the intended size, and a separable format. Exemplars that solve a problem end to end teach the model to answer instead of to branch.

solid answer

~50 s

Few-shot exemplars in a thought-generation prompt teach *form*, not content. A good one shows a partial state from a different instance of the same task family, then a small set of candidate next steps at exactly the granularity you want — for a turn-based strategy game, a mid-game position followed by three distinct candidate move sets, never a full playthrough. Three properties matter. The exemplar must **stop mid-solution**, so the model learns that a thought is an increment rather than an answer. Its candidates must be **visibly distinct from each other**, so the model copies the habit of offering real alternatives rather than variations. And the candidates must be **separable**, one per line or delimited, so the set can be split into branches. Draw exemplars from a different instance than the one being solved: reusing the current problem anchors the model on a particular continuation and quietly narrows the tree.

code

markdown · 6 lines
markdown
Position: white to move, knight on f5, black king castled short.

Candidate next moves:
1. Nxg7 — sacrifice to open the king's cover
2. Qd2 — build up before committing material
3. c4 — expand on the opposite wing and delay the attack

go deeper

for a junior

Remember that the exemplar shows a partial step and several options, never a finished answer, and that it should come from a different problem than the one you are solving.

for a middle

Explain that exemplars convey granularity, mutual distinctness and output format more reliably than instruction text, and that a worked-solution exemplar collapses the tree by teaching the model to finish.

for a senior

Diagnose from the symptom: complete answers mean a finishing exemplar, echoed content means anchoring, unsplittable output means a weak separator. Also account for exemplar tokens being re-sent at every expansion.

for a principal

Treat exemplars as a governed asset for a task family — versioned, rotated to avoid anchoring, and budgeted, since their token cost is multiplied by the entire search rather than paid once.

## Why exemplars behave differently here In ordinary prompting, a few-shot exemplar shows an input and the answer you want. In thought generation the wanted output is not an answer at all — it is a *set of candidate partial steps*. Everything about exemplar design follows from that one difference, and the most common mistake is importing the ordinary habit unchanged. ## What the exemplar must teach **That a thought stops early.** The exemplar's output ends mid-solution, on purpose. If your exemplar shows a problem being carried through to a final answer, the model has been taught that a good response finishes the job, and it will do exactly that at the first state — collapsing the tree into a single completed attempt. Nothing else in the prompt reliably overrides a worked-solution exemplar; instructions lose to demonstrations. **How big one step is.** Granularity is very hard to convey in words ("a moderate step") and very easy to convey by example. If a thought should be one move set rather than one move, or one paragraph rather than one sentence, show that size and the model will match it. This is the cheapest granularity control available. **That candidates are alternatives, not variations.** Show several candidates per exemplar and make them obviously different from one another — different plans, not different phrasings. The model copies the *relationship between the candidates*, not just their individual form. A three-candidate exemplar whose options are near-identical is actively harmful: it demonstrates the exact failure you are trying to avoid. **The count and the separator.** If you want k candidates in a fixed, splittable format, the exemplar should show k candidates in that format. Delimiters and one-per-line structure are learned far more reliably from a demonstration than from a description. ## Where exemplars should come from Use a **different instance of the same task family** than the one being solved. Two reasons. First, leakage: an exemplar drawn from the live problem can hand the model part of the answer, and any evaluation of the system then measures the exemplar rather than the reasoning. Second, anchoring: candidates that resemble the exemplar's content will be over-produced, which narrows the branch set at exactly the moment you wanted it wide. A related trap is recycling the tree's own best branch so far as an exemplar. It seems economical and it is a diversity killer — every subsequent expansion is nudged toward the shape of the branch you already favour, which is a feedback loop that quietly turns the search into a single line. ## Practical construction A workable exemplar for a turn-based strategy game with a small, enumerable action space looks like: a compact description of a mid-game position; then three candidate move sets labelled and separated, each pursuing a visibly different intent — press the flank, consolidate defensively, trade material for tempo. It does not evaluate them, does not pick one, and does not continue the game. Judging the candidates is a separate stage with its own prompt, and mixing the two into one exemplar teaches the model to prematurely commit. Keep the number of exemplars small. Every exemplar is re-sent on every generation call at every state of the tree, so exemplar tokens are multiplied by the size of the search. Two well-chosen exemplars usually convey step size, distinctness and format; a long gallery mostly buys cost and increases the chance the model latches onto exemplar content. ## How to tell an exemplar is misfiring - The model returns a complete solution instead of a step → an exemplar somewhere finishes a problem. - Candidates are the right count but the wrong size → the exemplar's step granularity does not match what you asked for in words; trust the exemplar, fix the exemplar. - Candidates keep echoing the exemplar's domain or specific move → anchoring; rotate exemplars or draw them from a more distant instance. - Candidates cannot be split apart reliably → the exemplar's separator is not distinctive enough to be learned. In every case the fix is in the demonstration, not in more instruction text. Exemplars are the strongest signal in the prompt, which is precisely why a careless one does so much damage.

  • Why is reusing the tree's current best branch as an exemplar a bad idea?
    It creates a feedback loop. Every later expansion is nudged toward the shape of the branch you already favour, so candidates drift toward variations of it and the effective branching factor falls. The search then converges on its own early preference rather than exploring, which is the exact behaviour a thought tree exists to avoid.
  • How many exemplars should a thought-generation prompt carry?
    Few — typically one or two. The prompt is re-sent at every state of the tree, so exemplar tokens are multiplied by the size of the search, and a long gallery raises the chance the model copies exemplar content rather than exemplar form. Two exemplars are usually enough to convey step size, mutual distinctness and output format.
  • What if the model matches the exemplar's format but not the step size you described in words?
    Trust the demonstration over the description and fix the exemplar. Models weight a concrete example far more heavily than an adjective like 'moderate' or 'small'. Rewrite the exemplar so its candidates are literally the size you want, and the instruction text can then stay short.

An exemplar here is like showing a chess student three plausible replies to a position rather than a whole annotated game: the point is to model what considering options looks like, not to hand over the result.

saying these in an interview costs you the question

  • Uses a fully worked solution as the exemplar
  • Draws the exemplar from the problem currently being solved
  • Shows one candidate per exemplar instead of several
  • Makes the exemplar's candidates near-identical to each other
  • Packs scoring or a chosen winner into the generation exemplar

context