skip to content

Thought Generation

Producing the branches: prompting for several candidate next steps, sampling them independently versus proposing them in one pass, and choosing how large a single 'thought' should be. Interviewers ask how you get genuine diversity instead of three rewordings of the same idea.

on this pageshow

questions

4

How do you choose how large a single thought should be in a Tree of Thought?

level: middleimportance: must knowfreq 52%

answer

  1. how big is one branch?
  2. two failure modes, one on each side
  3. notation-level differences are not alternatives
  4. could you judge a partial answer here?
  5. smallest unit whose promise is assessable

basics

~20 s

Size a thought so the model can produce several meaningfully different versions of it, and so a partial solution built from it can already be judged promising or hopeless. Too fine and siblings look identical; too coarse and there is almost nothing to branch over.

solid answer

~50 s

Thought granularity is the unit of one branch, and it sits between two failure modes. Make thoughts too fine — a single algebraic manipulation inside a proof sketch — and the candidates at each state are near-identical, the tree grows very deep, and almost every step is too small to tell you anything about where the branch is heading. Make them too coarse — a whole lemma, or a whole essay section — and you get only a handful of distinguishable options, each expensive to produce, and abandoning one throws away a lot of work. The working rule: **a thought should be the smallest unit whose promise you could plausibly comment on before the answer is complete**. It should also match the task's state space — rich open-ended tasks tolerate large, semantically loaded thoughts, while tightly constrained tasks want small steps because the constraint itself supplies the distinctions.

go deeper

for a junior

Know that a thought is one step of a partial solution and that you get to choose how big that step is. Be able to give an example of a too-small step and a too-large one.

for a middle

Explain both failure modes and the judgeability rule: the unit should be the smallest increment whose partial result can already be called promising or doomed. Connect thought size to search depth and total call count.

for a senior

Demonstrate calibration from real instances rather than intuition, and name the diagnostic signals — paraphrase siblings, branches never abandoned, errors found inside a thought — that tell you the unit is mis-sized.

for a principal

Frame granularity as a cost-of-exploration decision for a task family, and be prepared to conclude that a task has no assessable intermediate state and therefore should not be searched as a tree at all.

## What granularity means here A "thought" in Tree of Thought is one increment of partial solution: what the model adds at a state to move one level deeper. Granularity is how much that increment contains. For a proof sketch it could be a single algebraic manipulation, a full step of the argument, or an entire lemma. For a piece of writing it could be a phrase, a sentence, or a whole paragraph. Nothing in the method fixes the size — you choose it, and the choice shapes the search more than almost any other knob. ## The two failure modes **Too fine.** Suppose a thought is one algebraic manipulation. Two consequences follow. First, the k candidates at a state end up trivially different from one another — rearrange this term, rearrange that one — so branching buys almost no genuine alternatives; you are exploring notation, not ideas. Second, the depth needed to reach a complete solution explodes, and since cost scales with breadth times depth, you pay a lot for that notation. There is a third, subtler cost: a partial state one manipulation further along usually looks exactly as promising as its parent, so nothing distinguishes branches early enough to act on. **Too coarse.** Now suppose a thought is a whole lemma. The model produces perhaps two or three genuinely distinct lemmas, each long, each expensive, and each carrying a lot of internal reasoning that was never itself branched over — a mistake buried in the middle of a lemma is invisible until the whole thing is done. Backtracking becomes brutal: discarding one candidate discards a great deal of work. At the limit, a thought the size of the whole answer reduces the tree to "generate a few complete solutions and pick one," which is a different and much weaker technique wearing Tree of Thought's clothes. ## The sizing rule The most useful heuristic is *judgeability*: **a thought should be the smallest unit such that a partial solution ending there can already be called promising, neutral, or doomed.** Below that size, the increments are noise. That rule is task-relative, which is why granularity is not a universal setting. A second, complementary rule is *distinguishability*: at the intended size, can the model produce k candidates that a reader would call genuinely different approaches rather than rewordings? If not, the unit is too small — or the task has too little branching structure at that level to be worth a tree. ## Matching granularity to the state space The shape of the task pushes the answer around: - **Rich, open-ended spaces** (drafting, design, hypothesis generation) support large thoughts, because a large increment still leaves the model with many genuinely different options. A whole candidate paragraph of a story is a reasonable thought; a single word is not. - **Small, constrained spaces** (puzzles with legal-move structure, slot-filling under hard constraints) want smaller thoughts, because the constraints themselves generate the distinctions and because a large increment risks violating a constraint deep inside itself, where no branch point can catch it. - **Verifiable spaces** — anywhere an external check exists — want thoughts aligned to the checkpoints where that check can run. If a step can be validated, do not bundle two of them into one thought; you lose the ability to localize the failure. ## Practical calibration A workable procedure: solve two or three instances of the task by hand, write down the steps you naturally paused at, and use that as the unit. Human pause points tend to sit exactly where a decision was made and where the work so far could be assessed — which is what the judgeability rule asks for. Then sanity-check the arithmetic. If the typical solution needs D thoughts at your chosen size and you intend breadth b, the search is on the order of b·D generations before any pruning. If that number is untenable for the task's value, the honest response is a coarser thought, a narrower breadth, or not branching at all — not a silent hope that pruning saves you. ## Signals that the size is wrong - Sibling candidates read as paraphrases across the whole tree → thoughts are too fine (or the space is too narrow for branching). - Branches are rarely abandoned, because by the time one looks bad it is nearly the answer → thoughts are too coarse. - Errors are consistently found *inside* a single thought rather than between thoughts → the unit is bundling decisions that deserved their own branch point. Granularity is best treated as a per-task-family parameter you calibrate once on a handful of real instances, not as a constant carried between tasks.

  • How does thought size interact with the total cost of the search?
    Roughly, cost scales with breadth times depth, and depth is the solution length divided by the thought size. Halving the thought size doubles the depth and therefore doubles the generation count at the same breadth — while usually making sibling candidates more similar. That is why over-fine granularity is expensive twice over: more calls, less exploration per call.
  • Does the right granularity change between an exploratory phase and a refinement phase of the same task?
    Often, yes. Early on, large thoughts — whole approaches — separate genuinely different plans cheaply. Once a plan is chosen, the interesting decisions are local, and smaller thoughts let the search branch where the remaining risk actually lives. Varying the unit by depth is legitimate; it just has to be a deliberate choice rather than prompt drift.
  • What if a task has no natural intermediate state at all?
    Then the tree has nothing to branch over, and that is a real finding rather than a tuning problem. Tasks whose answers are produced in one shot, with no assessable partial form, do not benefit from a thought tree; you are better off generating several complete answers and choosing among them, or using a single linear reasoning pass.

saying these in an interview costs you the question

  • Believes smaller thoughts are always better because search is finer
  • Sets one thought size and reuses it across unrelated tasks
  • Makes a thought the whole solution and still calls it a tree
  • Ignores that depth times breadth is the real cost driver
  • Cannot say what makes two candidate thoughts genuinely different

context

open as a page

In Tree of Thought, when do you sample candidate thoughts independently instead of proposing them in one call?

level: middleimportance: must knowfreq 62%

basics

~20 s

Sample independently when the thought space is rich and open-ended, because separate draws naturally differ. Propose all candidates inside one call when the space is small and constrained, so each new candidate can see the earlier ones and avoid repeating them.

open as a page

In a Tree of Thought prompt, what should a few-shot exemplar of a thought demonstrate?

level: juniorimportance: should knowfreq 38%

basics

~20 s

An exemplar should show the shape of one step, not a finished solution: a partial state, several candidate continuations of the intended size, and a separable format. Exemplars that solve a problem end to end teach the model to answer instead of to branch.

open as a page

Your Tree of Thought branches are three rewordings of one idea — how do you get real diversity?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Ask for candidates that must differ along a named axis, show the model the siblings it already produced so it can avoid them, and drop near-duplicates before expanding. Then measure the distinct-candidate rate, because paraphrase siblings mean you are paying for width you never got.

open as a page