How do you choose how large a single thought should be in a Tree of Thought?
answer
- how big is one branch?
- two failure modes, one on each side
- notation-level differences are not alternatives
- could you judge a partial answer here?
- smallest unit whose promise is assessable
basics
~20 sSize a thought so the model can produce several meaningfully different versions of it, and so a partial solution built from it can already be judged promising or hopeless. Too fine and siblings look identical; too coarse and there is almost nothing to branch over.
solid answer
~50 sThought granularity is the unit of one branch, and it sits between two failure modes. Make thoughts too fine — a single algebraic manipulation inside a proof sketch — and the candidates at each state are near-identical, the tree grows very deep, and almost every step is too small to tell you anything about where the branch is heading. Make them too coarse — a whole lemma, or a whole essay section — and you get only a handful of distinguishable options, each expensive to produce, and abandoning one throws away a lot of work. The working rule: **a thought should be the smallest unit whose promise you could plausibly comment on before the answer is complete**. It should also match the task's state space — rich open-ended tasks tolerate large, semantically loaded thoughts, while tightly constrained tasks want small steps because the constraint itself supplies the distinctions.
go deeper
Know that a thought is one step of a partial solution and that you get to choose how big that step is. Be able to give an example of a too-small step and a too-large one.
Explain both failure modes and the judgeability rule: the unit should be the smallest increment whose partial result can already be called promising or doomed. Connect thought size to search depth and total call count.
Demonstrate calibration from real instances rather than intuition, and name the diagnostic signals — paraphrase siblings, branches never abandoned, errors found inside a thought — that tell you the unit is mis-sized.
Frame granularity as a cost-of-exploration decision for a task family, and be prepared to conclude that a task has no assessable intermediate state and therefore should not be searched as a tree at all.
## What granularity means here A "thought" in Tree of Thought is one increment of partial solution: what the model adds at a state to move one level deeper. Granularity is how much that increment contains. For a proof sketch it could be a single algebraic manipulation, a full step of the argument, or an entire lemma. For a piece of writing it could be a phrase, a sentence, or a whole paragraph. Nothing in the method fixes the size — you choose it, and the choice shapes the search more than almost any other knob. ## The two failure modes **Too fine.** Suppose a thought is one algebraic manipulation. Two consequences follow. First, the k candidates at a state end up trivially different from one another — rearrange this term, rearrange that one — so branching buys almost no genuine alternatives; you are exploring notation, not ideas. Second, the depth needed to reach a complete solution explodes, and since cost scales with breadth times depth, you pay a lot for that notation. There is a third, subtler cost: a partial state one manipulation further along usually looks exactly as promising as its parent, so nothing distinguishes branches early enough to act on. **Too coarse.** Now suppose a thought is a whole lemma. The model produces perhaps two or three genuinely distinct lemmas, each long, each expensive, and each carrying a lot of internal reasoning that was never itself branched over — a mistake buried in the middle of a lemma is invisible until the whole thing is done. Backtracking becomes brutal: discarding one candidate discards a great deal of work. At the limit, a thought the size of the whole answer reduces the tree to "generate a few complete solutions and pick one," which is a different and much weaker technique wearing Tree of Thought's clothes. ## The sizing rule The most useful heuristic is *judgeability*: **a thought should be the smallest unit such that a partial solution ending there can already be called promising, neutral, or doomed.** Below that size, the increments are noise. That rule is task-relative, which is why granularity is not a universal setting. A second, complementary rule is *distinguishability*: at the intended size, can the model produce k candidates that a reader would call genuinely different approaches rather than rewordings? If not, the unit is too small — or the task has too little branching structure at that level to be worth a tree. ## Matching granularity to the state space The shape of the task pushes the answer around: - **Rich, open-ended spaces** (drafting, design, hypothesis generation) support large thoughts, because a large increment still leaves the model with many genuinely different options. A whole candidate paragraph of a story is a reasonable thought; a single word is not. - **Small, constrained spaces** (puzzles with legal-move structure, slot-filling under hard constraints) want smaller thoughts, because the constraints themselves generate the distinctions and because a large increment risks violating a constraint deep inside itself, where no branch point can catch it. - **Verifiable spaces** — anywhere an external check exists — want thoughts aligned to the checkpoints where that check can run. If a step can be validated, do not bundle two of them into one thought; you lose the ability to localize the failure. ## Practical calibration A workable procedure: solve two or three instances of the task by hand, write down the steps you naturally paused at, and use that as the unit. Human pause points tend to sit exactly where a decision was made and where the work so far could be assessed — which is what the judgeability rule asks for. Then sanity-check the arithmetic. If the typical solution needs D thoughts at your chosen size and you intend breadth b, the search is on the order of b·D generations before any pruning. If that number is untenable for the task's value, the honest response is a coarser thought, a narrower breadth, or not branching at all — not a silent hope that pruning saves you. ## Signals that the size is wrong - Sibling candidates read as paraphrases across the whole tree → thoughts are too fine (or the space is too narrow for branching). - Branches are rarely abandoned, because by the time one looks bad it is nearly the answer → thoughts are too coarse. - Errors are consistently found *inside* a single thought rather than between thoughts → the unit is bundling decisions that deserved their own branch point. Granularity is best treated as a per-task-family parameter you calibrate once on a handful of real instances, not as a constant carried between tasks.
- How does thought size interact with the total cost of the search?Roughly, cost scales with breadth times depth, and depth is the solution length divided by the thought size. Halving the thought size doubles the depth and therefore doubles the generation count at the same breadth — while usually making sibling candidates more similar. That is why over-fine granularity is expensive twice over: more calls, less exploration per call.
- Does the right granularity change between an exploratory phase and a refinement phase of the same task?Often, yes. Early on, large thoughts — whole approaches — separate genuinely different plans cheaply. Once a plan is chosen, the interesting decisions are local, and smaller thoughts let the search branch where the remaining risk actually lives. Varying the unit by depth is legitimate; it just has to be a deliberate choice rather than prompt drift.
- What if a task has no natural intermediate state at all?Then the tree has nothing to branch over, and that is a real finding rather than a tuning problem. Tasks whose answers are produced in one shot, with no assessable partial form, do not benefit from a thought tree; you are better off generating several complete answers and choosing among them, or using a single linear reasoning pass.
saying these in an interview costs you the question
- Believes smaller thoughts are always better because search is finer
- Sets one thought size and reuses it across unrelated tasks
- Makes a thought the whole solution and still calls it a tree
- Ignores that depth times breadth is the real cost driver
- Cannot say what makes two candidate thoughts genuinely different