Your Tree of Thought branches are three rewordings of one idea — how do you get real diversity?
answer
- effective branching factor near one
- different wording is not a different plan
- let the siblings see each other
- name the axis they must differ on
- de-duplicate before you expand
basics
~20 sAsk for candidates that must differ along a named axis, show the model the siblings it already produced so it can avoid them, and drop near-duplicates before expanding. Then measure the distinct-candidate rate, because paraphrase siblings mean you are paying for width you never got.
solid answer
~50 sParaphrase siblings mean the effective branching factor is one: you pay for three branches and explore a single idea, and the waste compounds at every level. Fixes live in the generation step itself. First, **name the axis of difference** in the prompt — for candidate next paragraphs of a story bound by a pre-written plan, require that each candidate advance a *different* plan beat, not merely word the same beat differently. Second, **make candidates aware of each other**: either enumerate them in one call, or feed the siblings already generated back in with an instruction to add something unlike them. Third, **de-duplicate before expanding** — cheap lexical similarity catches copies, and a short model pass catches semantic ones. Fourth, **partition the option space** when you know its structure, assigning each branch a category. Finally, track the post-de-duplication candidate count as a metric; without it you cannot tell branching from repetition.
go deeper
Recognise the symptom: if the candidate next steps say the same thing in different words, the tree is not really branching. Know that the prompt should ask for options that differ in a specific, stated way.
Explain why independent draws cannot avoid each other, and describe two concrete fixes — enumerating candidates in one call so they are mutually aware, and de-duplicating before expansion.
Show the diagnosis end to end: instrument the effective branching factor, distinguish lexical from semantic duplicates, and partition a known option space rather than hoping the model spreads out. Say plainly when a state has only one real move.
Own the economics. Wasted breadth multiplies with depth, so treat effective branching factor as a first-class budget metric and be ready to argue that a task whose states admit one real alternative should not be searched as a tree.
## The symptom and why it is expensive A thought tree with breadth three that returns three rewordings of one idea has an **effective branching factor of one**. Every downstream stage inherits the problem: the ranking step is asked to choose between clones, and the traversal explores what looks like a wide frontier but is one line of reasoning in triplicate. The waste is multiplicative — at depth four with breadth three, you are paying roughly eighty-one times a single-path cost for the coverage of a single path. This is the most common reason a Tree of Thought implementation shows no improvement over simply reasoning once. The root cause is usually not the model being incapable of alternatives. It is that nothing in the generation step ever asked for a *different kind* of step, and nothing prevented the same high-probability continuation from being drawn repeatedly. ## Fix 1: name the axis of difference The weakest possible instruction is "give three different next steps," because "different" is satisfied by different wording. The strong version specifies *what* must differ. Consider generating five candidate next paragraphs for a short story that must satisfy a pre-written plan constraint. Instead of "five different paragraphs," require that each candidate take the story to a different one of the plan's remaining beats, or that each adopt a different narrative move — advance the conflict, reveal backstory, shift viewpoint, compress time. Now the candidates are forced apart at the level of decisions rather than sentences, and the constraint is checkable. This works because it converts an unstructured request into a partition of the option space, which is the next fix in stronger form. ## Fix 2: make candidates aware of each other Independent draws cannot avoid each other — that is the definition of independence. Two remedies: - **Enumerate in a single call**, so each candidate is written conditioned on the earlier ones and an instruction to differ has something to bind to. - **Iterate with siblings in the prompt**: generate one, then ask for a next candidate that is materially unlike the ones shown. This costs more calls but gives the tightest control, and it lets you stop as soon as a new candidate is only a paraphrase — a natural, adaptive breadth. ## Fix 3: de-duplicate before expanding Whatever the generation strategy, treat de-duplication as a hard step in the loop rather than an optimization. Exact and near-exact copies fall out of simple normalized string comparison. Semantic near-duplicates — the same plan expressed differently — need either a similarity measure over representations of the candidates or a short model pass asked to cluster them and keep one representative per cluster. Collapsing a five-candidate set to three real options and expanding only those is a straight cost saving, and it makes the tree's shape honest. ## Fix 4: partition the space when you know its structure When you can name the categories of a good next step for the task family, assign one per branch instead of hoping the model spreads out: for an outline, one structural option per branch; for a story continuation, one narrative move per branch. This turns diversity from a sampling accident into a design property, and it also gives you **coverage**, which is a different property from diversity. Three distinct candidates that all sit in one corner of the option space are diverse and still miss the region where the answer lives. Enumerating the axis explicitly is the only reliable way to cover it. ## Fix 5: measure it The metric to instrument is the **post-de-duplication candidate count** per expansion — the effective branching factor. Track its distribution across a batch of real tasks. If it hovers near one, branching is not happening and no downstream tuning will rescue it. Related signals: the fraction of expansions where every candidate survived de-duplication (suspiciously high can mean your duplicate detector is too strict), and whether solutions found by the tree ever descend through a non-first candidate. If the eventual answer always comes from the first branch, the other branches were decoration. ## Honest caveats Diversity is not free. Candidates forced apart along an axis include weaker options by construction — that is the point of a search, but it raises the burden on whatever prunes them. And on some tasks there genuinely are only one or two reasonable next steps; a low effective branching factor there is a property of the problem, and the correct response is to stop branching at that state rather than to manufacture artificial alternatives. Finally, note what does *not* fix this: turning up randomness in the sampling produces surface variation — different phrasings, occasionally worse coherence — while leaving the underlying plan the same. Diversity of wording is not diversity of thought, and only structural instructions and de-duplication address the latter.
- How would you measure whether branching is actually happening?Instrument the effective branching factor: candidates remaining after de-duplication, per expansion, across a batch of real tasks. If it sits near one, the tree is a chain. A second signal is whether the final answer ever descends through a non-first candidate; if it never does, the extra branches were decorative and can be cut without loss.
- Is a low effective branching factor always a bug?No. Some states genuinely admit only one sensible next step — a forced move, a single viable continuation under hard constraints. The right response there is to stop branching at that state and continue linearly, not to manufacture weaker alternatives to fill a quota. The bug is a low branching factor everywhere, on tasks that plainly have alternatives.
- What is the difference between diverse candidates and good coverage of the option space?Diversity is pairwise dissimilarity among the candidates you produced; coverage is whether they collectively reach the regions where good answers live. Three mutually distinct options clustered in one corner are diverse and miss the answer. Coverage is why explicitly partitioning a known option space beats relying on the candidates simply being unlike each other.
saying these in an interview costs you the question
- Thinks raising sampling randomness produces genuinely different plans
- Counts branches generated instead of branches that differ
- Skips de-duplication and expands paraphrases as separate nodes
- Asks only for 'three different options' without naming the difference
- Assumes every state must have the same number of real alternatives