How much of a Tree of Thought run's compute should state evaluation consume?
answer
- evaluation is often the bigger half
- branching factor times samples per state
- marginal question, not a fixed percentage
- generation-limited versus evaluation-limited
- cheap filter first, expensive judgement on survivors
basics
~20 sThere is no fixed share; it is an allocation choice. Spend on evaluation up to the point where another unit of judgement improves the final answer more than another branch would, and measure that frontier per workload rather than assuming generation should dominate.
solid answer
~50 sEvaluation is often the larger line item, which surprises people. Every candidate state gets judged, judgements are usually sampled several times for stability, and scoring prompts carry the state plus the task criteria — so evaluator calls can outnumber generation calls by the sampling factor. The right allocation is empirical: hold total spend fixed and vary the split between exploring more states and judging the existing ones more carefully. On tasks where candidates are cheap and mostly wrong, judgement wins and deserves the larger share. On tasks where the evaluator has little signal, extra judging buys nothing and the money is better spent elsewhere entirely. The standard structural lever is a cascade — a free deterministic check, then a cheap or small-model pass, then expensive judgement only on survivors — which usually buys back most of the cost without touching accuracy.
go deeper
Know that scoring states is itself a cost — every candidate gets judged, often several times — so evaluation is not a rounding error next to generating the candidates.
Be able to do the arithmetic: evaluator calls scale as branching factor times samples per state, which is why evaluation frequently outweighs generation, and be able to name a cheap-filter-first cascade as the standard reduction.
Show how you would locate the bottleneck empirically — checking whether winning candidates were present but unselected — and how you would spend unevenly by depth and by proximity to the pruning cutoff instead of uniformly.
Own the allocation as a policy decision with the premise open to challenge: evaluation spend competes not only with generation but with a better model, a better prompt or a real verifier, and sometimes the right call is not to run a tree.
## Why evaluation is not the small line item The naive mental model is that a reasoning tree spends its budget generating thoughts and a little extra grading them. In practice the ratio often runs the other way. Count the calls. Every candidate state produced by every expansion has to be scored, so evaluator invocations scale with the same branching factor and depth as generation. Then multiply: value judgements are noisy, so they are typically sampled several times and aggregated, and comparative votes are sampled too. Finally, weigh the prompts — an evaluation prompt carries the state under judgement plus the task criteria and often the rubric, which is not obviously cheaper than the generation prompt that produced the state. An arrangement with a branching factor of five and five evaluation samples per state issues twenty-five evaluator calls per expansion against five generation calls. Evaluation is then the dominant cost by a wide margin, and any conversation about the run's economics that ignores it is talking about the smaller half. ## Framing it as an allocation, not a ratio There is no correct percentage, and an interviewer asking this is not looking for one. The useful framing is marginal: for a fixed total budget, would the next unit of compute do more good exploring an additional state or judging the states you already have more carefully? That depends on where the bottleneck sits, and there are two distinguishable regimes. **Generation-limited.** The correct approach is rarely among the candidates produced. Evaluation is doing its job — it correctly ranks what it is given — but what it is given is poor. Spending more on judgement is wasted; the budget belongs upstream. **Evaluation-limited.** Good candidates are present but the search does not reliably pick them, and runs frequently expand branches that were visibly worse than a sibling. Here more or better judgement converts directly into accuracy, and evaluation deserves the larger share. The diagnostic that separates them is worth naming: on solved problems, check whether the winning state was among the candidates generated at each node. If it usually was and the search still went elsewhere, you are evaluation-limited. If it usually was not, no amount of scoring will help. ## The lever that changes the frontier Before reallocating between the two, most systems can simply make evaluation cheaper for the same signal, by not applying uniform effort to every node. A cascade does this. A deterministic check runs first — free, and it removes states that are invalid on their face. A cheap pass runs next, using a smaller model, a shorter prompt, or a single sample rather than five. Only the states that survive and remain genuinely ambiguous receive full, multi-sample judgement. Because most states are either obviously hopeless or obviously fine, expensive judgement ends up applied to a small minority, and total evaluation spend falls substantially with little accuracy change. The second lever is spending unevenly by position rather than by state. Judgement near the root is worth more than judgement near the leaves: pruning a root branch discards an entire subtree, so an error there is expensive and irreversible, while a mistake one step from a leaf costs almost nothing. Allocating more samples shallow and fewer deep spends the same money where it changes more outcomes. A third is confidence-adaptive sampling. Start with one evaluation sample per state and pay for more only where the state sits near the pruning cutoff or where the first samples disagree. States that are obviously fine or obviously dead do not need five opinions. ## Knowing when the whole thing does not pay The most valuable judgement at this level is the one that questions the premise. A tree is worth its multiplied cost only when branching plus selection beats a single linear pass by enough to justify the difference. If tuning the evaluation split moves accuracy by a point or two while the run costs many times a straightforward attempt, the honest conclusion may be that this task does not reward search — and the budget belongs in a better prompt, a stronger model, or a real verifier instead of a wider tree. That also reframes the allocation question. Evaluation spend is not competing only with generation spend; it competes with every other way of improving the system for the same money. ## How to run this in practice Instrument the split so the two costs are visible separately in every run; teams that report only a total cannot reason about the trade at all. Then run a small sweep: hold the budget fixed, vary the samples-per-state and the branching factor together, and plot accuracy against the split. The curve is usually flat over a broad middle, which is good news — it means precision is not required, only avoiding the extremes of a single noisy sample per state or an evaluator so expensive that the tree can only afford to be narrow. Re-run the sweep when the model, the evaluator prompt or the workload mix changes, since all three move the frontier.
- How would you tell whether a run is generation-limited or evaluation-limited?Check, on problems you eventually solved, whether the winning state was among the candidates offered at each node. If good candidates were present and the search expanded worse siblings instead, judgement is the bottleneck and deserves more budget. If the winning line was rarely among the candidates at all, no amount of extra scoring helps and the spend belongs upstream in generation. Without this check, teams reallocate by intuition and usually guess wrong.
- Where in the tree is an extra evaluation sample worth the most?Near the root. Discarding a shallow branch throws away everything below it, so an evaluation error there is expensive and effectively irreversible, while a misjudgement one step from a leaf wastes almost nothing. Allocating more samples shallow and fewer deep spends the same total where it changes more outcomes. The same logic favours paying for extra samples on states sitting close to the pruning cutoff, where the decision is genuinely uncertain.
- When would you conclude the tree itself is not worth its evaluation cost?When the accuracy gain over a single linear pass is small relative to the multiple you are paying, and no split of the budget between generating and judging closes that gap. At that point the marginal money is better spent on a stronger model, a better prompt, or a real verifier than on a wider search. Recognising this is more valuable than tuning the ratio, because search only pays on tasks where selection among alternatives genuinely beats one careful attempt.
saying these in an interview costs you the question
- Assuming generation always dominates a reasoning tree's cost
- Reporting one total cost without splitting generate versus evaluate
- Applying full multi-sample judgement to every single state
- Treating the split as a constant rather than a per-workload measurement
- Never asking whether the search beats one linear pass at all