Why is using the same model to propose and score Tree of Thought states risky?
answer
- the judge shares the author's blind spots
- scores bunch high, nothing gets pruned
- cost of a tree, benefit of a chain
- hide the reasoning trace from the judge
- external checks beat self-opinion
basics
~20 sA model tends to rate its own reasoning favourably, so scores skew high and stay bunched together. The search then finds nothing to prune, keeps expanding weak branches, and spends the extra compute of a tree without gaining the accuracy that pruning was supposed to buy.
solid answer
~50 sWhen the proposer and the evaluator are the same model with the same context, the evaluation is not independent evidence — it is the same reasoning process asked a second time. Two effects follow inside the search. Scores compress toward the top of the scale, because errors the model could not see while generating are errors it cannot see while judging either, so genuinely bad states rarely get a low value. And the ordering between siblings becomes weak, because all candidates came from the same distribution and look equally plausible to their author. The search consequence is concrete: pruning stops biting, the frontier grows, and you pay for width without the selectivity that justified branching. Practical mitigations are to judge the state without the reasoning trace that produced it, to use a different model or configuration for scoring, and above all to anchor scores against any external signal the task offers.
go deeper
Know that if the same model both writes a step and grades it, the grade tends to be generous, so the scores are weaker evidence than they look.
Be able to explain the mechanism and its search consequence: shared blind spots compress scores toward the top, pruning stops biting, and the tree costs far more than a linear pass while adding little accuracy.
Show how you would detect it — a labelled probe set of known-good and known-bad states, plus the live score histogram — and how you would mitigate it with hidden traces, a separate judge, and external verification signals wherever the task provides them.
Own the position that evaluator independence is an architectural property of the system, and decide where the organisation invests in real verification signals versus accepting a cheap, biased judge with a documented ceiling.
## The setup that creates the problem The default Tree of Thought implementation uses one model twice: once to generate candidate next steps, once to score the states those steps produce. It is the cheapest arrangement and the natural place to start. It also makes the evaluator's judgement statistically dependent on the generator's, and search relies on that judgement being informative. The intuition is simple. If a model produced a step, it did so because that step looked correct to it. Asking the same model, moments later and often with the same reasoning in context, whether that step looks correct is not an independent test. Errors invisible during generation are usually invisible during evaluation, for the same reason: the model's understanding of the problem is what produced them. ## What it looks like inside the search Two symptoms show up, and both are observable without any special instrumentation. **Score compression at the top.** The distribution of values bunches into a narrow band near the high end. Very few states get decisively rejected. This is the more damaging symptom, because absolute thresholds depend on bad states actually scoring low — if nothing falls below the cutoff, pruning is decorative. **Weak separation between siblings.** All candidates at a node came from the same generator and share its assumptions, so to their author they look comparably reasonable. Rankings become close to arbitrary, which means the search's choice of what to expand carries little signal. The combined effect is the failure mode worth naming in an interview: the tree keeps its cost and loses its benefit. Branching is only worth its multiplied token spend because evaluation lets you concentrate effort on promising states and abandon the rest. An evaluator that approves nearly everything turns the search into breadth for its own sake — many times the cost of a single linear pass, with accuracy that barely moves. A related trap is that the search can lock onto one line of reasoning. If the generator commits to an approach early and the evaluator, sharing that commitment, scores every state along that line highly, alternatives that would have required abandoning the premise get scored down. The tree flattens into an expensive chain. ## How to detect it Do not take the scores at face value; test them against something you already know. Build a small set of states whose quality is known — some from problems that were solved correctly, some from branches that demonstrably failed — and run the evaluator over them. If the scores do not separate the two groups, the evaluator carries no usable signal for this task, whatever it says about individual nodes. Also look at the score histogram from real runs. A healthy evaluator produces spread, with a genuine low tail. A distribution squeezed into the top fifth of the scale is the compression symptom, and it usually means the threshold is either doing nothing or has been nudged so high that it is pruning on noise. ## What actually helps **Hide the trace.** Present the state to be judged on its own, without the generation reasoning that produced it. The rationale in context is a strong cue toward endorsement; removing it makes the judgement about the state rather than about the argument for it. **Change who judges.** A different model, or at minimum a different configuration, breaks some of the dependence. It is not a complete fix — models share training influences and often share blind spots — but it reliably widens the score distribution compared with pure self-scoring. **Prefer any external signal.** This is the strongest correction and the reason it dominates in practice. Where the task offers a check that is not the model's opinion — a test suite, a constraint solver, a compiler, a lookup against known reagents or known facts — that signal should dominate the score. Self-judgement fills the gaps it cannot cover, not the other way round. **Reframe as comparison.** Asking which of several states is most promising, rather than assigning each an absolute value, is somewhat more robust than absolute self-scoring, because relative judgements are easier. It does not remove shared blind spots, and it gives up the absolute floor that pruning wants. **Ask for the objection first.** Prompting the evaluator to state the strongest reason the state might fail, before it produces a verdict, raises the rate at which weak states are caught. It is a partial mitigation, not a substitute for an external check — the model still cannot object to errors it does not perceive. ## The honest summary Self-evaluation is not worthless; a model can often reject a state that violates a stated constraint or contradicts itself. It is simply weakest exactly where search needs it most — on subtle errors that the same model made a moment earlier. Design around that: use it as a cheap first filter, anchor it to external signals wherever the task provides them, and verify it separates known-good from known-bad states before trusting it to steer an expensive tree.
- Does simply swapping in a second model as evaluator fix it?It helps but does not solve it. A different model breaks the direct dependence on the generator's own reasoning and typically widens the score distribution, which restores some pruning power. But models trained on overlapping data share blind spots, so a class of errors survives both. Treat cross-model judging as variance reduction, and reserve real confidence for signals that are not another model's opinion — tests, solvers, constraint checks, ground-truth lookups.
- How would you show an interviewer that your evaluator has this problem?Run it over a labelled probe set: states drawn from branches that reached correct answers and states from branches that demonstrably failed. If the score distributions of the two groups overlap heavily, the evaluator is not separating quality regardless of how confident individual scores look. Plotting the score histogram from live runs is a quicker tell — a distribution crushed into the top of the scale with no low tail is the signature.
- If self-evaluation is this weak, why is it still the default?Because it is free to build, works on any task you can describe in words, and is genuinely useful for the errors it can see — violated constraints, self-contradiction, states that ignore the goal. Most tasks offer no external verifier at all, so the realistic choice is self-judgement or no evaluation. The senior move is to treat it as a cheap first filter with known limits, not as a trustworthy value function.
saying these in an interview costs you the question
- Assuming a model reliably catches its own reasoning errors
- Ignoring that self-scores cluster high and never prune
- Trusting confident-looking scores without a probe set
- Using self-evaluation where a real verifier exists
- Believing more branches compensate for a biased evaluator