In Tree of Thought, when does an external verifier beat an LLM evaluator?
answer
- Ground truth beats judgement
- Verifiers refute, models rank
- Binary check, no ordering signal
- Hard filter first, model scores survivors
- Run it against a shadow, never production
basics
~20 sAn external verifier — a compiler, a test suite, a schema check, a query planner, a solver — beats a model evaluator whenever a candidate state can be checked mechanically. It refutes soundly, deterministically and cheaply, but it rarely ranks the candidates that pass.
solid answer
~60 sPut a verifier in the evaluator slot whenever the domain gives you ground truth. Planning a database migration, for example, you can render each candidate plan against a shadow schema and run the planner over the resulting queries: a step that fails to parse, drops an index the workload needs, or produces a sequential scan on a hot path is *known* bad, not *judged* bad. That gives you three things a model evaluator cannot: soundness on rejection, determinism across runs, and independence from the model that proposed the candidate, which removes self-preference bias. The catch is coverage and granularity. A verifier is usually partial — it says pass or fail, not 'this is the more promising of two valid plans' — and it often cannot judge an incomplete state at all. The standard resolution is a hybrid: the verifier is a hard filter that prunes invalid children, and a model scores only the survivors for ordering. Feeding the verifier's error text back into the proposer as context for the next attempt is the other half of the win.
go deeper
Know that the scoring step does not have to be the model — a compiler, a linter or a test run can decide whether a candidate step is valid.
Explain what a verifier makes sound and deterministic, and why a binary pass/fail cannot order the candidates that all pass.
Argue the hybrid concretely: verifier as a cheap hard filter, model scoring only the survivors, failure text fed back to the proposer, and every check run against a disposable shadow environment.
Treat the verifier as a system you now operate — its flakiness, latency and coverage gaps become search-quality risks, and 'what does it fail to check' belongs in the design review, not the retro.
## The evaluator slot is pluggable Because a tree run separates proposal from evaluation, whatever fills the evaluator slot only needs to accept a serialized state and return something the control policy can compare. Nothing requires it to be a model. In any domain with a checker, using that checker is the strongest available upgrade to a tree search, because it replaces a judgement with a fact. The usual candidates: a compiler or type checker for code steps; a unit or property test suite for a completed function; a schema or config linter for infrastructure changes; a query planner run against a shadow copy of a schema for migration steps; a constraint or LP solver for scheduling; a simulator for control policies; a parser for anything with a grammar. ## What a verifier buys **Soundness on rejection.** When a compiler says the code does not compile, that is not an opinion. A model evaluator can be talked into approving something malformed, especially when the malformed thing reads fluently. **Determinism.** The same state produces the same verdict every time, which makes the search reproducible and makes a regression attributable to a proposal change rather than to scoring noise. **Independence.** The verifier has no stake in which model wrote the candidate, so the self-preference bias that afflicts model-judged pipelines disappears from the pruning decision. **Cost profile.** Running a linter or planning a query is typically far cheaper and faster than an extra model call, so verifier-first pruning can shrink the frontier before you spend model tokens on it. ## What a verifier cannot do **Rank.** Most checkers are binary. Once ten candidate migration steps all plan successfully, the planner has nothing more to say about which is the better next move — and ranking is exactly what a search needs to guide expansion. This is the reason 'just use tests instead of a model' is not a complete answer. **Judge partial states.** A verifier usually wants a well-formed artefact. Halfway through a plan there may be nothing to compile and no query to plan, so the evaluator slot is empty precisely where the search needs guidance most. Some domains work around this with a relaxed check (does the partial plan still satisfy the invariants so far?), but the relaxation is itself engineering work. **Cover intent.** Passing tests, a clean lint and a good query plan together do not mean the plan achieves the goal. A verifier constrains the space; it does not confirm the destination. ## The hybrid that people actually ship Run the verifier first as a **hard filter**. Children that fail are pruned immediately and cheaply, with the failure text recorded. Children that pass go to a model evaluator that only has to answer the softer question — which of these valid options looks most promising — which is a question models are comparatively good at, and now they are only asked about states that are known well-formed. The compute saving is real: you stop paying model tokens to score candidates a parser could have rejected. The second half of the hybrid is the feedback path. A verifier failure carries a *reason* — a compiler error, a failing assertion, an index the planner ignored. Feeding that text into the proposer when it retries that region of the tree converts a dead branch into information. This is not the same as a general critique loop: the signal is a fact from a tool, and it is scoped to one node's regeneration. ## The operational catches **Side effects.** Running real artefacts is the whole point, and it is also the danger. Verification must happen against a disposable copy — a shadow schema, a scratch branch, an ephemeral container — with no path to production data. Never verify a destructive candidate by executing it somewhere that matters. **Latency and concurrency.** A verifier that takes seconds per call may be slower than the model it replaced, and running dozens concurrently may need process or container isolation. This can quietly become the bottleneck of a level. **Flakiness.** A flaky test suite in the evaluator slot is worse than a noisy model, because you will trust it. Verifier flakiness must be detected and treated like an unscored node, not like a rejection. **Partiality bias.** If the verifier can only judge a subset of states, and unjudgeable states default to 'pass', the search drifts toward exactly the candidates nobody checked. Make the unjudgeable case an explicit, logged category. ## The interview answer Name a concrete verifier for a concrete domain, say what it makes sound, admit that it filters rather than ranks, and describe the hybrid plus the sandbox. Candidates who claim a verifier replaces the model evaluator entirely have usually not watched a frontier of ten equally-valid candidates with no ordering signal.
- Your migration verifier passes every candidate at depth three. What has the search lost, and what do you do?It has lost its guidance signal — a filter that rejects nothing provides no ordering, so expansion becomes uninformed breadth. Either strengthen the check so it discriminates (assert on the plan shape, on estimated rows scanned, on lock acquisition), or fall back to a model evaluator over the survivors for ranking. Recording that the verifier was non-discriminating at that depth is what lets you notice it at all.
- How do you evaluate a partial state that has nothing complete enough to verify?Relax the check to the invariants that must hold at every prefix — the plan so far still touches no locked table, the code so far still parses, the schedule so far has no conflict — and treat 'not yet checkable' as an explicit logged category rather than a silent pass. If unjudgeable states default to passing, the search will drift toward the ones nobody checked.
- What is the risk of feeding a verifier's raw error text back into the proposer?Two risks. The text is untrusted output from a tool and may be long, noisy, or contain data you would rather not put in a prompt, so it should be truncated and sanitized. And over-anchoring on one error can push the proposer into narrowly patching that symptom instead of choosing a different step, which is a common way a branch stops exploring and starts thrashing.
saying these in an interview costs you the question
- Claims a verifier can replace the model evaluator entirely
- Expects a pass/fail check to rank valid candidates
- Verifies a candidate by executing it against real data
- Defaults unverifiable partial states to pass
- Treats a flaky test suite as trustworthy ground truth