When is a learned value model worth training to score Tree of Thought states?
answer
- free, flexible, or fast — pick per node
- automatic labels are the deciding condition
- training cost amortizes only under volume
- learned scorers drift when the generator changes
- start prompted, log outcomes, train later
basics
~20 sTrain one when the same task runs constantly, ground-truth outcomes are cheap to collect, and prompted judging is either too noisy or too expensive per node. Otherwise a rule-based check or a prompted evaluator is faster to build and easier to change.
solid answer
~50 sA learned value model — a small scorer trained to predict, from a partial state, how likely continuation is to succeed — earns its keep under three conditions together. First, **volume**: the same task shape runs often enough that a fixed training cost amortizes. Second, **cheap labels**: the outcome is automatically checkable, so you can roll out states to completion and label them without human annotation. Third, **a real gap**: prompted evaluation is measurably noisy or its per-node cost dominates the run. A learned scorer is orders of magnitude cheaper per call and can be calibrated against real outcomes, which prompted judging cannot. Against that, it needs a training pipeline, it goes stale when the generator or task distribution shifts, and it fails opaquely. Where a deterministic rule already settles feasibility — a constraint violated, a budget exceeded — use the rule first regardless; it is free and always right.
go deeper
Know the three options — deterministic rules, asking a model, and a trained scorer — and that the trained one is cheap per call but expensive to build and keep working.
Be able to explain what a learned value model predicts from a partial state, where its labels come from, and why automatically checkable outcomes make training practical while human annotation usually does not.
Demonstrate the decision rule under real constraints: node volume, label cost, and measured evaluator noise. Name staleness as the operational failure mode and describe the monitoring that compares predicted scores against realized outcomes.
Own the build-versus-prompt call as a long-lived commitment: a value model is a second trained system with its own data pipeline, retraining trigger and on-call surface, and that ongoing cost has to be weighed against simply spending more on prompted evaluation.
## Three families of evaluator Scoring a partial reasoning state can be done in three ways, and they sit on a clear cost-and-effort ladder. **Rule-based or heuristic evaluators** are deterministic code: does this partial plan violate a hard constraint, does the intermediate step exceed the budget, does the partial program still parse? They cost nothing per node, never hallucinate, and are unambiguous when they fire. Their weakness is coverage — they answer feasibility questions, not promise questions. A partial solution can satisfy every rule and still be going nowhere. **Prompted evaluators** ask a general model to judge the state. They cover anything you can describe in words, need no training data, and can be changed by editing a prompt. They are also the expensive option: every node costs a model call, often several because judgements are sampled and aggregated, and the resulting scores are not calibrated to anything. **Learned value models** are trained scorers — typically a small model, or a regression or classification head over an encoder — that map a partial state to a predicted outcome. In the reasoning literature this family includes process reward models, which score intermediate steps rather than only final answers. Inference is cheap and fast; the cost has moved to building and maintaining the thing. ## What makes a learned evaluator pay Three conditions have to hold together. **The task recurs.** Training only makes sense if the same distribution of states appears again and again. A one-off analysis will never repay the pipeline. **Labels are cheap and automatic.** This is the decisive condition in practice. If you can run a partial state forward to completion and check the result mechanically, you can mint labels at scale with no annotators. A concrete case: scoring partial Python programs by predicted unit-test pass rate. Complete each partial program many ways, run the test suite, record the pass fraction, and train a regression head on (partial program → pass rate). The labels are exactly the quantity the search cares about, and no human touched them. **There is a measurable gap to close.** Either prompted judging is too noisy — its rankings disagree with known-good outcomes — or evaluation calls dominate the run's cost. If neither is true, a learned model buys nothing. ## What it costs you A learned evaluator introduces a second system to own. It needs training data collection, a training loop, versioning, and an offline eval that says whether a new checkpoint is actually better. It also **goes stale**. A value model learns the distribution of states produced by one generator on one task family. Swap the generating model, change the prompt that produces candidate thoughts, or shift the workload, and the states it sees at inference stop resembling its training data. Silent degradation follows: scores keep coming back with the same confident shape while their correlation with real outcomes has collapsed. Monitoring has to compare predicted scores against realized outcomes on a sample of completed searches, continuously. And it fails opaquely. A prompted evaluator that is wrong usually explains itself badly enough that a human can spot the error. A regression head returns 0.41 and offers nothing. ## The practical arrangement Most mature systems layer rather than choose. A deterministic rule pass runs first and eliminates states that are outright invalid — free, and it removes the easiest work from the expensive stages. Whatever survives gets the learned scorer if one exists, or a prompted evaluator if not. Some designs keep a prompted evaluator as a second opinion at high-stakes nodes only, near the root or where the learned scores cluster tightly and give the search nothing to separate. The migration path matters as much as the endpoint. Start prompted, because it works on day one and tells you whether evaluation signal helps at all. Log every state along with the outcome the search eventually reached. That log *is* your training set, gathered as a side effect of running the system. Only when the volume justifies it do you train, and even then the prompted evaluator stays as the fallback and the yardstick. ## How to argue it in an interview The answer an interviewer wants is not "learned models are better." It is a decision rule: rules where truth is deterministic, prompting where the judgement is linguistic and volume is low, learned scorers where outcome labels are automatic and node counts are high enough that per-call cost dominates. Naming the staleness failure mode and the monitoring that catches it is what separates a senior answer from a textbook one.
- Where does the training data for such a value model come from?From the search's own history. Log every partial state that was scored along with what eventually happened on that branch — did the run reach a verified-correct answer or not. When the outcome is automatically checkable, this is free supervision accumulated while the prompted system runs. The subtlety is bias: your log only contains states the search chose to expand, so states pruned early are under-represented and the model is weakest exactly where the search already distrusts.
- How would you notice a learned evaluator has gone stale?Compare predictions against realized outcomes continuously, not just at training time. Sample completed searches, take the score the evaluator gave each state, and check whether it still correlates with whether that branch succeeded. Watch the score distribution too — a collapse toward the middle, or an upward drift, usually means the incoming states no longer resemble the training distribution. Changing the generator or the thought-generation prompt should trigger a re-check by default.
- Would you ever keep a heuristic evaluator once a learned one exists?Yes, in front of it. Deterministic checks — constraint violations, parse failures, budget overruns — are free, always correct, and cover cases where a learned model can only guess. Running them first also removes the easy rejections before you spend anything, so the learned scorer only sees states where judgement is actually required. The two are complementary layers, not competing choices.
saying these in an interview costs you the question
- Assuming a trained scorer is always better than prompting
- Ignoring that a value model drifts when the generator changes
- Planning to hand-label intermediate states at scale
- Skipping the free deterministic checks entirely
- Never measuring whether evaluator scores predict real outcomes