What do you log per node so a bad Tree of Thought run can be replayed?
answer
- Attribute failure: propose, score, or prune
- Keep raw output, not just parsed
- Parent pointers rebuild the shape
- Version the template, not just the prompt
- Cache responses to replay for free
basics
~20 sRecord one row per node: id, parent id, depth, the rendered prompt and template version, model and sampling parameters, the raw proposer and evaluator responses, the parsed score, tokens, latency, retries and the prune reason. That log makes a run replayable and its failure attributable.
solid answer
~50 sTreat the tree as an edge list you can reconstruct offline. Each node record carries its id and parent id, its depth, the exact rendered prompt (or a hash plus the template version and the state it was rendered from), the model and sampling parameters used, the raw response before parsing, the parsed thought, the evaluator's raw output and parsed score, token counts, latency, retry count, and whether the node was pruned and by which rule. A run-level record holds the configuration, the budget, which limit tripped and the final selection path. The payoff is that a wrong answer becomes a diagnosable question with three distinct answers: the right step was never proposed (a generation problem), it was proposed but scored badly (an evaluation problem), or it scored well and was pruned anyway (a control-policy problem). Without raw responses you cannot tell those apart. Sample full payloads rather than storing every one if volume or data sensitivity demands it, but always keep the metadata for every node.
go deeper
Know that each node should be logged with an id, its parent, the prompt sent, the response received and the score, so the tree can be looked at afterwards.
Explain how parent pointers reconstruct the tree and why raw responses must be kept beside parsed ones — a discarded malformed answer looks like a missing idea otherwise.
Attribute a bad run to generation, evaluation or control using the log, and show how a prompt-hash response cache replays a historical run offline to test a policy change at zero token cost.
Own the tradeoff between diagnosability and data exposure: metadata for every node always, sampled and redacted payloads with a retention window, and versioned templates so cost and quality regressions are diffs rather than investigations.
## The failure you are instrumenting for A tree run produces one answer from dozens of calls. When the answer is wrong, the useful question is not 'was the model bad' but *where in the pipeline the right path was lost*. There are exactly three places, and each has a different fix: 1. **Generation** — the correct next step was never proposed at any node. Fix the proposer template, the sampling temperature, or the branching width. 2. **Evaluation** — the correct step was proposed but scored below its siblings. Fix the scoring contract, swap in a verifier, or change what state the evaluator sees. 3. **Control** — the correct step was proposed and scored well, and the search still did not follow it. Fix the traversal or pruning policy. A log that cannot distinguish these three is not instrumentation, it is noise. That requirement alone dictates most of the schema: you must keep raw proposer output (for 1) and raw evaluator output with parsed scores (for 2) and the pruning decisions with their reasons (for 3). ## The node record A workable schema, one row per node: - **identity** — node id, parent id, depth, and the run id; - **input** — the rendered prompt, or a content hash plus the template id and version plus the serialized state, so the prompt can be reproduced exactly; - **model config** — model name, temperature, top-p, max output tokens, and any seed; - **output** — the raw response text, and the parsed artefact (the step); - **evaluation** — the evaluator's raw response, the parsed score, the evaluator's model config, and whether the score came from a model or a verifier; - **economics** — input tokens, output tokens, cached-prefix tokens if reported, latency, retry count; - **fate** — expanded, pruned (with the rule that pruned it), terminal, or unscored due to error. And one run-level record: configuration, budget, which limit stopped the run, wall-clock, total spend by role, and the final selected path as a list of node ids. ## Why template versions matter more than they look Prompts change. If a node record stores only the rendered prompt, you can replay it but you cannot answer 'did this run use the new evaluator wording'. If it stores only the template id, you cannot see what the state rendered to. Store both — the version and the rendered text (or its hash plus the state) — and comparisons across runs become mechanical. A cost or quality regression after a prompt edit then shows up as a diff between two run summaries rather than an investigation. ## Replay, cheaply Because every call is a pure function of its prompt, a keyed store of prompt hash to response gives you deterministic offline replay: re-run the orchestrator against the cache and it walks the identical tree without spending a token. That is how you test a change to the *control policy* — new pruning rule, different width — against a real historical run at zero model cost, since those changes reuse the same proposals. It stops working the moment you change a prompt (the hashes miss), which is exactly the right boundary: prompt changes need fresh calls, policy changes do not. ## Reconstructing and reading the tree Parent pointers alone reconstruct the shape. Once you can render the tree, a few views do most of the diagnostic work: the score distribution per depth (flat scores mean the evaluator is not discriminating); the ratio of pruned to expanded nodes (too high means the search is narrow and brittle, too low means you are paying for breadth you never use); spend by depth (usually shows a deep level dominating); and the selected path highlighted against its rejected siblings, which is the view a human actually reads when arguing about a wrong answer. ## Volume, cost and sensitivity Raw prompts and responses are the bulk of the data and the part with the compliance weight — they contain whatever user or business content the task carried. Two habits keep this manageable: log metadata (ids, scores, tokens, latency, fate) for every node unconditionally, since it is small and it is what most dashboards need; and log full payloads for a sample of runs, plus all runs that ended badly, with an explicit retention window and redaction of fields you should not keep. Storing the state that produced a prompt rather than the fully rendered prompt is also a useful compression, provided the renderer is versioned. ## The habit that separates people who have shipped this Assign the node id at *request* time, from the parent id and the request index, not from response order. Everything above depends on ids being stable and reproducible; ids derived from arrival order make two replays of the same run disagree about which node was which, and quietly destroy the comparison you built the log for. ## What interviewers listen for The three-way failure attribution, raw responses kept alongside parsed ones, prune reasons recorded rather than implied, template versioning, and a realistic position on payload volume and sensitivity. 'We log the final answer and the token count' is the answer of someone who has not yet had to explain a bad tree.
- Which recorded field tells you whether the correct step was ever proposed at all?The raw proposer responses for every expanded node, before parsing. Parsed steps are not enough, because a malformed response containing the right idea gets discarded at parse time and would look identical to a node that never suggested it. Keeping the raw text is what separates 'the model never thought of it' from 'we dropped it'.
- How does a prompt-hash-to-response cache let you test a change for free, and where does that stop?Every call is a pure function of its prompt, so replaying the orchestrator against cached responses reproduces the identical tree at zero token cost. That covers control-policy changes — different width, new pruning rule, different traversal — because they reuse the same proposals. It stops the moment you edit a template: the hashes miss, and the change genuinely needs fresh calls.
- What do you do about raw prompts and responses when the task carries sensitive data?Split the tiers. Metadata — ids, scores, tokens, latency, fate — is small and safe, so log it for every node always. Raw payloads get sampled, redacted on the fields you must not retain, and given an explicit retention window, with full capture forced for runs that failed. Storing the serialized state plus a versioned renderer instead of the rendered prompt shrinks the sensitive surface further.
saying these in an interview costs you the question
- Logs only the final answer and total token count
- Keeps parsed scores but discards raw model responses
- Prunes nodes without recording which rule pruned them
- Assigns node ids from response arrival order
- Stores prompts with no template version, so runs cannot be compared