skip to content

Thought Trace Design

What a good thought step looks like: enough reasoning to ground the next action, not so much that the model talks itself into a hallucination. Interviewers ask about the verbosity trade-off, since longer traces cost tokens and can add confident fiction.

on this pageshow

questions

4

In a ReAct agent, what does a well-formed thought step actually contain?

level: middleimportance: must knowfreq 66%

answer

  1. input to the next decision, not a log
  2. goal, gap, rationale
  3. three clauses, not three paragraphs
  4. apply the deletion test per sentence
  5. if it can't change the action, cut it

basics

~20 s

A well-formed ReAct thought states where the task stands against the goal, names the one piece of information still missing, and justifies the next action as the way to get it. Text that does not change which action follows is decoration.

solid answer

~50 s

I treat the thought as the bridge between the last observation and the next action, so it needs three things and little else: where we stand against the goal, what specific information is still missing, and why the chosen action closes that gap. The practical test is deletion — if I remove a sentence and the agent would still pick the same tool with the same arguments, that sentence was not doing work. A good thought also reasons from what the last observation actually said rather than from free-floating recall, which keeps the step grounded. On a procurement task, "I have the vendor name but not the supplier id; the vendor lookup returns supplier ids, so query it by name" is a complete thought. Restating the entire history, narrating persona, or pre-announcing the final answer adds tokens and dilutes the signal the model attends to on the next step.

code

markdown · 9 lines
markdown
Weak thought (decoration, no constraint):
"Let me carefully consider this procurement request. Purchasing involves
approvals, budgets and supplier relationships. I should probably look
something up about the vendor."

Strong thought (three clauses, fully determines the call):
"Have: vendor name, request date. Missing: supplier id, required by the
approval check. The vendor lookup maps name to supplier id, so call it
with the vendor name I already have."

go deeper

for a junior

Be able to say plainly that the thought is where the agent decides what to do next, and that it should name what is missing and which tool will supply it.

for a middle

Explain the three clauses — position, gap, rationale — and demonstrate the deletion test: any sentence that would not change the chosen tool or its arguments should be removed.

for a senior

Show judgment about grounding: thoughts must reason from what the last observation actually returned, and any unverified belief has to be labelled so a later step can check it rather than inherit it as fact.

for a principal

Own the policy question. Decide what thought discipline applies per task class, how you measure whether reasoning text improves action selection at all, and how that policy trades off against per-step token cost across a fleet of agents.

## What the thought is for In a ReAct-style agent the model alternates between writing reasoning text and emitting a tool call. The reasoning text — the thought — is not output for the user and it is not a log line for you. It is *input to the model's own next decision*: it sits in the transcript and is re-read when the model chooses the next action. That single fact determines everything about how a thought should be written. A thought earns its place only if it makes the next action better chosen or better parameterised. Everything else is cost. ## The three clauses A thought that does its job almost always carries three clauses, and they can be very short. **1. Position against the goal.** What has been established so far, stated in terms of the task, not the transcript. "I have the contract's start date and the vendor name" — not "I called two tools." **2. The missing information.** The single unknown that currently blocks progress, named concretely. "I still need the supplier id." Vagueness here is the leading cause of a badly chosen action: if the thought says "I need more details about the vendor," the model has given itself no constraint, and it will pick a plausible-looking tool with plausible-looking arguments. **3. The rationale for the action.** Why *this* action supplies *that* information. "The vendor lookup returns supplier ids keyed by vendor name, so I will call it with the vendor name I already have." This clause is what makes the argument values fall out of the reasoning rather than being invented at emission time. ## The deletion test The cleanest editorial rule is a counterfactual: delete a sentence and ask whether the agent would still select the same tool with the same arguments. If yes, the sentence was decoration. Applied honestly, this removes most of what long traces contain — recaps of observations that are still visible verbatim a few lines above, restatements of the system instructions, hedged musing about approaches the agent is not going to take, and confidence performance ("Great, I now have everything I need!"). ## Ground the thought on the observation A thought that follows an observation should reason from what the observation actually returned. "The lookup returned two vendors with that name, so the id is ambiguous and I need the region to disambiguate" is grounded: every claim traces to text the agent can see. "The vendor is probably the European entity" is not — nothing observed said so. Grounding is not stylistic; it is the main defence against the trace accumulating invented facts that later steps treat as settled. A useful habit is to make thoughts distinguish what was *observed* from what is *believed*, so a belief can be marked as needing verification rather than silently promoted to fact. ## What does not belong in a thought - **The answer, written early.** Once a thought states a conclusion, the remaining steps tend to rationalise toward it rather than test it. - **Facts the agent has not fetched.** Asserted policy limits, schema field names, ids, thresholds. If it was not observed, it is a hypothesis and should be labelled as one. - **Full re-narration of prior steps.** The transcript already contains them; repeating them buys nothing and pushes the useful content further from the generation point. - **Tool syntax.** How to format the call is the action's job, not the thought's. ## Why short thoughts usually win Every thought token is paid once when written and again on every subsequent step that re-reads the transcript, so per-step verbosity compounds across a long run. More importantly, a long thought raises the chance that at least one sentence asserts something unsupported, and unsupported assertions are sticky: the agent reads its own prose back as if it were evidence. The failure mode of terse thoughts is a poorly chosen action, which the next observation usually exposes. The failure mode of florid thoughts is a confidently wrong plan that survives several steps. The second is far more expensive to recover from. ## A worked contrast Weak: "Let me carefully consider the procurement request. There are many aspects to purchasing workflows, including approvals, budgets and supplier relationships. The user seems to want information about a supplier, and suppliers are usually identified by internal identifiers in most ERP systems. I should probably look something up." Strong: "Have: vendor name, request date. Missing: supplier id, which the approval check requires. Vendor lookup maps name to supplier id, so call it with the vendor name." The second is a third the length and completely determines the next call.

  • How would you tell, at review time, whether a thought was actually doing work?
    Run the deletion counterfactual on real traces: strip a thought's sentences and re-sample the action. If the tool and its arguments are unchanged, that text was decoration. At scale I compare tool-selection accuracy and argument correctness between a terse and a verbose thought policy on the same eval set — if the verbose one does not move either metric, the extra prose is pure cost.
  • Should the thought ever contain a belief the agent has not verified?
    Yes, but it must be marked as one. "Likely the European entity — unverified" is safe because later steps can see it is a hypothesis and choose a verifying action. The dangerous form is the same belief stated flatly, because the agent reads its own prose back as established fact and stops questioning it. Explicit hedging is cheap insurance.
  • Where should the thought stop and the action begin?
    The thought owns intent — which gap is being closed and why this capability closes it. The action owns encoding — the tool name and the argument values. Blurring them wastes tokens on syntax the model has to produce anyway, and it tempts the thought into inventing argument values before the reasoning has justified them.

saying these in an interview costs you the question

  • Thinks the thought is a log line for humans, not model input
  • Writes the final answer into the thought before any evidence arrives
  • States facts the agent has not observed as if established
  • Re-narrates the whole transcript in every thought step
  • Believes longer reasoning is always higher-quality reasoning

context

open as a page

How does a ReAct thought step turn into a hallucination the agent acts on?

level: seniorimportance: must knowfreq 58%

basics

~20 s

A thought is free text the model writes and then re-reads as if it were evidence. When it asserts something no observation supports — a policy limit, a field name, an id — later steps inherit that invention as established fact and act on it without ever checking.

open as a page

In a long ReAct run, why should thoughts restate the goal and remaining steps?

level: seniorimportance: should knowfreq 44%

basics

~20 s

On long runs the original instruction sits far behind a wall of observations, and its influence on the next token weakens. Periodically re-writing the goal and what remains puts the objective back near the point of generation, which is the cheapest defence against drift.

open as a page

How verbose should ReAct thought traces be when running agents at high volume?

level: principalimportance: should knowfreq 37%

basics

~20 s

There is no universal answer — set it per task class and prove it with an ablation. Thought tokens are billed on write and again on every later step that re-reads them, so verbosity compounds across a run while adding little to action quality on simple steps.

open as a page