How verbose should ReAct thought traces be when running agents at high volume?
answer
- billed on write, billed again on every re-read
- allocate verbosity, do not set it globally
- hard steps earn prose, mechanical steps do not
- measure tokens per successful task
- longer traces invite unsupported assertions
basics
~20 sThere is no universal answer — set it per task class and prove it with an ablation. Thought tokens are billed on write and again on every later step that re-reads them, so verbosity compounds across a run while adding little to action quality on simple steps.
solid answer
~50 sI treat trace verbosity as a policy with a measurable price, not a style preference. A token written in an early thought sits in the transcript and is re-read on every subsequent step, so per-step verbosity scales roughly with the square of the run length in a long trajectory — a terse "need the supplier id next" versus a 200-word essay per step is not a small difference by step thirty. Against that, verbosity buys something real only on steps with genuine ambiguity: multiple constraints to reconcile, several plausible tools, arguments that must be derived. So I run the ablation — terse policy versus verbose policy on the same eval suite — and look at tool-selection accuracy, argument correctness and end-task success, not at how thoughtful the traces read. In practice cheap steps get terse thoughts, hard steps get room, and the default leans terse because long thoughts also invite unsupported assertions.
go deeper
Know that the agent's reasoning text costs tokens like everything else, and that padding each step adds up quickly over a long run.
Explain the compounding: a thought written early is re-read on every later step, so per-step verbosity scales badly with run length, while mechanical steps gain nothing from prose.
Show you would settle it empirically — an ablation of terse versus verbose on a fixed eval suite, scored on tool-selection accuracy, argument correctness and repeated end-task success rather than on how the traces read.
Own the policy across task classes and its review cadence: where verbosity is granted by exception, what it costs per successful task at fleet volume, how it trades against latency and fabrication risk, and why the answer must be re-measured on every model change.
## Why this is a cost question, not a style question In a ReAct loop the transcript grows monotonically: each step's thought joins the context that every later step reads. A thought written at step one in a thirty-step run is processed roughly thirty times. That makes per-step verbosity superlinear in run length — doubling the average thought length in a long trajectory costs far more than double. At single-request scale nobody notices; at fleet scale, across thousands of runs a day, it is one of the larger line items you control without changing the model or the task. Providers' prompt-reuse mechanisms reduce the cash cost of re-reading a stable prefix, but they do not reduce the two other effects: the growing transcript still consumes the window, and the model still has to attend across all of it. So the argument for terseness survives even where the re-read is discounted. ## What verbosity actually buys It is not true that shorter is always better. Reasoning text earns its cost when the step is genuinely hard: - **Multiple constraints to reconcile.** A step that must satisfy a date window, an entitlement and a budget cap benefits from writing them down before choosing. - **Ambiguous tool choice.** Two tools with overlapping purposes; the thought is where the disambiguation happens. - **Derived arguments.** When the argument is computed from several observations rather than copied from one. And it stops buying anything on the many steps that are mechanical: fetch the next record, retry with the corrected field, read the item the plan already named. Those steps have one obvious action, and prose about them is pure tax. That asymmetry is the actual design insight — verbosity should be allocated, not set globally. Cheap, determined steps get one line; genuinely branching steps get room. ## The ablation is the only honest answer Interviewers ask this question to see whether you measure. The experiment is straightforward: hold the task suite and the tools fixed, vary only the thought-length policy in the instructions, and compare on metrics that reflect decisions rather than prose quality: - **Tool-selection accuracy** — did it pick the right capability? - **Argument correctness** — were the parameters right? - **End-task success**, run repeatedly, since agent runs are nondeterministic and a single pass will mislead you. - **Step count and redundant actions** — verbose traces sometimes reduce flailing, which can offset their token cost. - **Tokens and wall-clock per successful task**, which is the number that actually matters; tokens per run is the wrong denominator because a cheap run that fails costs infinity. Run it per task class. It is common for the terse policy to win on retrieval-shaped work and lose on multi-constraint planning work, which is precisely why a single global setting is the wrong artefact. ## The quality risk of long thoughts Cost is only half the case. Longer thoughts have more surface area for assertions the agent has not verified, and those assertions are re-read later as if they were evidence. Long thoughts also tend to reach a conclusion early, after which the remaining steps rationalise rather than test. And in a long transcript, verbose reasoning dilutes the observations — the actual evidence sits further apart, surrounded by the model's own narration. "Signal dilution" is the standard framing: you are lowering the ratio of grounded text to generated text in the very context the next decision reads. ## Latency and the human factor Thought tokens are generated serially before the action can be emitted, so verbosity is also a per-step latency cost on an interactive agent. Multiply by the number of steps and a chatty policy can add meaningful wall-clock to a user-facing task. There is a counter-pressure worth naming honestly: humans debugging agents like long traces. Rich reasoning is easier to review, and teams often mandate verbosity for observability rather than for accuracy. That is a legitimate goal but the wrong lever — the cheaper answer is structured logging of the decision inputs around the loop, not paying the model to narrate itself into every subsequent step's context. ## The policy I would actually set Default terse: goal, gap, rationale in one or two lines. Grant extra length by exception, on the task classes where the ablation shows it pays. Cap it, because unbounded thoughts are where fabrication accumulates. Re-run the ablation when the model changes, because the tradeoff is model-dependent and today's answer does not carry forward — newer models often need less explicit scaffolding to reach the same action quality, which shifts the optimum toward terseness over time.
- Which metric would you use to compare a terse and a verbose thought policy?Tokens and wall-clock per *successful* task, not per run — a cheap policy that fails more often is not cheap. Alongside that, tool-selection accuracy and argument correctness show whether the reasoning is doing decision work, and repeated runs are essential because agent trajectories are nondeterministic and a single pass will point you the wrong way.
- Doesn't prompt reuse make the re-read cost argument moot?It reduces the cash cost of a stable prefix, not the underlying problem. The transcript still consumes the context window, the model still attends across all of it, and long thoughts still dilute observations and invite unsupported claims. Caching changes one term in the cost equation; it does not change the quality argument for terseness.
- Teams often mandate verbose traces for debuggability. How would you handle that?Grant the goal, refuse the lever. Debuggability is better served by structured logging around the loop — the observation, the selected tool, the arguments, the outcome — which costs nothing at inference and is queryable. Paying the model to narrate for human readers means those tokens are re-read by every later step and can steer behaviour, which is a strange price for observability.
saying these in an interview costs you the question
- Treats trace length as a style preference rather than a measured cost
- Ignores that early thought tokens are re-read on every later step
- Sets one global verbosity policy for every task class
- Measures cost per run instead of per successful task
- Assumes more reasoning text always improves action quality