skip to content

A ReAct agent takes 12 turns at 8k context each; how does plan-and-execute change its cost and latency?

level: seniorimportance: should knowfreq 52%

answer

  1. the transcript ramps, it does not repeat flat
  2. one expensive call versus twelve
  3. serial round trips set the wall clock
  4. independent steps collapse to the critical path
  5. a stale plan wastes everything after it

basics

~20 s

It replaces twelve full-context reasoning calls with one planning call plus twelve cheap executions, cutting billed input tokens and letting independent steps run concurrently instead of in twelve serial round trips — but a wrong plan wastes the entire run.

solid answer

~50 s

Work the arithmetic. Twelve ReAct turns averaging 8k context bill roughly 96k input tokens, and because the transcript accumulates, the later turns carry the bulk of that — turn twelve pays for everything before it. Latency is twelve serial model round trips plus twelve tool calls, because step *n+1* is not even chosen until step *n* returns. Plan-and-execute restructures this into one planning call of maybe 3k on a strong model, then twelve executions that need a small model or none, each seeing only its own slice rather than the whole history. Independent steps run concurrently, so wall-clock collapses toward the critical path instead of the sum. Two honest caveats. Prompt caching narrows the gap: the shared ReAct prefix is discounted on repeat, so the naive token count overstates the real bill. And the saving is conditional — if the plan is wrong at step three, you paid for planning *and* for executions that produced nothing useful. This reflects practice as of mid-2026.

go deeper

for a junior

Know that a ReAct run resends its whole history each turn, so twelve turns cost far more than twelve times the first turn, and that each turn is one more round trip.

for a middle

Do the arithmetic out loud: the ramping context, one expensive planning call versus twelve, and why independent steps can only be parallelized when the plan exists upfront.

for a senior

Show operational judgment — cache hit rates, observation bloat, plan completion rate, and the cheaper levers (trimming results, smaller models on mechanical turns) you would try before rewriting the control loop.

for a principal

Frame it as expected cost under uncertainty, not sticker price: planning cost divided by plan survival rate, plus the blast radius of an executor that keeps firing side effects on stale premises.

## Why the naive token count is wrong in both directions The first thing to get right in this answer is that ReAct's cost is not twelve times one turn. Each turn resends the accumulated transcript — every prior thought, action and observation — so input tokens grow with turn index. A run described as "12 turns at 8k average context" is really a ramp: turn one might be 2k, turn twelve 15k. That shape matters, because it means long ReAct runs get *more* expensive per unit of progress the longer they go, and because a single oversized observation early in the run is paid for on every subsequent turn. It is also wrong in the other direction if you ignore caching. Providers discount tokens that repeat a previously-seen prefix, and a ReAct transcript is almost pure prefix repetition — turn twelve's prompt is turn eleven's prompt plus a bit. Where prompt caching applies, the marginal cost of an extra ReAct turn is much closer to the new tokens than to the full context. A candidate who quotes 96k billed tokens with no caveat has not run one of these in production. ## The plan-and-execute cost shape One planning call reads the task, the tool catalogue and whatever context is needed to decide the sequence — call it 3k in, some hundreds out, on your most capable model. Then twelve executions, each of which needs only its own step's instructions plus whatever inputs it depends on. Two consequences: - **The expensive model is called once.** Execution can run on a small model, or on no model at all where the step is a mechanical tool call with substituted arguments. - **Context does not accumulate.** Each execution's prompt is bounded by its own step, not by the run's history. This is the difference between a linear-ish total and a quadratic-ish one. ## Latency ReAct's wall clock is inherently serial: `12 × (model round trip + tool round trip)`. There is no way around it, because the decision for step *n+1* does not exist until step *n*'s observation exists. Faster models shorten each term; they do not change the shape. Plan-and-execute's wall clock is `one planning round trip + the critical path through the plan`. If eight of the twelve steps are mutually independent lookups, they finish in roughly the time of the slowest one. On fan-out-heavy work this is the single biggest latency lever available, and it is unavailable to ReAct by construction. The analogy interviewers respond to is deep-space robotics: you send a rover a sequence of moves rather than one instruction at a time, precisely because each round trip costs minutes of light delay. When the round trip is expensive relative to the work, committing to a plan wins. ## What the plan costs you Three things, and a strong answer names all three. **Wasted work on stale plans.** If reality diverges at step three, everything after it may be worthless, and you have also paid the planning call. The expected cost of plan-and-execute is not the plan cost — it is the plan cost divided by the probability that the plan survives to completion. On a domain where plans hold 90% of the time, that is a small tax; at 50% it can erase the entire saving. **Fixed overhead on trivial tasks.** Every request pays for planning, including the ones that would have taken ReAct two turns. **Larger blast radius when wrong.** ReAct notices a bad first result immediately and stops. An executor mid-plan may keep firing side-effecting steps against premises that no longer hold, which is a correctness and safety cost, not just a token one. ## How to actually decide Measure rather than argue. The numbers you want are: distribution of turn counts per task, distribution of context size at each turn (to see how badly observations bloat), cache hit rate on the ReAct prefix, and — the one people skip — the rate at which a generated plan completes without needing revision. That last number determines whether the arithmetic favouring plan-and-execute is real or theoretical. Also check the cheap levers before restructuring control flow. Trimming oversized tool results, letting the agent hold references and fetch on demand instead of pasting whole documents, and using a smaller model for the mechanical turns often recover most of the cost gap while keeping ReAct's adaptivity. Restructuring the control strategy is the bigger hammer; reach for it when the task shape genuinely supports an upfront plan, not merely because the token bill is high.

  • Where does prompt caching change this arithmetic?
    Substantially in ReAct's favour. Each turn's prompt is the previous turn's prompt plus a delta, so the repeated prefix is discounted and the marginal cost of an extra turn approaches the new tokens rather than the full context. That does not help latency at all — the round trips are still serial — and it evaporates whenever something rewrites earlier context and invalidates the prefix.
  • Which strategy has the worse tail latency, and why?
    ReAct, on fan-out-shaped work. Its wall clock is the sum of every round trip because no step can start before the previous observation exists, so p99 is driven by turn count times the slowest link. Plan-and-execute converts independent steps into concurrent ones, so its tail tracks the critical path. On strictly sequential tasks the two converge.
  • What single number decides whether the plan-and-execute saving is real?
    The share of generated plans that run to completion without needing revision. The saving is roughly the token gap multiplied by that rate, minus the wasted executions on the rest. At 90% completion the strategy is clearly cheaper; near 50% you are paying for planning and for discarded work, and ReAct's step-by-step grounding is likely the better buy.

Committing to a plan makes sense when round trips are expensive — you send a Mars rover a sequence of moves rather than one instruction at a time, because waiting on each reply costs minutes of light delay.

saying these in an interview costs you the question

  • Counts only output tokens and forgets every turn resends the transcript
  • Assumes plan-and-execute is cheaper regardless of how often plans go stale
  • Ignores that ReAct's serial round trips dominate wall-clock latency
  • Quotes raw token totals with no mention of prompt caching
  • Restructures control flow before trimming oversized tool results

context