skip to content

ReAct vs Plan-and-Execute

The two dominant agent control strategies: ReAct decides one step at a time from what it just observed, while plan-and-execute writes the whole plan first and then runs it. Being able to argue the cost, latency and reliability trade-offs between them is a standard senior-level question.

part ofAI agentsoverview, primer and where to startread it →
on this pageshow

questions

5

In a ReAct agent loop, what do the thought, action and observation steps each do?

level: juniorimportance: must knowfreq 70%

answer

  1. reason, act, observe, repeat
  2. the loop looks before each leap
  3. tool output is fed back into the prompt
  4. thought is reasoning, never executed
  5. exit when it answers instead of acting

basics

~20 s

Thought is the model's reasoning about what to do next, action is the tool call it emits, and observation is the real tool result appended back into the prompt. The loop repeats until the model answers instead of acting.

solid answer

~50 s

ReAct interleaves reasoning with acting in one loop. On each turn the model writes a short **thought** — its rationale for the next move given everything it has seen so far — then emits an **action**, which is a tool call with arguments. The runtime executes that tool and appends the **observation**, the raw result, to the running transcript. The model reads the enlarged transcript and produces the next thought, and so on. The loop ends when the model returns a final answer instead of another action. The interleave is the whole point: because the next thought is conditioned on data that was actually observed rather than on a guess, a surprising result changes course immediately. On current frontier models the thought is often the model's native reasoning output rather than literal text you prompt for, but the control structure is identical.

go deeper

for a junior

Be able to name the three steps and say plainly that the observation is the tool's real result fed back in. Say that the loop ends when the model answers instead of calling another tool.

for a middle

Explain why reasoning after each observation is what makes the loop grounded, and that the transcript accumulates so later turns cost more input tokens than earlier ones.

for a senior

Show you have operated one: talk about oversized observations poisoning the window, thought text as your only trace of intent when a run misbehaves, and the serial round trip per step.

for a principal

Own the framing that ReAct trades tokens and latency for adaptivity, and that the trade only pays when the next step genuinely cannot be known before the previous result arrives.

## The loop in one sentence ReAct — short for *Reason + Act* — is a control loop in which a language model alternates between producing reasoning and calling tools, with every new round of reasoning conditioned on the actual results of the previous calls. It comes from the ReAct paper (Yao et al., 2022), which showed that interleaving a reasoning trace with actions beat both pure chain-of-thought reasoning, which can reason beautifully but cannot check anything against the world, and action-only agents, which call tools without articulating why and therefore recover badly from surprises. ## Thought The thought is the model's own reasoning for the current turn: what it has learned, what is still missing, and therefore what to do next. It is generated text (or native reasoning content), not something the runtime executes. Its job is to force the model to commit to a rationale before it commits to an action, which measurably reduces flailing — the model that has just written "the first lookup returned nothing for this ID, so the ID is probably in the legacy format" is far more likely to make a sensible next call than one that jumps straight to another tool. A second, underrated job of the thought is auditability. When a run goes wrong, the thought text is the only record of *why* the agent did what it did. Traces without it show you a sequence of calls with no explanation. ## Action The action is the structured tool invocation: a tool name plus arguments. This is the only part of the turn with side effects — it hits a search index, a database, a shell, an HTTP API. In modern systems it is emitted as a structured call rather than parsed out of free text, which removes a whole class of parsing failures the original paper had to deal with. The critical property is that exactly one decision point precedes it. The model chose this action knowing everything observed up to now and nothing about what comes after. ## Observation The observation is what the tool actually returned — search hits, a row set, an error message, a stack trace — placed back into the conversation as input for the next turn. It is ground truth, not model output, and that is precisely what makes ReAct grounded. If the tool returns an error or an empty result, the model sees the error and can react to it on the very next turn. Observations are also where the loop gets expensive. A single verbose tool result — a 40k-token HTML page, a full log file — permanently occupies the transcript for the rest of the run. ## Why the interleave matters Consider a multi-hop question: which of our three warehouses can fulfil an order today? The agent cannot know which warehouse to query second until it has seen the first one's stock level. Any strategy that fixes the whole call sequence in advance must guess. ReAct simply does not have to: each decision is made after the evidence for it arrives. The cost of that adaptivity is one model round trip per step. ## What a transcript looks like A typical run reads: thought → action → observation → thought → action → observation → … → final answer. Each turn resends everything before it, so the prompt grows monotonically and later turns are the expensive ones. Nothing in the loop is parallel by construction; the model decides, waits for the tool, then decides again. ## The modern wrinkle On frontier models, reasoning is adaptive and always available, so you rarely prompt for a literal "Thought:" prefix any more — the model's reasoning content plays that role and the runtime carries it between turns. Treat the three step names as a description of the control flow, not as a prompt template you must reproduce verbatim. ## What interviewers listen for They want you to say that the observation is real tool output, that the thought is reasoning and not executed, and that the loop's defining property is that reasoning happens *after* each observation rather than all at the front. Candidates who describe ReAct as "the model writes a plan and follows it" have described the opposite strategy.

  • What actually breaks if you strip the thought and let the model emit tool calls only?
    You lose two things. Accuracy drops on multi-hop tasks, because the model no longer commits to a rationale before choosing a call and tends to repeat or misorder lookups. And you lose the audit trail: a trace of bare calls tells you what happened but never why, which makes failures far harder to diagnose. Verbose thoughts cost tokens, so the practical answer is short rationales, not none.
  • How does the prompt grow across a long ReAct run?
    Monotonically. Every turn resends the accumulated thoughts, actions and observations, so turn ten pays for turns one through nine as input tokens. Large observations dominate — one oversized tool result stays in the window for the rest of the run. This is why long ReAct loops feel cheap per turn and expensive in aggregate.
  • Is ReAct the same thing as chain-of-thought prompting?
    No. Chain-of-thought produces reasoning only; nothing external is consulted, so the model can reason its way confidently to a wrong fact. ReAct alternates reasoning with real tool calls and feeds the results back, so its conclusions are grounded in observations. ReAct's reasoning steps are chain-of-thought-like, but the acting half is what makes it an agent loop.

saying these in an interview costs you the question

  • Says ReAct writes the full plan upfront and then executes it
  • Thinks the observation is the model's own prediction, not real tool output
  • Believes the thought text is executed as code by the runtime
  • Assumes each turn starts from a fresh prompt with no prior history
  • Confuses ReAct with plain chain-of-thought reasoning

context

open as a page

How does plan-and-execute differ from ReAct in when the LLM makes decisions?

level: middleimportance: must knowfreq 78%

basics

~20 s

ReAct calls the model once per step and picks each action from the latest observation. Plan-and-execute calls the model once upfront to write the whole step list, then runs those steps with little or no further reasoning.

open as a page

A ReAct agent takes 12 turns at 8k context each; how does plan-and-execute change its cost and latency?

level: seniorimportance: should knowfreq 52%

basics

~20 s

It replaces twelve full-context reasoning calls with one planning call plus twelve cheap executions, cutting billed input tokens and letting independent steps run concurrently instead of in twelve serial round trips — but a wrong plan wastes the entire run.

open as a page

Why plan a freight-routing agent's legs upfront but run ReAct inside each leg?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Because predictability is uneven across the task. The nine-leg route sequence is stable, auditable and parallelizable, so plan it. Each leg's actual execution is messy and data-dependent, so give it a short bounded reasoning loop.

open as a page

How do you decide which tasks get a fixed workflow, plan-and-execute, or ReAct?

level: principalimportance: should knowfreq 35%

basics

~20 s

Match the strategy to how much of the step sequence you can know in advance. Fully known steps get hard-coded code, a knowable shape with variable content gets an upfront plan, and genuinely open-ended work gets a ReAct loop.

open as a page