How does plan-and-execute differ from ReAct in when the LLM makes decisions?
answer
- decide once, or decide every step
- the model is scheduler vs compiler
- a plan is an artifact you can review
- upfront plans cannot see tool output
- adaptivity bought with tokens and latency
basics
~20 sReAct calls the model once per step and picks each action from the latest observation. Plan-and-execute calls the model once upfront to write the whole step list, then runs those steps with little or no further reasoning.
solid answer
~50 sThe difference is where the decision points sit. In **ReAct** every step is a fresh decision: the model reasons, acts, sees the observation, and only then decides step two. Nothing is committed in advance, so the agent adapts to whatever it finds, at the price of one model round trip per step and a transcript that grows all run. In **plan-and-execute** the model makes one big decision at the front — a written, ordered list of steps with their tools and arguments — and an executor then runs it. Executions may need no reasoning model at all, or a small cheap one. That buys determinism (the same plan runs the same way), auditability (a human or a policy check can read the plan before anything executes), parallelism for independent steps, and far fewer expensive calls. What it costs is grounding: the plan is written before any tool has returned anything, so it can only be as good as what the model could know upfront.
go deeper
Be able to state the core contrast: ReAct decides one step at a time from the last result, plan-and-execute writes all the steps first and then runs them.
Explain what moving the decision point buys and costs — determinism, cost and parallelism on one side, grounding and adaptivity on the other — and how a plan expresses a step-to-step dependency.
Argue the choice from a concrete task: how much of the sequence is knowable upfront, how often plans go stale in that domain, and what a wrong plan costs in side effects before anyone notices.
Own the position that this is a per-task-class decision rather than a product-wide architecture, and that the hybrid exists because predictability is usually uneven across a workflow.
## Two shapes of the same loop Every tool-using agent has to answer one question repeatedly: what next? The two dominant control strategies differ only in *when* that question is put to the model. **ReAct** asks it every step. Reason, act, observe, reason again — the decision for step *n+1* is made after step *n*'s result is in hand. **Plan-and-execute** asks it once, before anything runs, and gets back a complete step list; a much dumber executor then walks the list. Variants of the second shape appear under several names in the literature — Plan-and-Solve prompting, ReWOO (Reasoning WithOut Observation), LLMCompiler — and they differ in details, but all share the property of deciding the sequence before observing anything. ## Where the model sits In ReAct the model is *in* the control loop; it is the scheduler. In plan-and-execute the model is a compiler that emits a program, and something deterministic runs the program. That single relocation is where every downstream difference comes from. ## What an upfront plan buys **Determinism and reviewability.** A plan is an artifact. You can log it, diff two runs, show it to a person or a policy engine before a single side effect fires, and reject it if it names a tool the caller is not allowed to use. A ReAct agent's intentions only exist one turn at a time. **Cost.** One planning call on a strong model, then executions that need little or no reasoning. You have replaced N expensive calls over a growing transcript with one expensive call plus N cheap ones. **Parallelism.** A written plan makes dependencies explicit, so independent steps can run concurrently. ReAct is serial by construction: it cannot start step three before step two's observation exists, because step three has not been chosen yet. **Bounded shape.** You know before execution how many steps there will be and what they touch. That makes cost, latency and blast radius predictable, which matters enormously for anything customer-facing. ## What an upfront plan assumes A plan written before any observation encodes the model's *beliefs* about the world. Three assumptions are doing quiet work: 1. **The tools behave as the model imagines.** A planner that has only tool descriptions can happily emit a step no tool can actually perform, or arguments in a shape the API rejects. 2. **Values the plan needs are knowable, or can be deferred.** Step four often needs something only step two can return. Good plan formats handle this with variable substitution — step four references step two's output as a placeholder that the executor fills in — which is exactly ReWOO's contribution. A planner that instead *invents* the value has hallucinated a fact into your control flow. 3. **The world does not change under the plan.** The longer the horizon, the weaker this holds. A plan for nine sequential steps in a system with real failure rates is a plan that will be stale before it finishes. When a step does turn out impossible, you need machinery to notice and revise — which is a subject in its own right, and a real cost of choosing this strategy rather than a free property of it. ## Where ReAct wins outright Any task where the *shape* of the work is unknowable upfront. Debugging a failing service: you cannot enumerate the queries before seeing the first log line. Exploratory research: what you read second depends on what the first source said. In these, an upfront plan is a guess dressed as structure, and its apparent determinism is misleading — it is deterministic about the wrong steps. ## Where plan-and-execute wins outright Tasks whose *shape* is stable while the *content* varies. A monthly reconciliation across four systems is the same seven steps every time with different identifiers; planning it once and executing is faster, cheaper, reviewable, and parallelizable. If you find yourself running ReAct on such a task, you are paying twelve reasoning calls to re-derive a sequence you already know. ## The honest summary ReAct buys adaptivity with tokens, latency and nondeterminism. Plan-and-execute buys determinism, cost and parallelism with the risk that its assumptions were wrong. Neither is a better architecture in the abstract; the question is how much of the step sequence you can honestly know before the first tool returns. That is also why the hybrid — plan the coarse steps, run a small ReAct loop inside each — is so common in production: it puts the deterministic strategy where the work is predictable and the adaptive one where it is not.
- A plan's step four needs a value only step two can return. How do plan formats handle that?With variable substitution: step four references step two's output as a placeholder, and the executor fills it in at runtime. That keeps the plan writable before anything runs while still expressing the dependency. The failure mode to watch for is a planner that instead invents a plausible-looking value and hardcodes it — a hallucination baked into control flow, which executes silently and produces confidently wrong results.
- Which strategy parallelizes better, and why?Plan-and-execute. Because the whole step list exists before execution, independent steps are visible and can run concurrently — that is the core idea behind compiler-style planners. ReAct is serial by construction: step three has not even been chosen until step two's observation arrives, so there is nothing to parallelize across turns.
- What does a plan-and-execute agent silently assume about your tools?That they can do what the planner imagined and accept the arguments it invented. A planner sees descriptions, not behaviour, so it can emit steps no tool supports or argument shapes the API rejects — and the mismatch only surfaces at execution. Tight, honest tool descriptions and validating the plan against real schemas before running it are the standard mitigations.
- Does plan-and-execute mean no model calls during execution?Not necessarily. Some executors are purely mechanical; many call a small, cheap model per step to fill arguments or summarize a result, and hybrids run a full short reasoning loop inside each step. The defining property is not the absence of model calls during execution — it is that the step sequence was committed before the first tool ran.
saying these in an interview costs you the question
- Says ReAct has no plan at all rather than an implicit one re-derived each turn
- Claims plan-and-execute is always cheaper, ignoring runs where the plan is wrong
- Assumes an upfront plan is grounded in tool behaviour it has never observed
- Treats the choice as a framework feature instead of a control-flow decision
- Thinks plan-and-execute forbids any model call during execution