skip to content

Planning and Reasoning

How an agent decides its next move — decomposing a goal into steps, interleaving reasoning with actions in a ReAct loop, and replanning when a step fails. Expect questions about where LLM planners break down: long horizons, irreversible actions, and plans that read as coherent but are not grounded in what the tools can actually do.

part ofAI agentsoverview, primer and where to startread it →
on this pageshow

explore

questions

14

In a ReAct agent loop, what do the thought, action and observation steps each do?

level: juniorimportance: must knowfreq 70%

answer

  1. reason, act, observe, repeat
  2. the loop looks before each leap
  3. tool output is fed back into the prompt
  4. thought is reasoning, never executed
  5. exit when it answers instead of acting

basics

~20 s

Thought is the model's reasoning about what to do next, action is the tool call it emits, and observation is the real tool result appended back into the prompt. The loop repeats until the model answers instead of acting.

solid answer

~50 s

ReAct interleaves reasoning with acting in one loop. On each turn the model writes a short **thought** — its rationale for the next move given everything it has seen so far — then emits an **action**, which is a tool call with arguments. The runtime executes that tool and appends the **observation**, the raw result, to the running transcript. The model reads the enlarged transcript and produces the next thought, and so on. The loop ends when the model returns a final answer instead of another action. The interleave is the whole point: because the next thought is conditioned on data that was actually observed rather than on a guess, a surprising result changes course immediately. On current frontier models the thought is often the model's native reasoning output rather than literal text you prompt for, but the control structure is identical.

go deeper

for a junior

Be able to name the three steps and say plainly that the observation is the tool's real result fed back in. Say that the loop ends when the model answers instead of calling another tool.

for a middle

Explain why reasoning after each observation is what makes the loop grounded, and that the transcript accumulates so later turns cost more input tokens than earlier ones.

for a senior

Show you have operated one: talk about oversized observations poisoning the window, thought text as your only trace of intent when a run misbehaves, and the serial round trip per step.

for a principal

Own the framing that ReAct trades tokens and latency for adaptivity, and that the trade only pays when the next step genuinely cannot be known before the previous result arrives.

## The loop in one sentence ReAct — short for *Reason + Act* — is a control loop in which a language model alternates between producing reasoning and calling tools, with every new round of reasoning conditioned on the actual results of the previous calls. It comes from the ReAct paper (Yao et al., 2022), which showed that interleaving a reasoning trace with actions beat both pure chain-of-thought reasoning, which can reason beautifully but cannot check anything against the world, and action-only agents, which call tools without articulating why and therefore recover badly from surprises. ## Thought The thought is the model's own reasoning for the current turn: what it has learned, what is still missing, and therefore what to do next. It is generated text (or native reasoning content), not something the runtime executes. Its job is to force the model to commit to a rationale before it commits to an action, which measurably reduces flailing — the model that has just written "the first lookup returned nothing for this ID, so the ID is probably in the legacy format" is far more likely to make a sensible next call than one that jumps straight to another tool. A second, underrated job of the thought is auditability. When a run goes wrong, the thought text is the only record of *why* the agent did what it did. Traces without it show you a sequence of calls with no explanation. ## Action The action is the structured tool invocation: a tool name plus arguments. This is the only part of the turn with side effects — it hits a search index, a database, a shell, an HTTP API. In modern systems it is emitted as a structured call rather than parsed out of free text, which removes a whole class of parsing failures the original paper had to deal with. The critical property is that exactly one decision point precedes it. The model chose this action knowing everything observed up to now and nothing about what comes after. ## Observation The observation is what the tool actually returned — search hits, a row set, an error message, a stack trace — placed back into the conversation as input for the next turn. It is ground truth, not model output, and that is precisely what makes ReAct grounded. If the tool returns an error or an empty result, the model sees the error and can react to it on the very next turn. Observations are also where the loop gets expensive. A single verbose tool result — a 40k-token HTML page, a full log file — permanently occupies the transcript for the rest of the run. ## Why the interleave matters Consider a multi-hop question: which of our three warehouses can fulfil an order today? The agent cannot know which warehouse to query second until it has seen the first one's stock level. Any strategy that fixes the whole call sequence in advance must guess. ReAct simply does not have to: each decision is made after the evidence for it arrives. The cost of that adaptivity is one model round trip per step. ## What a transcript looks like A typical run reads: thought → action → observation → thought → action → observation → … → final answer. Each turn resends everything before it, so the prompt grows monotonically and later turns are the expensive ones. Nothing in the loop is parallel by construction; the model decides, waits for the tool, then decides again. ## The modern wrinkle On frontier models, reasoning is adaptive and always available, so you rarely prompt for a literal "Thought:" prefix any more — the model's reasoning content plays that role and the runtime carries it between turns. Treat the three step names as a description of the control flow, not as a prompt template you must reproduce verbatim. ## What interviewers listen for They want you to say that the observation is real tool output, that the thought is reasoning and not executed, and that the loop's defining property is that reasoning happens *after* each observation rather than all at the front. Candidates who describe ReAct as "the model writes a plan and follows it" have described the opposite strategy.

  • What actually breaks if you strip the thought and let the model emit tool calls only?
    You lose two things. Accuracy drops on multi-hop tasks, because the model no longer commits to a rationale before choosing a call and tends to repeat or misorder lookups. And you lose the audit trail: a trace of bare calls tells you what happened but never why, which makes failures far harder to diagnose. Verbose thoughts cost tokens, so the practical answer is short rationales, not none.
  • How does the prompt grow across a long ReAct run?
    Monotonically. Every turn resends the accumulated thoughts, actions and observations, so turn ten pays for turns one through nine as input tokens. Large observations dominate — one oversized tool result stays in the window for the rest of the run. This is why long ReAct loops feel cheap per turn and expensive in aggregate.
  • Is ReAct the same thing as chain-of-thought prompting?
    No. Chain-of-thought produces reasoning only; nothing external is consulted, so the model can reason its way confidently to a wrong fact. ReAct alternates reasoning with real tool calls and feeds the results back, so its conclusions are grounded in observations. ReAct's reasoning steps are chain-of-thought-like, but the acting half is what makes it an agent loop.

saying these in an interview costs you the question

  • Says ReAct writes the full plan upfront and then executes it
  • Thinks the observation is the model's own prediction, not real tool output
  • Believes the thought text is executed as code by the runtime
  • Assumes each turn starts from a fresh prompt with no prior history
  • Confuses ReAct with plain chain-of-thought reasoning

context

open as a page

Why do agents write the plan to an external todo file instead of keeping it in context?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A written plan survives what the conversation does not. Context gets truncated, summarized or crowded out on long runs, so an external plan file keeps the goal and per-step status stable and re-readable, and doubles as an audit trail of what the agent actually did.

open as a page

How does plan-and-execute differ from ReAct in when the LLM makes decisions?

level: middleimportance: must knowfreq 78%

basics

~20 s

ReAct calls the model once per step and picks each action from the latest observation. Plan-and-execute calls the model once upfront to write the whole step list, then runs those steps with little or no further reasoning.

open as a page

How does an LLM agent detect that its plan has gone stale mid-execution?

level: middleimportance: must knowfreq 62%

basics

~20 s

Detection comes from checking each step against an expected outcome instead of assuming success. Three signals dominate: an explicit tool error, tool output that contradicts an assumption the plan was built on, and a postcondition check on the step's result that fails.

open as a page

How do you pick subtask granularity when an agent decomposes a goal?

level: middleimportance: must knowfreq 65%

basics

~20 s

Size each subtask so its completion can be checked objectively and it still fits one focused stretch of work. Too coarse and nobody can tell whether it succeeded; too fine and per-step overhead costs more than the work itself.

open as a page

When should an agent repair a single plan step instead of replanning fully?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Repair locally when the broken assumption belongs only to the failing step. Rewrite the remaining suffix when downstream steps depended on it. Replan wholly when the goal's feasibility or a global constraint changed. Scope follows the blast radius of the invalidated assumption, not the severity of the error.

open as a page

When should an agent backtrack to an earlier step rather than repair the failing one?

level: middleimportance: should knowfreq 38%

basics

~20 s

Backtrack when the failure's cause is an earlier step's output rather than the failing step itself — a stale reading, a wrong document, a bad intermediate result. Repairing at the point of failure only patches the symptom; the agent must invalidate the upstream result and re-derive everything that consumed it.

open as a page

A ReAct agent takes 12 turns at 8k context each; how does plan-and-execute change its cost and latency?

level: seniorimportance: should knowfreq 52%

basics

~20 s

It replaces twelve full-context reasoning calls with one planning call plus twelve cheap executions, cutting billed input tokens and letting independent steps run concurrently instead of in twelve serial round trips — but a wrong plan wastes the entire run.

open as a page

Why plan a freight-routing agent's legs upfront but run ReAct inside each leg?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Because predictability is uneven across the task. The nine-leg route sequence is stable, auditable and parallelizable, so plan it. Each leg's actual execution is messy and data-dependent, so give it a short bounded reasoning loop.

open as a page

How do you stop an agent from thrashing between replans of the same failed step?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Count replan attempts per failing subgoal, not just globally, and cap them at a small number such as three. Require each new plan to differ observably from the last and to state what the previous failure taught. On exhaustion, escalate with the failed plans and the deviation evidence attached.

open as a page

When does an agent plan need a dependency DAG instead of a linear todo list?

level: seniorimportance: should knowfreq 48%

basics

~20 s

A linear list is enough while every step genuinely depends on the one before it. Model an explicit dependency graph once steps are independent, once you want to run branches concurrently, or once a step's real prerequisites are not the item directly above it.

open as a page

How do you tell truly parallel subtasks from false parallelism in an agent plan?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Two subtasks are only independent if neither reads what the other writes and they touch no shared mutable resource. Absence of an obvious ordering is not evidence of independence — check data flow, shared write targets and shared bottlenecks before branching.

open as a page

How do you decide which tasks get a fixed workflow, plan-and-execute, or ReAct?

level: principalimportance: should knowfreq 35%

basics

~20 s

Match the strategy to how much of the step sequence you can know in advance. Fully known steps get hard-coded code, a knowable shape with variable content gets an upfront plan, and genuinely open-ended work gets a ReAct loop.

open as a page

How do you keep an agent's replanning anchored to the original goal over a long run?

level: principalimportance: should knowfreq 40%

basics

~20 s

Treat the goal and its hard constraints as an immutable artifact that is re-injected verbatim into every replanning cycle and never summarized away. Require each replan to restate the goal and justify how the new plan serves it, and check the final result against the original success criteria.

open as a page