skip to content

ReAct

The ReAct pattern: alternate a thought, an action against a tool, and the observation that comes back, until the agent has enough to answer. It is the reference loop behind most agent frameworks, so interviewers use it to test how a real agent turn is assembled.

on this pageshow

explore

questions

26

In a ReAct agent, what happens in one pass of the reason-act-observe loop?

level: juniorimportance: must knowfreq 78%

answer

  1. three named parts, one repeated cycle
  2. one action per pass, then stop
  3. the environment supplies the third part
  4. transcript grows, next thought reads it
  5. a pass with no action ends the run

basics

~20 s

Each pass has three parts: the model writes a short reasoning step about what it needs next, emits one action against a tool, and gets the tool's result back as an observation. That observation joins the running context, and the next pass begins.

solid answer

~40 s

A ReAct pass is thought, then action, then observation. The model first writes a brief reasoning step naming what it still needs and why; it then emits a single action, meaning a tool call with arguments; the environment runs that call and hands back a result, which is appended to the running transcript as an observation. The next pass reasons over everything so far, including that new observation, so each cycle starts from strictly more information than the last. A shopping agent might reason "I need candidates under the budget", search, observe ten products with prices, then reason "three are under budget, open the cheapest" and act again. The loop keeps turning until the model produces an answer instead of an action, or the surrounding harness stops it.

go deeper

for a junior

Be able to name the three parts in order and say plainly that the observation comes back from the tool, not from the model. Walking one concrete four-step example is worth more than a definition.

for a middle

Explain that the transcript accumulates, so pass N is conditioned on every earlier thought, action and observation, and that generation must halt at the action boundary so the model cannot fabricate results.

for a senior

Show what you inspect when a run goes wrong: read the trace pass by pass and find the first thought whose stated reason does not follow from the observation above it. That is usually where the run went off.

for a principal

Own the argument for when this per-step feedback is worth its per-pass model call at all, versus a fixed pipeline for work whose steps never vary. Interleaving buys adaptivity, and you pay for it in latency and tokens on every task.

## The shape of the loop ReAct is the pattern of alternating **reasoning** with **acting**. Rather than thinking once and then executing a whole plan, or executing blindly with no thinking, the agent takes one small reasoning step, one action, and reads one result, over and over. The three parts of a pass are conventionally called *thought*, *action* and *observation*. The structural fact that makes the loop work is that it is a **single growing transcript**. Nothing is discarded between passes by default: pass three sees the thoughts, actions and observations of passes one and two. The model is not remembering in any special sense; the history is simply still in its context, and the next token it generates is conditioned on all of it. ## One pass, step by step **Thought.** The model writes a short natural-language step: what it knows, what is still missing, and what it will do about that. This is the model committing, in words, to a rationale before it commits to an action. **Action.** The model emits exactly one action — a named operation plus arguments — and then stops generating. Control leaves the model here. This is the crucial handoff: the model does not get to imagine the result, because the loop pauses and someone else produces it. **Observation.** The harness executes the action against the real environment (a search index, an HTTP API, a database, a shell) and appends whatever comes back to the transcript, labelled as an observation. Then the model is invoked again, and pass two begins from the enlarged transcript. ## Why this ordering matters Two things follow from the ordering. First, **every action is preceded by a stated reason**, so the trace explains itself: reading the transcript tells you not only what the agent did but what it believed at the time. That is what makes ReAct runs debuggable in a way that a bare sequence of tool calls is not. Second, **every reasoning step after the first is grounded in real returned data**. The agent is not extrapolating from its training data about what a search would probably return; it is reading what the search actually returned. This is why a ReAct agent can recover from a wrong guess: the observation contradicts the guess, and the next thought can say so. ## A worked example A shopping agent asked to buy a mechanical keyboard under 120 dollars might run: - Thought: I need keyboards under 120 dollars, so I should search the catalogue. - Action: search the catalogue for mechanical keyboards. - Observation: a list of twelve products with prices, several over the budget. - Thought: three are under 120 dollars; the cheapest may be out of stock, so open its product page. - Action: open that product page. - Observation: in stock, ships in two days, 94 dollars. - Thought: this satisfies the constraint; place the order. - Action: place the order. - Observation: order confirmed. - Thought: done. Final answer: ordered the 94-dollar model. Four passes, each one action wide. No pass tried to do two things at once, and no pass acted on an assumption that a later observation had to undo. ## Where a pass can end the loop A pass does not have to contain an action. When the model believes it has enough, it produces a final answer rather than an action, and the loop exits. That is the normal, in-band termination signal. Practical systems pair it with external limits — a cap on the number of passes, a cost or time budget — because a model that never emits an answer would otherwise spin forever. ## Common misreadings The most frequent misunderstanding is treating the loop as "plan everything, then execute". ReAct deliberately does not do that: the plan is only ever one step deep, re-derived after each observation. A second misreading is thinking the model produces the observation. It does not. If the model writes both an action and its result in one go, the run is hallucinating tool output, and the harness must be structured so that generation halts at the action boundary. A third is assuming a pass must contain a tool call; a pass whose thought concludes the task ends the run instead. ## What interviewers are checking They want to hear the three parts in order, the fact that the observation comes from the environment rather than the model, and that the transcript accumulates so each cycle reasons over more evidence than the last. Candidates who can walk a concrete four-step example are doing better than those who recite the three words.

  • Who produces the observation, and why does it matter that the model does not?
    The harness produces it by actually executing the action against the environment and appending the result. It matters because the observation is the only part of the transcript the model cannot invent. If generation is allowed to run past the action boundary, the model will happily write a plausible result itself, and the whole run becomes ungrounded fiction that still looks well formed.
  • Does every pass have to contain an action?
    No. A pass whose reasoning concludes the task produces a final answer instead of an action, and that is the loop's normal exit. A pass that emits an answer without ever having acted is the degenerate case: the agent has answered from memory rather than evidence, which is usually a defect on tasks that require live data.
  • Why one action per pass rather than several?
    Because the point of the pattern is that the next decision is conditioned on the last result. Batching actions gives up that feedback for the batched steps: they were all chosen from the same, older state. Loops do issue independent calls together for latency, but any step whose choice depends on a previous result has to wait for that observation.

It is navigating by looking up at each junction rather than memorising the whole route: decide the next turn, take it, see where you actually ended up, then decide again.

saying these in an interview costs you the question

  • Says the model writes the observation itself
  • Describes ReAct as planning all steps upfront
  • Thinks each pass starts from a fresh, empty context
  • Claims every pass must contain a tool call
  • Confuses the reasoning step with the final answer

context

open as a page

In ReAct, what is an observation and who is allowed to write it?

level: juniorimportance: must knowfreq 58%

basics

~20 s

An observation is the tool's actual output, appended to the transcript by the runtime after an action. The model must never produce it: generation stops at the action, the real result is inserted, and only then does the model continue.

open as a page

In a ReAct loop, what does the runtime do with an action before the tool runs?

level: juniorimportance: must knowfreq 65%

basics

~20 s

The runtime resolves the action name against its registered tool catalogue, deserializes the arguments and validates them against that tool's schema, then executes. Anything that fails resolution or validation never reaches the tool — an error is returned instead.

open as a page

What does a Reflexion-style agent carry into its next attempt after a failed one?

level: middleimportance: must knowfreq 62%

basics

~20 s

A short self-written note in plain language about why the attempt failed and what to do differently. That note is appended to the next attempt's context, so the retry starts from an explicit lesson instead of repeating the same action.

open as a page

What ends a ReAct loop, and which stop conditions can the model itself decide?

level: middleimportance: must knowfreq 62%

basics

~20 s

The model ends the loop by producing a final answer instead of an action. Because that signal is only the model's own judgement, the surrounding harness adds external stops it cannot override: a cap on passes, a cost or time limit, and a goal check on the result.

open as a page

Why does a ReAct agent reason before each action instead of acting directly?

level: middleimportance: must knowfreq 70%

basics

~20 s

Writing a reason first forces the next action to be chosen from the current evidence rather than from habit. Without it the agent emits plausible-looking call sequences that ignore what the last result said, and nothing in the run catches the divergence.

open as a page

How do you truncate a 40k-token ReAct observation without losing the answer?

level: middleimportance: must knowfreq 62%

basics

~20 s

Reduce by relevance, not by position. Extract the fields or rows the current step actually needs, and leave an explicit marker saying how much was dropped, so the model treats the observation as partial rather than complete.

open as a page

In a text-protocol ReAct prompt, why set a stop sequence at "Observation:"?

level: middleimportance: must knowfreq 65%

basics

~20 s

A model that has seen Thought/Action/Observation exemplars will happily write the observation itself. Halting generation at the "Observation:" label hands control back to your runtime, which calls the real tool and appends the true result.

open as a page

In a ReAct agent, what does a well-formed thought step actually contain?

level: middleimportance: must knowfreq 66%

basics

~20 s

A well-formed ReAct thought states where the task stands against the goal, names the one piece of information still missing, and justifies the next action as the way to get it. Text that does not change which action follows is decoration.

open as a page

Why is regex-parsing a ReAct Action line brittle compared to structured tool calls?

level: middleimportance: must knowfreq 60%

basics

~20 s

Free-text action lines have no reliable grammar, so a quote or bracket inside an argument breaks the regex — often truncating the value silently rather than erroring. Structured tool calls constrain output to the declared schema, so arguments arrive already parsed and typed.

open as a page

What should a ReAct agent do when a tool returns an empty or null observation?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Render the empty case as explicit text, never as blank output. A literal such as NO_RESULT, plus the arguments used, tells the model the tool ran and matched nothing, which is a fact to reason about rather than a gap to fill from memory.

open as a page

How does a ReAct thought step turn into a hallucination the agent acts on?

level: seniorimportance: must knowfreq 58%

basics

~20 s

A thought is free text the model writes and then re-reads as if it were evidence. When it asserts something no observation supports — a policy limit, a field name, an id — later steps inherit that invention as established fact and act on it without ever checking.

open as a page

Why does Reflexion split the agent into actor, evaluator and self-reflection roles?

level: middleimportance: should knowfreq 44%

basics

~20 s

Three different jobs with different failure modes: the actor produces the attempt, the evaluator judges it and returns a verdict or score, and the self-reflection step turns that thin verdict into concrete written advice the actor can act on next time.

open as a page

How many Thought/Action/Observation exemplars should a ReAct prompt carry?

level: middleimportance: should knowfreq 50%

basics

~20 s

A handful — typically two to six complete trajectories — is enough to teach the block grammar and the handoff rhythm. Add more only when a measured failure demands it, because exemplars sit in the prefix and are re-sent on every step of the loop.

open as a page

How do you tell that a self-refine loop has stopped improving output and is just changing it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Score every round against a signal outside the critic rather than trusting the critic's own satisfaction, keep the best-scoring version instead of the last one, and stop when the diffs become paraphrase-level churn or the critique starts repeating and contradicting itself.

open as a page

A ReAct agent repeats the same two searches until its step cap fires — why?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Because each pass re-derives the next action from a context that barely changed. When an observation adds nothing usable, the state that produced the last action is effectively reproduced, so the same action is chosen again, and the loop has no built-in notion of progress to notice it.

open as a page

How do you keep a ReAct thought grounded in what the observation actually said?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Require the thought to quote a short span or cite the observation id before it draws a conclusion, and make sure your reduction preserved the line the conclusion depends on. Ungrounded thoughts drift into plausible paraphrase within two or three steps.

open as a page

When is a hand-written ReAct text protocol still better than native tool calling?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Rarely, as of mid-2026 — mainly with models that lack reliable structured tool calling, with action spaces that are not function-shaped, or when you need a provider-portable transcript you fully control. Otherwise the structured interface removes the exemplars, the stop string and the parser.

open as a page

What belongs in the system block of a helpdesk ReAct agent's prompt?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The stable, per-agent material: role and scope, the catalogue of available tools with when each applies, the required output format, escalation and refusal rules, and the operating limits. Per-request data — the customer's ticket, retrieved documents, tool results — belongs downstream, not here.

open as a page

In a long ReAct run, why should thoughts restate the goal and remaining steps?

level: seniorimportance: should knowfreq 44%

basics

~20 s

On long runs the original instruction sits far behind a wall of observations, and its influence on the next token weakens. Periodically re-writing the goal and what remains puts the objective back near the point of generation, which is the cheapest defence against drift.

open as a page

An on-call agent keeps calling search_logs when get_metrics would answer. How do you fix it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Overlapping tool definitions are usually the cause: the model reads names and descriptions as prompt text and cannot tell which surface owns the question. Rewrite each description to state what it is for and when to prefer the other, and merge genuine near-duplicates.

open as a page

When should a ReAct agent dispatch several tool calls in one action step?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Fan out only when the calls are independent reads whose arguments do not depend on each other's results. Keep writes sequential: parallel writes have no ordering guarantee, and a partial failure leaves inconsistent state with nothing to roll back.

open as a page

When does an agent's self-critique stop being evidence that its output is correct?

level: principalimportance: should knowfreq 33%

basics

~20 s

Self-critique stops being evidence when the critic shares the actor's blind spot or caves under pushback. A verdict that flips to "looks correct" after one objection is measuring agreeableness, not correctness, so calibrate the critic against seeded known-bad outputs before trusting it.

open as a page

When should an orchestrator drive the ReAct loop instead of letting the model continue it?

level: principalimportance: should knowfreq 34%

basics

~20 s

Take loop control into a state machine when the run must be auditable, resumable, or constrained — regulated or irreversible operations, long-running work that must survive a restart, and tasks whose legal step order is known. Leave continuation to the model when the path genuinely varies per task.

open as a page

When does fine-tuning on ReAct traces beat few-shot prompting the loop?

level: principalimportance: should knowfreq 32%

basics

~20 s

When the workflow is narrow, high-volume and stable, exemplars have grown long without fixing format or tool-selection errors, and thousands of verified trajectories exist. Fine-tuning moves the scaffolding into the weights, buying a shorter prompt at the price of a retrain whenever the tools change.

open as a page

How verbose should ReAct thought traces be when running agents at high volume?

level: principalimportance: should knowfreq 37%

basics

~20 s

There is no universal answer — set it per task class and prove it with an ablation. Thought tokens are billed on write and again on every later step that re-reads them, so verbosity compounds across a run while adding little to action quality on simple steps.

open as a page