skip to content

Reason-Act-Observe Loop

The cycle itself — reason, act, observe, then reason again over what came back — and the conditions under which it stops. Interviewers ask what feeds the next thought and how you stop a loop that would otherwise never terminate.

on this pageshow

questions

5

In a ReAct agent, what happens in one pass of the reason-act-observe loop?

level: juniorimportance: must knowfreq 78%

answer

  1. three named parts, one repeated cycle
  2. one action per pass, then stop
  3. the environment supplies the third part
  4. transcript grows, next thought reads it
  5. a pass with no action ends the run

basics

~20 s

Each pass has three parts: the model writes a short reasoning step about what it needs next, emits one action against a tool, and gets the tool's result back as an observation. That observation joins the running context, and the next pass begins.

solid answer

~40 s

A ReAct pass is thought, then action, then observation. The model first writes a brief reasoning step naming what it still needs and why; it then emits a single action, meaning a tool call with arguments; the environment runs that call and hands back a result, which is appended to the running transcript as an observation. The next pass reasons over everything so far, including that new observation, so each cycle starts from strictly more information than the last. A shopping agent might reason "I need candidates under the budget", search, observe ten products with prices, then reason "three are under budget, open the cheapest" and act again. The loop keeps turning until the model produces an answer instead of an action, or the surrounding harness stops it.

go deeper

for a junior

Be able to name the three parts in order and say plainly that the observation comes back from the tool, not from the model. Walking one concrete four-step example is worth more than a definition.

for a middle

Explain that the transcript accumulates, so pass N is conditioned on every earlier thought, action and observation, and that generation must halt at the action boundary so the model cannot fabricate results.

for a senior

Show what you inspect when a run goes wrong: read the trace pass by pass and find the first thought whose stated reason does not follow from the observation above it. That is usually where the run went off.

for a principal

Own the argument for when this per-step feedback is worth its per-pass model call at all, versus a fixed pipeline for work whose steps never vary. Interleaving buys adaptivity, and you pay for it in latency and tokens on every task.

## The shape of the loop ReAct is the pattern of alternating **reasoning** with **acting**. Rather than thinking once and then executing a whole plan, or executing blindly with no thinking, the agent takes one small reasoning step, one action, and reads one result, over and over. The three parts of a pass are conventionally called *thought*, *action* and *observation*. The structural fact that makes the loop work is that it is a **single growing transcript**. Nothing is discarded between passes by default: pass three sees the thoughts, actions and observations of passes one and two. The model is not remembering in any special sense; the history is simply still in its context, and the next token it generates is conditioned on all of it. ## One pass, step by step **Thought.** The model writes a short natural-language step: what it knows, what is still missing, and what it will do about that. This is the model committing, in words, to a rationale before it commits to an action. **Action.** The model emits exactly one action — a named operation plus arguments — and then stops generating. Control leaves the model here. This is the crucial handoff: the model does not get to imagine the result, because the loop pauses and someone else produces it. **Observation.** The harness executes the action against the real environment (a search index, an HTTP API, a database, a shell) and appends whatever comes back to the transcript, labelled as an observation. Then the model is invoked again, and pass two begins from the enlarged transcript. ## Why this ordering matters Two things follow from the ordering. First, **every action is preceded by a stated reason**, so the trace explains itself: reading the transcript tells you not only what the agent did but what it believed at the time. That is what makes ReAct runs debuggable in a way that a bare sequence of tool calls is not. Second, **every reasoning step after the first is grounded in real returned data**. The agent is not extrapolating from its training data about what a search would probably return; it is reading what the search actually returned. This is why a ReAct agent can recover from a wrong guess: the observation contradicts the guess, and the next thought can say so. ## A worked example A shopping agent asked to buy a mechanical keyboard under 120 dollars might run: - Thought: I need keyboards under 120 dollars, so I should search the catalogue. - Action: search the catalogue for mechanical keyboards. - Observation: a list of twelve products with prices, several over the budget. - Thought: three are under 120 dollars; the cheapest may be out of stock, so open its product page. - Action: open that product page. - Observation: in stock, ships in two days, 94 dollars. - Thought: this satisfies the constraint; place the order. - Action: place the order. - Observation: order confirmed. - Thought: done. Final answer: ordered the 94-dollar model. Four passes, each one action wide. No pass tried to do two things at once, and no pass acted on an assumption that a later observation had to undo. ## Where a pass can end the loop A pass does not have to contain an action. When the model believes it has enough, it produces a final answer rather than an action, and the loop exits. That is the normal, in-band termination signal. Practical systems pair it with external limits — a cap on the number of passes, a cost or time budget — because a model that never emits an answer would otherwise spin forever. ## Common misreadings The most frequent misunderstanding is treating the loop as "plan everything, then execute". ReAct deliberately does not do that: the plan is only ever one step deep, re-derived after each observation. A second misreading is thinking the model produces the observation. It does not. If the model writes both an action and its result in one go, the run is hallucinating tool output, and the harness must be structured so that generation halts at the action boundary. A third is assuming a pass must contain a tool call; a pass whose thought concludes the task ends the run instead. ## What interviewers are checking They want to hear the three parts in order, the fact that the observation comes from the environment rather than the model, and that the transcript accumulates so each cycle reasons over more evidence than the last. Candidates who can walk a concrete four-step example are doing better than those who recite the three words.

  • Who produces the observation, and why does it matter that the model does not?
    The harness produces it by actually executing the action against the environment and appending the result. It matters because the observation is the only part of the transcript the model cannot invent. If generation is allowed to run past the action boundary, the model will happily write a plausible result itself, and the whole run becomes ungrounded fiction that still looks well formed.
  • Does every pass have to contain an action?
    No. A pass whose reasoning concludes the task produces a final answer instead of an action, and that is the loop's normal exit. A pass that emits an answer without ever having acted is the degenerate case: the agent has answered from memory rather than evidence, which is usually a defect on tasks that require live data.
  • Why one action per pass rather than several?
    Because the point of the pattern is that the next decision is conditioned on the last result. Batching actions gives up that feedback for the batched steps: they were all chosen from the same, older state. Loops do issue independent calls together for latency, but any step whose choice depends on a previous result has to wait for that observation.

It is navigating by looking up at each junction rather than memorising the whole route: decide the next turn, take it, see where you actually ended up, then decide again.

saying these in an interview costs you the question

  • Says the model writes the observation itself
  • Describes ReAct as planning all steps upfront
  • Thinks each pass starts from a fresh, empty context
  • Claims every pass must contain a tool call
  • Confuses the reasoning step with the final answer

context

open as a page

What ends a ReAct loop, and which stop conditions can the model itself decide?

level: middleimportance: must knowfreq 62%

basics

~20 s

The model ends the loop by producing a final answer instead of an action. Because that signal is only the model's own judgement, the surrounding harness adds external stops it cannot override: a cap on passes, a cost or time limit, and a goal check on the result.

open as a page

Why does a ReAct agent reason before each action instead of acting directly?

level: middleimportance: must knowfreq 70%

basics

~20 s

Writing a reason first forces the next action to be chosen from the current evidence rather than from habit. Without it the agent emits plausible-looking call sequences that ignore what the last result said, and nothing in the run catches the divergence.

open as a page

A ReAct agent repeats the same two searches until its step cap fires — why?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Because each pass re-derives the next action from a context that barely changed. When an observation adds nothing usable, the state that produced the last action is effectively reproduced, so the same action is chosen again, and the loop has no built-in notion of progress to notice it.

open as a page

When should an orchestrator drive the ReAct loop instead of letting the model continue it?

level: principalimportance: should knowfreq 34%

basics

~20 s

Take loop control into a state machine when the run must be auditable, resumable, or constrained — regulated or irreversible operations, long-running work that must survive a restart, and tasks whose legal step order is known. Leave continuation to the model when the path genuinely varies per task.

open as a page