skip to content

Why does a ReAct agent reason before each action instead of acting directly?

level: middleimportance: must knowfreq 70%

answer

  1. decoration or load-bearing?
  2. the channel evidence uses to reach the action
  3. without it, priors choose the call
  4. fluent, well-formed, wrong sequences
  5. one sensor is not two sensors

basics

~20 s

Writing a reason first forces the next action to be chosen from the current evidence rather than from habit. Without it the agent emits plausible-looking call sequences that ignore what the last result said, and nothing in the run catches the divergence.

solid answer

~50 s

The thought step is where the agent states what the last observation actually established and what is therefore still missing. That statement conditions the tokens that follow, so the action is selected against present evidence instead of a generic pattern the model has seen before. Consider a wildfire-alert agent: one sensor reports a temperature spike, and the reasoning step says "a single sensor spike could be a faulty probe, so check the neighbouring sensor before escalating" — which produces a second query rather than an immediate page-out. Strip the reasoning and the agent still emits confident, well-formed calls, but they stop tracking the evidence: it repeats steps, escalates on one noisy reading, or misses that the previous call already failed. The reasoning step also makes the run auditable, since every action carries the belief that produced it.

go deeper

for a junior

Say clearly that the agent states why it is about to do something before it does it, and that this is what lets the previous result change the next step.

for a middle

Explain the mechanism: the thought is generated first and the action tokens are conditioned on it, so without a thought the model falls back on the generic shape of such call sequences and stops tracking the evidence.

for a senior

Bring a failure you have seen — an agent that kept calling the next obvious step after an empty result — and say what in the trace told you the reasoning had stopped doing work.

for a principal

Own the cost side: reasoning on every pass is a per-request tax across the whole workload. Be ready to argue where you would spend it, where a fixed sequence is cheaper and safer, and how you would measure the difference rather than assert it.

## The claim being tested The interviewer is probing whether you understand that the reasoning step is *load-bearing* rather than decoration. It is easy to assume the thought is commentary for humans and the action is the real work. In fact the thought is the mechanism by which the observation influences the action. ## Why it works mechanically An agent's next action is generated conditioned on everything already in the transcript. When a thought sits between the last observation and the next action, that thought is written first, and the action tokens are then generated conditioned on it. The thought is effectively a compression step: it takes a possibly messy observation and states, in the agent's own words, what it means for the goal. The action is then chosen against that statement. Remove the thought and the action is generated directly after the observation. The model still sees the observation, but nothing forces it to commit to an interpretation, so it falls back on the strongest available prior — the generic shape of what a call sequence for this kind of task usually looks like. That is exactly the regime in which agents emit fluent, well-formed, wrong sequences. ## What degrades without it **The last result stops steering.** The most visible symptom is that the agent behaves the same whether the previous call succeeded, returned nothing, or errored. It proceeds to the step that "comes next" for tasks of that shape. **Wrong premises survive.** With a reasoning step the agent can write "the search returned nothing for this part number, so my assumption about the format is wrong". Without one there is no place in the trace for that correction to be made, and the incorrect assumption is carried into every later call. **Constraints get dropped.** Multi-constraint goals ("under budget", "only sensors in the same district", "read-only") are re-asserted in the thought. Actions that were never mediated by a thought tend to satisfy the salient constraint and quietly ignore the rest. **The trace becomes uninterpretable.** A sequence of bare calls tells you what happened but not why, so when a run goes wrong you cannot tell whether the agent chose badly or was misled by a bad result. ## The wildfire example, both ways A wildfire-alert agent reads a sensor showing 71 degrees where 40 is normal. With reasoning: "A single sensor reading 71 is far above baseline, but a failing probe produces exactly this signature. Check the nearest neighbouring sensor before escalating." The action is a second sensor read. The observation shows the neighbour also elevated, and only then does the agent page the duty crew — a decision now supported by two independent readings. Without reasoning: the observation is a high number, the highest-prior next action for an alert agent is to raise an alert, and the crew is dispatched on one possibly faulty probe. Nothing in that run was malformed. It was simply never made to ask whether one reading is enough. ## The honest limits Reasoning first is not a correctness guarantee. A confidently wrong thought produces a confidently wrong action, and a reasoning step can rationalise a bad choice as readily as it can catch one. Reasoning also costs tokens and latency on every single pass, which is a real budget item on long runs. And on genuinely mechanical steps — the third page of an obvious pagination loop — the thought adds little beyond cost. The defensible claim is narrower and stronger than "thinking makes it right": interleaving gives the run a place where the last observation can change the next decision, and a record of whether it did. Reason-only agents, which think without acting, drift because nothing external ever contradicts them; act-only agents, which call tools without reasoning, cannot notice that a result contradicted the plan. Alternating the two was the point of the pattern. ## What to say in an interview Lead with the mechanism (the thought conditions the action, so the observation reaches the decision), give one concrete before-and-after like the sensor case, then name the cost — tokens and latency per pass — and the limit, that a wrong thought still yields a wrong action. Candidates who present it as a guarantee are marked down; candidates who present it as the channel through which feedback reaches the next decision are not.

  • If reasoning is that useful, why not make every thought long and exhaustive?
    Because every thought is generated and then re-read on every later pass, so verbosity compounds in both cost and context pressure. Long thoughts also give the model more room to rationalise a choice it had already drifted toward. The useful thought is short and decision-shaped: what the last result established, what is missing, what to do next.
  • Where does interleaved reasoning fail to help at all?
    On confidently wrong premises. If the thought misreads the observation, the action follows the misreading, and the next thought reasons from that. Reasoning propagates errors as readily as it catches them, which is why production loops add external signals — verifiers, checks, or a human — rather than trusting the trace to police itself.
  • How would you show that the reasoning step is actually earning its cost?
    Run the same task suite with the reasoning step and with actions emitted directly, and compare task success, wrong-tool rate and step count, not just vibes from a few traces. On mechanical, low-branching tasks the gap is often small and the cheaper loop wins; on tasks where results genuinely change the next step, the gap is where the value is.

saying these in an interview costs you the question

  • Calls the reasoning step just commentary for humans
  • Claims reasoning first prevents hallucination
  • Says it makes the agent's behaviour deterministic
  • Ignores the token and latency cost of thinking every pass
  • Assumes a correct-looking thought means a correct action

context