skip to content

How does a ReAct thought step turn into a hallucination the agent acts on?

level: seniorimportance: must knowfreq 58%

answer

  1. the model re-reads its own prose as evidence
  2. observations and thoughts look identical in context
  3. a fabricated premise steers later actions
  4. hedge it or verify it
  5. walk each claim back to an observation

basics

~20 s

A thought is free text the model writes and then re-reads as if it were evidence. When it asserts something no observation supports — a policy limit, a field name, an id — later steps inherit that invention as established fact and act on it without ever checking.

solid answer

~50 s

The mechanism is self-citation. Reasoning text lands in the same transcript as real tool observations, and on the next step the model cannot tell its own prose from grounded data — everything is just prior context. So a compliance agent whose thought says "the policy allows 30 days for a response" before any policy lookup has run will happily plan, compute and answer on that number, and the fabrication never meets a tool that could contradict it. The defences are all about provenance. Require each factual claim in a thought to point at the observation it came from; make unverified beliefs explicitly hedged so a later step treats them as open; and bias the agent toward spending one cheap verifying action instead of asserting. Keeping thoughts short helps too, since most invented facts appear in the padding rather than in the sentence that picks the action.

go deeper

for a junior

Know that a thought is text the model wrote, not evidence, and that stating an unchecked fact there can send the agent down a wrong path.

for a middle

Explain why observations and thoughts are indistinguishable in the transcript, and give the concrete fix: attribute claims to observations, hedge anything unverified.

for a senior

Show that you would diagnose this in production — provenance-walk load-bearing claims back to observations, watch for mandatory lookups that never fired, and gate real side effects on grounded values.

for a principal

Own the policy: which task classes require provenance for every action-driving fact, what the verification tax costs against the risk of a confidently wrong side effect, and how that discipline is enforced and audited across teams.

## The mechanism: the model cites itself A ReAct transcript interleaves two very different kinds of text: observations, which came from the outside world through a tool, and thoughts, which the model generated. To the model on the next step there is no such distinction — both are tokens in the prior context, and both are treated as things that are already known. That is the whole failure. A fabricated claim written in step 2 is, by step 5, indistinguishable from a fact that was actually retrieved, and every subsequent step conditions on it. This is worse than a one-shot hallucination in a single answer. In a single answer a wrong fact produces one wrong sentence. In an agent loop a wrong fact steers the *action* selection: it decides which tool is called, with which arguments, and when the agent believes it is done. The error compounds through the trajectory rather than sitting still. ## The canonical shape Consider a compliance agent asked whether a vendor response is late. Somewhere in step 2 the thought reads: "The policy allows 30 days for a response, and the vendor replied on day 34, so this is a breach." No policy lookup has been performed. The number is a plausible prior — 30 days is a common contractual window — and it is stated flatly, with no hedge. From that point on: - The agent does not call the policy tool, because from its own perspective it already knows the answer. - It calls the escalation tool instead, which is a real side effect on a real system. - The final answer is confidently wrong *and internally consistent*, which makes it hard to spot in review. Notice that the tool layer never failed. Nothing threw an error, nothing timed out, no schema was violated. This is a reasoning-layer failure that the tool layer cannot catch, because the missing call is the failure. ## Why verbosity raises the odds Every additional sentence of reasoning is another opportunity to state something the model does not know. Terse thoughts that name a gap and pick an action have almost no surface area for invention; a 200-word essay about the domain has a great deal. Long thoughts also tend to *pre-commit* to a conclusion, and once a conclusion is on the page the remaining steps drift toward rationalising it rather than testing it. Empirically, the invented facts in agent traces cluster in the padding — the background narration, the confident summary, the "as we know" clauses — not in the terse sentence that selects the tool. ## Defences that actually work **Provenance discipline.** Instruct the agent that any factual claim in a thought must be attributable to a specific observation, and that unattributable claims must be written as hypotheses: "the policy window is likely 30 days — not yet verified." A hedged claim is safe because later steps can see it is open. The same claim stated flatly is a trap. **Verify-before-assert.** When a thought names its own uncertainty, the right next action is usually the cheap check rather than the expensive commitment. "I believe the window is 30 days but have not read the policy; the policy lookup costs one call, so read it" is exactly the behaviour to reward. Making the cheap verifying action available and obviously cheap is a design decision, not a prompting flourish. **Gate side effects on grounded facts.** Any action with a real consequence — escalation, refund, ticket creation, write to a system of record — should depend on values that trace to an observation. If the agent cannot point at where the number came from, the action does not fire. **Keep thoughts short.** Fewer sentences, fewer chances to invent. This is the cheapest lever and it costs nothing at inference time. ## How you detect it after the fact Reviewing traces for this failure is mostly provenance auditing: take each load-bearing factual claim in the final answer and walk backwards until you find the observation that produced it. If the walk terminates in a thought rather than an observation, that fact was invented. This is mechanisable — extract claims, match them against observed text, flag the unmatched ones — and it is a far better signal than judging the final answer alone, because a fabricated premise can still produce an answer that looks reasonable. A second signal is the *absent call*: for task classes where a specific lookup is always required, check whether it appears in the trajectory at all. A missing mandatory call frequently means the model decided it already knew the answer. ## What this is not This is not about whether the stated reasoning faithfully reflects the model's internal computation — that is a separate, measured phenomenon. Here the concern is narrower and more practical: the trace contains assertions with no external support, and the agent's own later steps consume them as if they had support.

  • How would you audit a batch of production traces for this specific failure?
    Provenance-walk the load-bearing claims. For each factual assertion the final answer depends on, trace backwards until you hit the observation that produced it; if the chain terminates in a thought, the fact was invented. Complement that with an absent-call check: for task classes where a lookup is mandatory, flag trajectories where it never appears, since that usually means the agent decided it already knew.
  • Doesn't a tool error eventually expose the fabrication?
    Often not, because the characteristic failure is a call that never happens. Nothing errors when the agent skips the policy lookup — it simply proceeds on an invented number and takes a downstream action that succeeds mechanically. Tool-layer health tells you nothing about reasoning-layer grounding, which is why side effects should be gated on values traceable to an observation.
  • Does forcing the agent to hedge every claim hurt performance?
    Blanket hedging degrades decisiveness and adds tokens, so target it: require attribution only for claims that will drive an action or appear in the answer. Domain framing and restatements of the request need no hedge. The rule to enforce is narrow — an unattributable number, identifier, threshold or policy statement must be marked unverified before anything acts on it.

saying these in an interview costs you the question

  • Assumes the model can distinguish its own reasoning from tool output
  • Thinks hallucination only affects the final answer, not action choice
  • Believes longer reasoning reduces fabrication
  • Relies on tool errors to catch invented facts
  • Never checks whether a mandatory lookup actually happened

context