In ReAct, what is an observation and who is allowed to write it?
answer
- who typed this line?
- environment text, not model text
- the runtime appends it after the action
- label the source, the arguments, an id
- self-written observations are fabricated evidence
basics
~20 sAn observation is the tool's actual output, appended to the transcript by the runtime after an action. The model must never produce it: generation stops at the action, the real result is inserted, and only then does the model continue.
solid answer
~50 sIn a ReAct transcript the model writes thoughts and actions; the **runtime** writes observations. After the model emits an action, generation is halted at that boundary, the runtime executes the tool, and the tool's real output — reduced if necessary — is appended as the observation before the model resumes. This separation is the whole point: an observation is the only part of the transcript that is evidence rather than the model's own text. If the model is allowed to continue past the action, it will happily write a fluent, entirely fabricated observation and then reason on top of it, and nothing downstream can tell the difference. In practice each observation is labelled with its source — the tool name, a summary of the arguments used, and a step id — so a later thought can cite 'per observation 3' instead of restating the content, and so a human reading the trace can tell which lines came from the world and which from the model.
code
python · 7 linesdef append_observation(transcript: list[str], step: int, tool: str, args: dict, result: str) -> None:
transcript.append(f"Observation {step} ({tool}, args={args}):\n{result}")
transcript: list[str] = ["Thought 3: I need the on-call engineer.", "Action 3: roster_lookup(service='payments')"]
append_observation(transcript, 3, "roster_lookup", {"service": "payments"}, '{"engineer": "R. Okafor"}')
print(transcript[-1])go deeper
Be able to name the three line types — thought, action, observation — and say plainly that the runtime, not the model, writes observations.
Explain why a self-generated observation is indistinguishable from a real one in the transcript, and what labelling makes a trace auditable.
Talk about logging tool invocations independently of the transcript so fabricated observations are detectable, and about delimiting tool output so it is never read as instruction.
Own the provenance contract across the system: what evidence every asserted claim must trace back to, and how that trace is retained for audit and incident review.
## The three roles in a ReAct transcript A ReAct trace interleaves three kinds of line. A **thought** is the model reasoning in natural language about what to do next. An **action** is the model naming a tool and its arguments. An **observation** is what came back. Thoughts and actions are model output. Observations are not: they are produced outside the model by whatever runs the tool, and inserted into the transcript so the next step can read them. That asymmetry is what makes the pattern work. Everything the model writes is a prediction; the observation is the single channel through which the outside world enters the reasoning. Blur the boundary and the loop becomes a monologue that merely looks grounded. ## Why the model must not write its own observations A transcript full of `Thought: ... / Action: ... / Observation: ...` is a pattern the model can continue in either direction. Left to generate freely, it will produce the action *and* a plausible observation — a well-formatted search result, a realistic-looking row, a stack trace that has never existed. It then reasons on that fabricated evidence with full confidence, and the resulting trace is indistinguishable from a real one at a glance. The mechanical prevention is that generation halts at the action boundary: the runtime detects the completed action, stops the model, executes the tool, appends the true result, and resumes. Any agent that has ever produced a beautifully formatted answer with no corresponding tool invocation in its logs has hit exactly this failure. ## What belongs in a well-formed observation At minimum, the payload the tool returned. In practice, four things: 1. **A step id.** Numbering observations lets later thoughts refer to them ('per observation 3, the account was closed in March') instead of copying the content forward, which keeps long transcripts from re-stating the same payload at every step. 2. **The source tool.** Two searches with different tools produce differently trustworthy text; the model and any human reader need to know which produced this. 3. **The arguments actually used.** This is the cheapest debugging aid in the whole loop. When an observation is surprising, the first question is always whether the query was what the thought intended. 4. **A reduction marker, if any.** If the payload was shortened, the observation says so, so the model does not read a slice as the whole. ## Provenance is also a safety property Observation text is not authored by you. It is a web page, a document, a log line, a database row — arbitrary content that ends up sitting inside the model's context. Labelling it clearly as tool output rather than instruction is the baseline hygiene that keeps the model from treating retrieved text as a directive. At the level of the loop, the discipline is simple: observations are delimited, attributed, and never merged into the instruction region of the transcript. ## How it reads in a trace ``` Thought 3: I need the current on-call engineer for the payments service. Action 3: roster_lookup(service="payments", at="2026-08-19T22:10Z") Observation 3 (roster_lookup, service=payments): { "engineer": "R. Okafor", "shift_ends": "2026-08-20T06:00Z" } Thought 4: Per observation 3 the on-call is R. Okafor until 06:00Z, so I can page them directly. ``` Notice that thought 4 cites the observation rather than restating the JSON. That habit is what keeps a twenty-step transcript from growing quadratically, and it makes the trace auditable: every claim in the final answer traces back to a numbered observation or it is unsupported. ## What junior candidates commonly get wrong The frequent misconception is that the model 'calls' the tool. It does not; it emits a request, and something else executes it. A second is that the observation is the tool's raw response verbatim — often it is a reduced or reformatted version, and that transformation is a deliberate, lossy design decision rather than an accident. A third is treating the observation as trusted, authoritative text simply because it arrived from a tool.
- What happens if the model is allowed to keep generating past its action?It writes its own observation. The text is fluent and correctly formatted, so the following thoughts reason on evidence that never existed, and the final answer looks fully grounded. The tell is a trace whose observations have no matching tool execution in the runtime's logs — which is why every agent should log invocations independently of the transcript.
- Why number observations instead of letting later thoughts restate their content?Restating copies the payload forward at every step, so a long run re-pays for the same tokens repeatedly and drifts as each restatement paraphrases the last. A numbered reference — 'per observation 3' — is a few tokens, keeps the original text as the single source, and makes the final answer auditable claim by claim.
saying these in an interview costs you the question
- Saying the model executes the tool itself
- Letting the model generate the observation text
- Appending raw output with no tool or argument label
- Treating observation text as trusted instructions
- Restating full payloads in every later thought