In a text-protocol ReAct prompt, why set a stop sequence at "Observation:"?
answer
- who is allowed to write the observation
- text protocol, two speakers, one transcript
- generation must halt at the handoff
- otherwise the model invents tool results
- stop string must match exemplar labels
basics
~20 sA model that has seen Thought/Action/Observation exemplars will happily write the observation itself. Halting generation at the "Observation:" label hands control back to your runtime, which calls the real tool and appends the true result.
solid answer
~40 sPrompted ReAct is a **text protocol**: the model owns the `Thought:` and `Action:` lines, and the environment owns the `Observation:` line. Nothing in plain text marks that boundary as someone else's turn — the exemplars show complete triples, so the statistically likely continuation after an action is a plausible-looking observation the model invents. Configuring a stop string such as `\nObservation:` halts decoding exactly at the handoff. Your runtime then parses the action, executes the tool, appends the real observation, and re-sends the transcript for the next step. Without it the failure is silent and expensive: the trace reads like a perfect agent run, but the final answer rests on fabricated tool output. The string must match the exemplar labels exactly, including leading newline, colon and capitalization, or generation never halts.
code
markdown · 8 linesThought: I need the current status of ticket 4471.
Action: LookupTicket[4471]
Observation: Ticket 4471 was closed on 3 March after a refund was issued.
Thought: The ticket is already closed, so no action is needed.
Final Answer: Ticket 4471 is closed; a refund was issued on 3 March.
<!-- No tool ever ran. Everything from "Observation:" onward was generated
because no stop sequence halted decoding at the handoff. -->go deeper
Know that in a text-based ReAct prompt the model writes the thought and the action, and your code writes the observation. Say plainly that generation must be stopped before the observation so the real tool result can be inserted.
Be ready to explain the mechanism: the exemplars teach a complete triple, so the likeliest continuation after an action is an invented observation. Describe the stop string, where it sits, and why it must match the exemplar labels character for character.
Show you have caught this in production. Talk about the silent-failure shape, comparing tool-invocation counts against observation lines in traces, logging the raw completion before parsing, and handling actions that come back unparsable.
Own the framing question: whether the handoff should live in a fragile string convention at all, versus a structured tool-calling interface where the runtime enforces the boundary. Weigh portability and trace control against the parsing tax you are choosing to keep.
## What the text protocol actually is In its original, prompted form ReAct is not an API feature — it is a **convention encoded in plain text**. Each step of the agent is one turn of a small script with three labelled blocks: - `Thought:` — free-form reasoning, written by the model. - `Action:` — a tool invocation in whatever grammar the prompt declares, written by the model. - `Observation:` — the result that came back from the world, written by *your runtime*. Two of the three blocks belong to the model; the third belongs to the environment. The whole correctness of the loop depends on the model never writing the third one. ## Why the model writes the observation if you let it A language model does one thing: continue the text. Your few-shot exemplars show complete trajectories — thought, action, observation, thought, action, observation, final answer. From the model's point of view, nothing distinguishes the observation line from the two lines above it. So after it emits an action line, the highest-probability continuation is a newline, the literal token `Observation:`, and a fluent, well-formed, entirely invented result. It then reasons confidently on top of that invention, emits another action, invents that observation too, and closes with a final answer. This failure mode is nastier than a crash because it is **silent and well-formed**. The transcript looks like a textbook agent run. Reviewers reading the trace see thoughts grounded in observations. Nothing is grounded in anything; no tool was ever executed. In a domain where the answer is checkable — a price, a ticket status, a row count — you find out later and expensively. ## What a stop sequence does Every mainstream inference API lets you supply one or more **stop strings**: decoding halts as soon as one is produced. For prompted ReAct the conventional choice is the observation label with a leading newline. The loop then runs like this: 1. Send the transcript so far, with the stop string configured. 2. The model emits a thought and an action, starts the observation label, and generation halts. 3. The runtime parses the action out of the completion and executes the real tool. 4. The runtime appends the true `Observation: ...` line to the transcript. 5. Repeat, until the model emits a final answer instead of an action. The stop sequence is what makes step 3 possible at all. It is the seam between the model's turn and the world's turn. ## Choosing the string A few practical rules: - **Match the exemplars exactly.** If the exemplars write `Observation:` but you stop on `observation:` or `Observation :`, generation never halts and you are back to fabricated results. - **Include the leading newline.** Stopping on the bare word risks halting mid-thought if the model uses the word conversationally. - **Stop at the right boundary.** A stop on `Action:` cuts off the call itself; a stop on `Thought:` truncates the reasoning. The observation label is the only correct seam. - **Strip defensively.** Providers differ on whether the matched stop text is echoed back in the completion, so your parser should tolerate both. - A second stop string on the exemplar separator is cheap insurance if the model tends to start a fresh fabricated example. ## What a stop sequence is not It is not a safety control, not a cost control, and not a validator. A maximum-token cap truncates on length and will happily land in the middle of an action line, which is worse than useless. The stop sequence also cannot rescue a malformed action: if the model writes an action your parser cannot read, you still need a repair path — re-prompt with the parse error, or fail the step. ## Where it disappears When you use a provider's structured tool-calling interface instead of a hand-written text protocol, the API itself halts generation at the tool-call boundary and returns the call to you as data. There is no stop string to configure and no label to match, because the handoff is expressed in the message structure rather than in prose. The stop sequence is an artefact of running ReAct in plain text — which is exactly why interviewers use it to check whether you understand what the text protocol is actually doing. ## Detecting the bug In a running system, instrument the seam. Count real tool invocations and compare them with the number of observation lines in the transcript; any transcript with more observations than executions contains at least one fabrication. Log the raw completion before parsing, so you can see whether the model ran past the boundary. Asserting "every observation in this transcript was written by the runtime" is a cheap, high-value invariant.
- What breaks if your exemplars label results differently from the string you stop on?Nothing halts. The stop condition never matches, so the model runs straight past the boundary and writes its own observation, and possibly several further steps. Because the output still looks well formed, the failure surfaces only as wrong answers. Treat the label in the exemplars, the label your runtime appends, and the stop string as one constant defined in a single place.
- Do you still need a stop sequence when the provider exposes native tool calling?No. With a structured tool-calling interface the API halts at the tool-call boundary itself and returns the call as data, so there is no label to match and no prose to parse. The handoff moves from a string convention into the message structure. You still need the loop and the decision about when to stop looping — you just no longer configure a stop string for it.
- Can a max-token limit substitute for the stop sequence?No. A token cap halts on length, not on meaning, so it typically truncates in the middle of an action line — leaving you an unparsable call rather than a clean handoff. It also cannot distinguish a short legitimate step from a long one. Keep the cap as a runaway guard and use the stop string for the protocol boundary.
saying these in an interview costs you the question
- Thinks naming an action in text actually executes the tool
- Believes stop sequences exist mainly to save output tokens
- Assumes a fabricated observation is obvious in the final answer
- Stops on the action label, cutting off the call itself
- Thinks a max-token cap does the same job