What ends a ReAct loop, and which stop conditions can the model itself decide?
answer
- two families: model-decided and harness-decided
- the answer replaces the action
- fails in both directions
- exiting cleanly is not being right
- what you return when the cap fires
basics
~20 sThe model ends the loop by producing a final answer instead of an action. Because that signal is only the model's own judgement, the surrounding harness adds external stops it cannot override: a cap on passes, a cost or time limit, and a goal check on the result.
solid answer
~50 sThere are two families of stop condition. **In-band**: the model emits an answer rather than an action, which is the loop's designed exit. **Out-of-band**: the orchestrator halts the run regardless of what the model wants — a maximum number of passes, a wall-clock or spend limit, an explicit goal predicate the harness evaluates against the environment, or an operator cancelling. The in-band signal alone is not trustworthy in either direction. A model can stop too early, answering on pass one from memory without ever calling a tool, and it can also never stop, alternating between actions until something external fires. So a production loop treats the model's answer as a *proposal*: it is checked — by a verifier, a schema, a test, or a human — while the caps guarantee the run is bounded no matter how the model behaves.
go deeper
Know that the loop stops when the model gives a final answer instead of another action, and that a maximum number of passes exists as a backstop so a run cannot spin forever.
Separate the two families cleanly — the model's in-band answer versus harness-enforced caps and goal checks — and explain why the in-band signal fails in both directions, stopping too late and stopping too early.
Show judgement about the exhaustion path: what you return when a cap fires, how you flag partial results, and how you detect premature finishes by comparing declared success against an independent check of the environment.
Own the policy: where caps are set across a workload from observed pass-count distributions, what an incomplete run costs the business versus a confidently wrong one, and which actions an agent may never finish unsupervised.
## Two families of stop condition Every ReAct loop needs an answer to "when does this stop?", and there are exactly two places the answer can come from. **In-band, decided by the model.** A pass produces a final answer rather than an action. This is the pattern's designed exit and the only one that reflects task understanding: the agent believes it has enough. Some loops make this explicit rather than implicit, giving the agent a dedicated finishing action it must call, which has the useful property that finishing becomes an observable event with arguments rather than the absence of a tool call. **Out-of-band, decided by the harness.** The loop is bounded from outside by things the model has no say over: a maximum pass count, a wall-clock deadline, a spend limit, a goal predicate the harness itself evaluates against the environment (did the order actually get placed? does the test suite pass?), or an operator cancelling the run. A correct answer to this question names both families and says why neither alone is sufficient. ## Why the model's own signal is not enough The in-band signal fails in both directions. **Never stopping.** The model keeps emitting actions because each pass looks locally reasonable. Nothing in the loop has an opinion about total progress, so a run can spend fifty passes circling. Without an external cap, that is unbounded cost. **Stopping too early.** The more insidious failure. A ReAct agent asked for a live figure returns a fully formed answer on pass one having called nothing, because the question resembled things it could answer from its training data. The transcript looks clean — a thought and an answer — and the answer is fluent and specific. It is also unevidenced, and the agent will report success. This is why success rate cannot be measured by "did the loop exit normally": premature exits exit normally. ## Termination is not the same as correctness The distinction interviewers listen for: **the loop terminating is a control-flow fact; the answer being right is a separate claim requiring separate evidence.** These come apart constantly. A run that hits its pass cap has failed to terminate in-band but may still have gathered enough to return something useful as a partial result. A run that exits cleanly on the model's own say-so may be entirely fabricated. Treat the model's answer as a proposal that gets checked — against a schema, a test, a verifier, a second reading of the environment, or a human — rather than as proof of completion. ## What to do when a cap fires Hitting an external limit is not automatically an error to throw away. The reasonable options, in rough order of preference: - **Return a partial result** with an explicit "incomplete" flag and the trace, so the caller can decide. - **Escalate to a human**, handing over the transcript so far, which is the right default when the remaining step is irreversible. - **Hard fail**, appropriate when a partial result would be misleading or acted on automatically. The pattern to avoid is silently returning the best-so-far as though it were a completed answer, which converts an honest timeout into a confident wrong result. ## Choosing the caps Caps are set from observed traces, not intuition. Look at the pass-count distribution for successful runs on your task suite: if the median success takes four passes and the tail takes nine, a cap of ten bounds cost while barely touching legitimate work, and a cap of thirty mostly funds runaway runs. Setting the cap too tight is not free either — it cuts off the genuinely hard cases, which are frequently the ones that mattered. On irreversible actions, the interesting limit is often not the pass count at all but whether the agent should be allowed to finish unsupervised. A run that will place an order or dispatch a crew is one where the last step deserves confirmation regardless of how many passes it took. ## Answering well Name the in-band exit first, since it is the pattern's own mechanism. Then name the external bounds and explain that they exist because the model's judgement about being finished is exactly the thing under test. Close with the premature-answer case: it is the failure most candidates forget, and mentioning it signals you have watched real runs rather than only read about the pattern.
- A run hits its maximum pass count. What should it return?Rarely a bare error. Prefer a partial result explicitly flagged incomplete, carrying the trace so the caller can judge what was established, or escalate to a human when the remaining step is irreversible. The one thing to avoid is returning the best-so-far as if it were a finished answer, which turns an honest timeout into a confident wrong result.
- Why do some loops give the agent an explicit finishing action instead of just letting it answer?Because it turns termination into an observable event rather than the absence of one. A finishing call has a name and arguments, so the harness can validate the claimed result against a schema or predicate, log it, and distinguish a deliberate finish from a malformed pass. It also makes premature finishing something you can count and alert on.
- How would you detect premature termination across many runs?Compare declared success against an independent check of the environment, and watch the action count: runs that finish with zero or one tool call on tasks that require live data are the prime suspects. On an eval suite with known-correct outcomes, premature exits show up as confidently wrong answers reached in far fewer passes than the reference trajectory.
saying these in an interview costs you the question
- Says the step cap is the only stop condition needed
- Treats a clean exit as proof the answer is correct
- Forgets the model can answer without ever acting
- Returns best-so-far at the cap as a completed result
- Picks caps by intuition rather than from observed traces