What stop conditions end an agent loop besides the model returning no tool call?
answer
- two families of exit
- the model's own exit signal
- harness-owned caps sit underneath
- iterations, wall clock, tokens, cost
- return the stop reason, not just text
basics
~20 sAgent loops end naturally when the model returns a final answer with no tool call, or calls an explicit done tool. They end forcibly on a max-iteration cap, a wall-clock deadline, a token or cost budget, or an unrecoverable error.
solid answer
~50 sThere are two families. **Natural stops** come from the model: it returns a text-only turn with no tool call, or it invokes an explicit `done`-style tool whose schema carries the final answer. The done tool is unambiguous and gives typed output, but the model can forget to call it, so most harnesses keep the no-tool-call path as a fallback. **Forced stops** come from the harness, because the model's judgement is not a safety property: a maximum iteration count, a wall-clock deadline (a hung tool burns time without burning iterations), a token budget (history grows every turn, so cost climbs superlinearly with steps), a session cost cap in currency, and error stops after repeated tool failures or an external cancellation. Whichever fires, the loop should return a status naming the condition rather than just an answer — the distribution of stop reasons is one of your best health metrics.
code
python · 14 linesimport time
def run_agent(step, max_iters=20, deadline_s=120.0, token_budget=40_000):
started, used, last = time.monotonic(), 0, None
for _ in range(max_iters):
last = step() # one model call + any tool execution
used += last["tokens"]
if last["done"]:
return {"status": "complete", "output": last["output"], "tokens": used}
if used >= token_budget:
return {"status": "budget_exhausted", "output": last["output"], "tokens": used}
if time.monotonic() - started >= deadline_s:
return {"status": "timeout", "output": last["output"], "tokens": used}
return {"status": "max_iterations", "output": last and last["output"], "tokens": used}go deeper
Know that an agent loop repeats model call, tool call, model call, and that it needs both a completion signal from the model and a hard cap from the code. Be able to name max iterations as the simplest cap.
Explain the difference between a natural stop (no tool call, or an explicit done tool) and a forced stop, and say which resource each cap bounds — iterations bound model calls, the clock bounds latency, tokens bound cost.
Show you have operated one: describe the exit record you return, why the stop reason is logged as a metric, and what a spike in max-iteration exits tells you about your task distribution or a looping agent.
Own the policy across a fleet: where caps are set (per task class, per tenant), who can raise them, how cost caps interact with model routing, and the argument that termination guards bound spend and blast radius while correctness needs separate verification.
## Termination is a design decision, not a default An agent loop is a `while` loop around a model call: the harness sends the conversation, the model either answers or requests a tool, the harness executes the tool, appends the result, and calls the model again. Nothing in that structure guarantees an end. The model decides when it is finished, and a model that is uncertain, distracted, or working from a vague goal can keep asking for tools indefinitely. The first wave of autonomous-agent projects became notorious for exactly this: a task loop that re-decomposed its own goal forever, spending money and producing nothing. Every serious harness since then treats "how does this stop?" as a first-class part of the design. ## Natural stops: the model signals completion **No tool call.** The simplest contract: if the assistant turn contains only text and no tool request, the harness treats that text as the final answer and exits. It costs nothing to implement and works with any model. Its weakness is ambiguity — models sometimes emit a conversational text turn in the middle of a task ("Let me check the logs next."), and a naive harness reads that as completion and returns a half-finished answer. **An explicit done tool.** You give the model a `submit_result` / `finish` tool with a JSON schema, and completion means calling it. This is unambiguous, and it lets you *require* structure: the final answer field, a confidence value, a list of things it could not resolve. Downsides: it is one more tool competing for selection, and the model may simply never call it — so production harnesses usually accept both signals, treating a text-only turn as completion too, or nudging once ("call submit_result when you are done") before accepting it. ## Forced stops: the harness owns the bound Natural stops are the happy path. Forced stops exist because the happy path is not guaranteed. - **Max iterations.** A cap on trips through the loop. The cheapest and most universal guard; it bounds the number of model calls but nothing else. - **Wall-clock deadline.** A tool that hangs on a slow network consumes no iterations at all, so an iteration cap alone does not bound latency. Anything user-facing needs a clock. - **Token budget.** Each turn resends the whole conversation, so per-step cost grows as the transcript grows — twenty steps costs far more than twice ten. A token counter is the only guard that tracks the thing that actually scales. - **Cost cap.** A currency-denominated ceiling per session or per tenant, sitting above tokens, so a switch to a pricier model does not silently multiply spend. - **Error stops.** N consecutive tool failures, an authentication error, a schema the model cannot satisfy — retrying these forever is pure waste. - **External cancellation.** The user closed the tab or the request was aborted; the loop must observe cancellation between steps. Each of these catches a different failure. Iterations catch a chatty agent, the clock catches a hung tool, tokens catch a transcript that ballooned from one huge tool result, cost catches an expensive model, error stops catch a broken dependency. Shipping only one of them means the other four failure modes run unbounded. ## The exit contract The loop's return value should be a small record, not a bare string: a **status** (`complete`, `max_iterations`, `timeout`, `budget_exhausted`, `error`, `cancelled`), whatever partial output exists, and the usage actually consumed. Callers need the status to decide whether to display, retry, escalate, or degrade — and you need it in telemetry. A dashboard of stop reasons is remarkably diagnostic: if ten percent of runs end at `max_iterations`, either the cap is too tight for the task distribution or the agent is looping, and both are worth investigating before users notice. ## What termination is not A clean stop is not a correct answer. The model can call the done tool while confidently reporting work it never did; a run that ends inside its budget is not thereby a successful run. Termination guards bound *cost and blast radius*. Correctness is the job of verifiers, evals and, for irreversible actions, human approval — different mechanisms entirely.
- Why is a max-iteration cap not enough on its own?It bounds only the number of model calls. A tool that blocks for two minutes consumes one iteration but two minutes of latency, and a single huge tool result can blow the token budget inside three steps. Iterations, wall clock, tokens and cost each bound a different resource, so a production loop checks all of them.
- What breaks if the harness treats every text-only turn as completion?The model sometimes narrates mid-task — "Now I'll check the deploy logs" — with no tool call attached. A harness that exits there returns a confident-sounding partial answer and the caller has no way to tell. Mitigations: require an explicit done tool, or re-prompt once asking whether the task is finished before accepting the text as final.
- How would you tell a genuine completion from a premature one in telemetry?Record the stop reason plus a cheap completeness check: did the run touch the tools the task requires, does the final output satisfy the result schema, did any required field come back empty. Then sample runs that stopped early and grade them. A rising rate of short `complete` runs with no tool usage is the signature of a model bailing out.
saying these in an interview costs you the question
- Assumes the model reliably stops on its own
- Treats max iterations as the only guard needed
- Thinks an explicit done tool proves the answer is correct
- Returns only text, never the reason the loop stopped
- Bounds steps but never tokens, time or cost