skip to content

Termination, Budgets and Loop Detection

Every agent loop needs an exit. Beyond "the model stopped calling tools" you need iteration caps, token and cost budgets, and a way to notice the agent that keeps calling the same tool forever.

part ofAI agentsoverview, primer and where to startread it →
on this pageshow

questions

4

What stop conditions end an agent loop besides the model returning no tool call?

level: middleimportance: must knowfreq 62%

answer

  1. two families of exit
  2. the model's own exit signal
  3. harness-owned caps sit underneath
  4. iterations, wall clock, tokens, cost
  5. return the stop reason, not just text

basics

~20 s

Agent loops end naturally when the model returns a final answer with no tool call, or calls an explicit done tool. They end forcibly on a max-iteration cap, a wall-clock deadline, a token or cost budget, or an unrecoverable error.

solid answer

~50 s

There are two families. **Natural stops** come from the model: it returns a text-only turn with no tool call, or it invokes an explicit `done`-style tool whose schema carries the final answer. The done tool is unambiguous and gives typed output, but the model can forget to call it, so most harnesses keep the no-tool-call path as a fallback. **Forced stops** come from the harness, because the model's judgement is not a safety property: a maximum iteration count, a wall-clock deadline (a hung tool burns time without burning iterations), a token budget (history grows every turn, so cost climbs superlinearly with steps), a session cost cap in currency, and error stops after repeated tool failures or an external cancellation. Whichever fires, the loop should return a status naming the condition rather than just an answer — the distribution of stop reasons is one of your best health metrics.

code

python · 14 lines
python
import time

def run_agent(step, max_iters=20, deadline_s=120.0, token_budget=40_000):
    started, used, last = time.monotonic(), 0, None
    for _ in range(max_iters):
        last = step()                      # one model call + any tool execution
        used += last["tokens"]
        if last["done"]:
            return {"status": "complete", "output": last["output"], "tokens": used}
        if used >= token_budget:
            return {"status": "budget_exhausted", "output": last["output"], "tokens": used}
        if time.monotonic() - started >= deadline_s:
            return {"status": "timeout", "output": last["output"], "tokens": used}
    return {"status": "max_iterations", "output": last and last["output"], "tokens": used}

go deeper

for a junior

Know that an agent loop repeats model call, tool call, model call, and that it needs both a completion signal from the model and a hard cap from the code. Be able to name max iterations as the simplest cap.

for a middle

Explain the difference between a natural stop (no tool call, or an explicit done tool) and a forced stop, and say which resource each cap bounds — iterations bound model calls, the clock bounds latency, tokens bound cost.

for a senior

Show you have operated one: describe the exit record you return, why the stop reason is logged as a metric, and what a spike in max-iteration exits tells you about your task distribution or a looping agent.

for a principal

Own the policy across a fleet: where caps are set (per task class, per tenant), who can raise them, how cost caps interact with model routing, and the argument that termination guards bound spend and blast radius while correctness needs separate verification.

## Termination is a design decision, not a default An agent loop is a `while` loop around a model call: the harness sends the conversation, the model either answers or requests a tool, the harness executes the tool, appends the result, and calls the model again. Nothing in that structure guarantees an end. The model decides when it is finished, and a model that is uncertain, distracted, or working from a vague goal can keep asking for tools indefinitely. The first wave of autonomous-agent projects became notorious for exactly this: a task loop that re-decomposed its own goal forever, spending money and producing nothing. Every serious harness since then treats "how does this stop?" as a first-class part of the design. ## Natural stops: the model signals completion **No tool call.** The simplest contract: if the assistant turn contains only text and no tool request, the harness treats that text as the final answer and exits. It costs nothing to implement and works with any model. Its weakness is ambiguity — models sometimes emit a conversational text turn in the middle of a task ("Let me check the logs next."), and a naive harness reads that as completion and returns a half-finished answer. **An explicit done tool.** You give the model a `submit_result` / `finish` tool with a JSON schema, and completion means calling it. This is unambiguous, and it lets you *require* structure: the final answer field, a confidence value, a list of things it could not resolve. Downsides: it is one more tool competing for selection, and the model may simply never call it — so production harnesses usually accept both signals, treating a text-only turn as completion too, or nudging once ("call submit_result when you are done") before accepting it. ## Forced stops: the harness owns the bound Natural stops are the happy path. Forced stops exist because the happy path is not guaranteed. - **Max iterations.** A cap on trips through the loop. The cheapest and most universal guard; it bounds the number of model calls but nothing else. - **Wall-clock deadline.** A tool that hangs on a slow network consumes no iterations at all, so an iteration cap alone does not bound latency. Anything user-facing needs a clock. - **Token budget.** Each turn resends the whole conversation, so per-step cost grows as the transcript grows — twenty steps costs far more than twice ten. A token counter is the only guard that tracks the thing that actually scales. - **Cost cap.** A currency-denominated ceiling per session or per tenant, sitting above tokens, so a switch to a pricier model does not silently multiply spend. - **Error stops.** N consecutive tool failures, an authentication error, a schema the model cannot satisfy — retrying these forever is pure waste. - **External cancellation.** The user closed the tab or the request was aborted; the loop must observe cancellation between steps. Each of these catches a different failure. Iterations catch a chatty agent, the clock catches a hung tool, tokens catch a transcript that ballooned from one huge tool result, cost catches an expensive model, error stops catch a broken dependency. Shipping only one of them means the other four failure modes run unbounded. ## The exit contract The loop's return value should be a small record, not a bare string: a **status** (`complete`, `max_iterations`, `timeout`, `budget_exhausted`, `error`, `cancelled`), whatever partial output exists, and the usage actually consumed. Callers need the status to decide whether to display, retry, escalate, or degrade — and you need it in telemetry. A dashboard of stop reasons is remarkably diagnostic: if ten percent of runs end at `max_iterations`, either the cap is too tight for the task distribution or the agent is looping, and both are worth investigating before users notice. ## What termination is not A clean stop is not a correct answer. The model can call the done tool while confidently reporting work it never did; a run that ends inside its budget is not thereby a successful run. Termination guards bound *cost and blast radius*. Correctness is the job of verifiers, evals and, for irreversible actions, human approval — different mechanisms entirely.

  • Why is a max-iteration cap not enough on its own?
    It bounds only the number of model calls. A tool that blocks for two minutes consumes one iteration but two minutes of latency, and a single huge tool result can blow the token budget inside three steps. Iterations, wall clock, tokens and cost each bound a different resource, so a production loop checks all of them.
  • What breaks if the harness treats every text-only turn as completion?
    The model sometimes narrates mid-task — "Now I'll check the deploy logs" — with no tool call attached. A harness that exits there returns a confident-sounding partial answer and the caller has no way to tell. Mitigations: require an explicit done tool, or re-prompt once asking whether the task is finished before accepting the text as final.
  • How would you tell a genuine completion from a premature one in telemetry?
    Record the stop reason plus a cheap completeness check: did the run touch the tools the task requires, does the final output satisfy the result schema, did any required field come back empty. Then sample runs that stopped early and grade them. A rising rate of short `complete` runs with no tool usage is the signature of a model bailing out.

saying these in an interview costs you the question

  • Assumes the model reliably stops on its own
  • Treats max iterations as the only guard needed
  • Thinks an explicit done tool proves the answer is correct
  • Returns only text, never the reason the loop stopped
  • Bounds steps but never tokens, time or cost

context

open as a page

How does an advisory token budget an agent paces against differ from an enforced cap?

level: seniorimportance: should knowfreq 44%

basics

~20 s

An advisory budget is told to the model so it can pace itself and wrap up gracefully; an enforced cap is applied by the harness and cuts the run off wherever it happens to be. Advisory budgets shape behaviour, enforced caps bound spend.

open as a page

How do you detect an agent loop that keeps repeating steps without making progress?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Hash each step — tool name, arguments and the observation it returned — and count repeats. When the same hash recurs a few times, or a window of recent hashes cycles, the agent is stuck. Trip a no-progress stop instead of waiting for the iteration cap.

open as a page

When an agent exhausts its budget mid-task, should it return partial results or continue?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

Return the partial result with an explicit exhausted marker whenever partial work is usable — a bug list missing the last level beats nothing at all. Continue only when the task is checkpointable and someone has knowingly authorised the extra spend.

open as a page