skip to content

How do you keep an OpenAI tool-calling agent loop from running forever?

level: seniorimportance: should knowfreq 52%

answer

  1. Exit on absence of calls, not on hope
  2. Steps are the cheapest guard
  3. Cost grows superlinearly with steps
  4. Same call twice means stuck
  5. Check the forcing setting first

basics

~20 s

Loop until a response has no tool_calls and finish_reason is stop, and bound it: a maximum step count, a token or cost budget, per-tool timeouts, and duplicate-call detection. Force a prose answer with tool_choice none on the final step.

solid answer

~50 s

The natural exit is a response whose `message.tool_calls` is empty — usually with `finish_reason: "stop"` — so branch on that, not on a guess. But the natural exit is not guaranteed, so wrap it in hard bounds: a maximum iteration count, a cumulative token or cost budget summed from `usage` on each response, and a wall-clock deadline for the whole task. Add per-tool timeouts so a hung dependency cannot stall the loop, and detect repeats — the same function name with identical arguments twice in a row means the model is stuck, and the cure is to feed that back as a tool result rather than run it again. On the last permitted iteration, send `tool_choice: "none"` so the model is forced to answer with what it has instead of requesting one more lookup. Check the trivial cause first: leaving `tool_choice: "required"` set on every request makes termination impossible by construction.

code

python · 34 lines
python
import json, os, time
from openai import OpenAI

client = OpenAI()
MODEL = os.environ["OPENAI_MODEL"]

def run_loop(messages, tools, handlers, max_steps=8, token_budget=60_000, deadline_s=60):
    seen, spent, start = set(), 0, time.monotonic()
    for step in range(max_steps):
        last = step == max_steps - 1 or time.monotonic() - start > deadline_s
        resp = client.chat.completions.create(
            model=MODEL, messages=messages, tools=tools,
            tool_choice="none" if last else "auto",
        )
        spent += resp.usage.total_tokens
        msg = resp.choices[0].message
        if not msg.tool_calls:
            return msg.content                      # natural exit
        messages.append(msg)
        for call in msg.tool_calls:
            key = (call.function.name, call.function.arguments)
            if key in seen:
                content = json.dumps({"error": "identical call already made; vary the arguments"})
            else:
                seen.add(key)
                try:
                    args = json.loads(call.function.arguments)
                    content = json.dumps(handlers[call.function.name](**args))
                except Exception as exc:
                    content = json.dumps({"error": str(exc)})
            messages.append({"role": "tool", "tool_call_id": call.id, "content": content})
        if spent > token_budget:
            break
    return None                                     # bound hit; report it

go deeper

for a junior

Know that the loop ends when the response contains no tool calls, and that you should always add a maximum number of iterations as a safety net.

for a middle

Explain the exit condition precisely, the step and token bounds, and the tool_choice mistake that makes a normal exit impossible.

for a senior

Show the full guard set in a production loop — budgets summed from usage, deadlines, per-tool timeouts that still emit a message, repeat detection, and a forced prose turn at the boundary — plus what you log when a bound fires.

for a principal

Own the policy across services: what a task is allowed to spend, how bounded loops are made observable, and when a runaway trace is treated as a tool-design defect rather than a cap to raise.

## The shape of the loop An agent loop is: send messages plus tools → if the reply has `tool_calls`, execute them and append results → repeat. The exit is a reply with no calls. Everything difficult about it is what happens when that exit does not arrive. ## The correct exit condition Branch on `message.tool_calls` being empty rather than on `finish_reason` alone. `"stop"` accompanies a normal prose answer, but you may also see `"length"` (output ceiling hit — not a completed answer, and worth surfacing as an error), and under a forced tool choice you will see `"tool_calls"` every single time. Code that only ever tests `finish_reason == "stop"` hangs on the first truncated response. ## Bound one: step count The simplest and most important guard. Pick a ceiling — often 5 to 15 depending on the task — and count iterations. When you hit it, stop calling the model in tool mode and produce something for the user: either the model's last prose, or one final request with `tool_choice: "none"` so it must summarize what it has learned. The second is nearly always better UX than a raw timeout, because the model has real intermediate results in the transcript. ## Bound two: token and cost budget Step count alone is a poor proxy for spend, because each iteration re-reads the whole transcript, which grows monotonically. Iteration ten costs far more than iteration one. Accumulate `usage.prompt_tokens` and `usage.completion_tokens` from each response, convert to money, and abort when the task exceeds its budget. This is what stops a single pathological request from consuming a day's quota. Note that with streaming you must explicitly ask for usage to be included, or you will have no counts to sum. ## Bound three: wall clock A deadline for the whole task, checked before each iteration, protects against the case where each individual step is legal but their sum blows past any latency the caller will tolerate. Pair it with per-tool timeouts — remembering that a timed-out call still needs a tool message carrying an error, or the transcript becomes unsendable. ## Bound four: repetition detection The most common non-terminating pattern is not infinite creativity, it is a stuck loop: the model calls `search("quarterly revenue")`, gets an unhelpful result, and calls `search("quarterly revenue")` again. Hash `(name, arguments)` per iteration and keep a set. On a repeat, do not execute — return a tool message saying so, e.g. "this exact call was already made and returned the result above; try different arguments or answer with what you have". That nudge breaks the cycle far more often than raising the step cap. ## Bound five: the configuration trap Before blaming the model, check `tool_choice`. If it is `"required"` on every request, prose is forbidden by construction and the loop cannot terminate normally — it will only ever stop when your step cap fires. Force on the first turn if you must, then fall back to `"auto"`. ## Errors belong in the transcript When a tool throws, put a short error message into that call's tool content rather than raising out of the loop. The model can then adapt: fix the arguments, choose a different tool, or explain the failure. Cap this — three consecutive failures of the same tool means stop and report, not keep trying. And never put credentials, internal hostnames, or raw stack traces into that content: it goes into the prompt and, via the model's answer, potentially in front of the user. ## Context growth A long loop grows the transcript until it approaches the context window; you will then start seeing failures or aggressive truncation. Mitigations are summarizing older turns into a compact note, dropping large tool outputs once they have been consumed, or storing bulky results out-of-band and passing a handle. Whatever you trim, keep the transcript legal: an assistant message with `tool_calls` and its matching tool messages must be removed together, never one without the other. ## What to report when you stop A loop that hits a bound should fail loudly in your telemetry — record the terminating bound, the step count, tokens spent, and the sequence of tool names. Those traces are how you discover that one tool's description is ambiguous, or that a particular class of question always exhausts the budget. Silent capping hides a design problem behind a bill.

  • Why is a step cap alone a weak cost control?
    Because each iteration re-sends the entire transcript, which grows with every tool result. Step ten can cost several times step one, so a cap of fifteen bounds the count but not the spend — a task with large tool outputs can blow a budget well inside it. Sum prompt and completion tokens from each response's usage object and abort on money, not on iterations. With streaming, request usage explicitly or you will have nothing to sum.
  • The model keeps issuing the same search call. What is the fix?
    Detect it by hashing function name plus the exact arguments string, and on a repeat return a tool message stating the call was already made and its result is above, with a suggestion to vary the arguments or answer directly. That feedback breaks the cycle far more reliably than raising the step cap. If a tool triggers it habitually, the description or its return shape is unclear — a result that says "no matches" beats an empty array the model reads as a transient failure.
  • Should tool exceptions be raised out of the loop or returned to the model?
    Returned, in almost all cases — a short error object as the tool message content lets the model correct arguments, fall back to another tool, or explain the failure to the user. Raise only for conditions the model cannot act on, such as an auth failure or an exhausted budget. Cap consecutive failures of the same tool, and scrub the content: it enters the prompt and can surface in the answer, so no stack traces, hostnames, or credentials.

saying these in an interview costs you the question

  • Exiting only on finish_reason 'stop' and hanging on truncation
  • Leaving tool_choice 'required' on every turn of the loop
  • Trusting the model to stop on its own with no step cap
  • Raising a tool exception out of the loop instead of feeding it back
  • Treating step count as a proxy for spend as the transcript grows

context