What stops a LangChain create_agent loop from running forever?
answer
- two stop paths, one graceful one hard
- counts steps, not turns
- the default is lower than it looks
- exception, not a partial answer
- set it per invocation, not at build time
basics
~20 sAn agent from create_agent normally stops when the model returns a message with no tool calls. The hard backstop is the graph recursion limit, default 25 super-steps, which raises GraphRecursionError rather than returning a partial answer.
solid answer
~50 sThere are two termination paths and they behave very differently. The **normal** one is semantic: the model returns an `AIMessage` with content and no `tool_calls`, and the loop ends. The **backstop** is structural: LangChain v1's `create_agent` returns a compiled LangGraph graph, and graph execution is capped by `recursion_limit`, which defaults to 25 and is set per invocation via config, for example `agent.invoke(state, {"recursion_limit": 50})`. Exceeding it raises `GraphRecursionError`. Two details bite in production. First, the limit counts execution super-steps, not agent turns — a model call plus a tool call is two steps, so the default allows roughly a dozen tool-using iterations. Second, hitting it is an **exception**, not a graceful "I ran out of steps" answer, so the caller gets nothing usable unless you catch it. If you want a bounded run that still returns something, add your own cap in a `before_model` hook or middleware that ends the run, and treat `recursion_limit` purely as a runaway guard.
code
python · 12 linesfrom langgraph.errors import GraphRecursionError
# agent = create_agent(model, tools)
try:
result = agent.invoke(
{"messages": [{"role": "user", "content": "research this"}]},
{"recursion_limit": 50},
)
final = result["messages"][-1].content
except GraphRecursionError:
final = "Stopped after too many steps; returning partial progress."go deeper
Know that an agent normally stops when the model replies without tool calls, and that a step limit exists as a safety net.
Explain that create_agent compiles to a graph whose recursion_limit defaults to 25, is set per invocation in config, and raises rather than returning a partial answer.
Demonstrate the operational split: recursion_limit as a runaway guard, your own turn/token/time budget as the product behaviour, and diagnosis by reading the message list for oscillation or retry storms.
Own the budget policy across services — per-run cost and latency ceilings, what a degraded answer looks like to the user, and when a task that keeps exhausting its budget should be redesigned as a bounded workflow instead.
## Two kinds of stopping An agent stops for one of two reasons, and conflating them is the classic interview miss. **Semantic termination** is the model deciding it is done: it returns an `AIMessage` with content and an empty `tool_calls` list, and the loop exits normally with a final answer. Everything about prompt design, tool descriptions and model choice feeds this path. Most "my agent loops forever" incidents are really "my agent cannot tell it has succeeded" — a tool that returns an ambiguous empty result, two tools with overlapping descriptions, or a goal with no stated success criterion. **Structural termination** is the framework refusing to continue. In LangChain v1, `create_agent` compiles to a LangGraph graph, and graph execution carries a `recursion_limit`. It defaults to 25 and is supplied at invoke time in the config dict, not at construction: `agent.invoke({"messages": [...]}, {"recursion_limit": 50})`. When the graph exceeds it, execution raises `GraphRecursionError`. ## Super-steps, not turns The limit counts graph super-steps. A tool-using turn is at minimum two of them: the model node runs, then the tool node runs. So a default of 25 corresponds to roughly twelve tool-calling iterations, not twenty-five. Teams that assume one step per turn size the limit at half what they intended and are surprised when a legitimately long research task dies mid-way. When you raise it, reason in steps and leave headroom. ## Why an exception is the wrong user experience `GraphRecursionError` propagates out of `invoke`. Nothing partial is returned by default, so a user waiting ninety seconds gets a stack trace instead of "here is what I found so far". For anything user-facing you want a softer cap layered above the hard one: - A `before_model` middleware hook can inspect the accumulated messages, count how many assistant turns have happened, and end the run with a synthesised final message once a budget is reached. - Cost and latency budgets often matter more than turn counts. Counting tool calls, accumulated tokens or wall-clock time and stopping on whichever trips first is closer to what you actually want to bound. - Prebuilt middleware exists for common cases — `SummarizationMiddleware` to compact history so long runs stay affordable, `HumanInTheLoopMiddleware` to pause for approval rather than letting the agent grind on. The rule of thumb: `recursion_limit` is a runaway guard that protects your bill and your process; your own budget check is the product behaviour. ## What the pre-1.0 world did The legacy `AgentExecutor` had `max_iterations` (default 15), `max_execution_time` for a wall-clock cap, and `early_stopping_method` to control what happened on exhaustion — notably a `"force"` mode that returned a canned stop message instead of raising. Interviewers often probe whether you know why the model changed: the executor owned the loop and could therefore own the graceful stop, whereas the v1 agent is a graph and the graph runtime's guard is a hard error, deliberately leaving graceful degradation to you. In LangChain 1.x this executor-era surface has moved out of the main package into the classic compatibility package; describing it as current API is a dating tell. ## Streaming changes the failure shape If you `stream` the agent instead of invoking it, the caller has already received the intermediate steps by the time the limit trips. That is usually a better experience — the user sees the work, then an error — and it is one more reason to stream long agent runs. It does not make the run recoverable, but it makes partial progress visible. ## Diagnosing a run that hits the limit Read the message list, not the exception. Three signatures recur: 1. **Oscillation** — the same two tools alternate with near-identical arguments. Usually overlapping descriptions, or a tool whose result does not actually answer what the model asked. 2. **Retry storm** — the same tool called repeatedly with slightly mutated arguments after each error. The model is guessing at a schema; tighten the schema and return an actionable error message. 3. **No stopping criterion** — the agent keeps gathering more evidence because nothing tells it when it has enough. Fix in the system prompt by stating what a complete answer looks like, or by using structured output so "done" has a shape. Raising `recursion_limit` fixes none of those; it just makes them more expensive. Raise it only when you have confirmed the task genuinely needs more steps.
- Your agent hits the recursion limit on a task you know needs eight tool calls. What do you check first?Read the message list rather than raising the limit. Eight calls is well inside a default of 25 super-steps, so the run is not simply long — look for oscillation between two tools with overlapping descriptions, or a retry storm where the same tool is called repeatedly after a schema error. Both are prompt or tool-metadata bugs; raising the cap only makes them cost more.
- How would you give the user a partial answer instead of an exception when the budget runs out?Layer a soft cap above the hard one: a before_model middleware hook counts assistant turns, tokens or elapsed time and ends the run with a synthesised final message once the budget trips, leaving recursion_limit purely as a runaway guard. Streaming also helps, because the caller has already received intermediate steps before any error surfaces.
- Why does the default of 25 allow fewer than 25 agent iterations?The limit counts graph super-steps, and a tool-using iteration costs at least two — the model node then the tool node. So 25 allows roughly a dozen iterations. Teams that reason in turns rather than steps consistently set the cap at about half of what they intended.
saying these in an interview costs you the question
- Assuming recursion_limit counts agent turns rather than graph steps
- Expecting a partial result instead of GraphRecursionError
- Raising the limit as the first fix for a looping agent
- Setting the cap at construction instead of per invocation
- Citing AgentExecutor max_iterations as the current v1 mechanism