What makes a Koog agent hit its maxIterations limit mid-run?
answer
- Liveness guard, not a spend cap
- Two families: wiring or model
- Deterministic failure points at the graph
- Intermittent failure points at a tool
- Trace which node ran each iteration
basics
~20 sUsually a graph that cannot reach nodeFinish — a missing or shadowed assistant-message exit edge — or a model stuck re-calling a tool that keeps failing. Koog aborts the run rather than returning a partial answer.
solid answer
~50 s`maxIterations` on `AIAgent` caps how many steps the agent may take through its strategy before Koog stops the run with an error; it is a liveness guard, not a cost budget. Two causes dominate. First, a wiring gap: some node has no reachable path to `nodeFinish` — typically a missing `onAssistantMessage` edge on `nodeLLMSendToolResult`, or a narrow edge shadowed by a catch-all declared above it — so the graph cycles forever. That is deterministic and reproducible with a unit test. Second, a model-side loop: a tool keeps returning an error or an empty result, the model keeps retrying it, and the tool-call edge fires every pass. That one is non-deterministic and shows up as a spike, not a constant failure. Diagnose by logging which node ran each iteration: a repeating node pair is wiring, a repeating tool with changing arguments is the model.
go deeper
Know that maxIterations bounds how many steps an agent takes and that exceeding it fails the run, most often because the strategy has no reachable path to nodeFinish.
Be able to name the classic wiring cause — a missing or shadowed assistant-message edge on the send-tool-result node — and explain why the compiler cannot catch it.
Separate the two families from the symptom: deterministic same-input failure is wiring, intermittent failure correlating with a flaky tool is a model loop, and a node trace per iteration is what tells them apart.
Treat cap hits as a monitored signal with three usual causes — a tool degraded, a model version shipped, a prompt edit removed an exit — and set the number from the graph's honest worst case rather than tuning it upward to silence alerts.
## What the cap is `AIAgent` takes a `maxIterations` parameter that bounds how many iterations the agent performs before Koog gives up and fails the run. It is deliberately blunt: when it trips you get an error, not a truncated answer, because a half-finished agent run silently presented as a result is worse than a visible failure. Be precise about what it is *not*. It is not a token budget and not a spend cap — a run can burn a lot of money in few iterations, or many cheap iterations. It is not a timeout either. It is a liveness guard: proof that the graph is making progress toward `nodeFinish`, and a hard stop when it is not. ## Cause one: the graph cannot terminate Koog type-checks that every edge is a legal hop. It does not and cannot prove that a path to `nodeFinish` exists for every value a node can emit. So the classic failure is structural: - `nodeLLMSendToolResult` has an `onToolCall` edge back to `nodeExecuteTool` but no `onAssistantMessage` edge to `nodeFinish`. The agent handles tool-free questions correctly and hangs on every question that touches a tool. - A broad or unconditional edge is declared above the narrow one that leads out, so the exit is shadowed — edges are taken in declaration order, first match wins, and the shadowed edge is legal, compiling dead code. - A custom node routes back into an earlier node on a condition that the graph never clears. The signature of this class of bug is determinism. The same input fails the same way every time, at roughly the same iteration count, regardless of model or temperature. That makes it cheap to reproduce: a strategy is an ordinary Kotlin object, so you can drive it in a unit test with a stubbed executor and assert the run reaches `nodeFinish`. ## Cause two: the model will not stop The second family is behavioural. The graph is wired correctly and does have an exit, but the model never emits the plain assistant message that takes it. Typical triggers: - A tool returns an error string every time — bad credentials, a 404, an empty result set — and the model keeps retrying it, sometimes with slightly different arguments, sometimes identically. - Two tools contradict each other, and the model oscillates between them looking for agreement. - The system prompt insists the model must verify before answering, and verification never succeeds, so the model obediently never answers. This one is intermittent. It correlates with a downstream dependency being unhealthy, with a specific class of input, or with a model version change. It does not reproduce on your laptop. ## Telling them apart Instrument the run so each iteration records which node executed and, for tool nodes, which tool with which arguments. Then read the trace: - The same two nodes alternating with no new tool activity, ending exactly at the cap: wiring. Fix the graph. - The same tool called repeatedly, arguments drifting or identical, results all errors: model-side loop. Fix the tool or the prompt, not the graph. - A node reached once and never left: no outgoing edge matched. That is wiring again, of the stranded rather than the cycling kind. That classification is what an interviewer at senior level is listening for — the ability to separate "my graph is wrong" from "my dependency is down" from a symptom that looks identical from the outside. ## Choosing the number Set the cap from the graph's shape, not from superstition. Count the maximum sensible tool round-trips for the task and leave headroom — a research agent that legitimately makes a dozen searches needs more than a classifier that calls one tool. Setting it very high to "stop the errors" converts a fast, loud failure into a slow, expensive one; setting it below the honest worst case truncates legitimate long runs, and users experience that as random failure on the hardest queries. Also treat cap hits as a monitored signal rather than noise. A steady low rate is your graph's natural tail. A step change means something moved: a tool started failing, a model version shipped, a prompt edit removed an early exit. Because the cause is nearly always one of those three, the alert is unusually actionable. ## What to do when it trips in production Fail the request cleanly and say so — do not present whatever text happened to be lying around as an answer. Log the node trace with the failure so the classification above can be done after the fact rather than by reproducing it. And if the run did real, non-idempotent work through tools before aborting, know what state it left behind; the cap stops iteration, it does not undo side effects the tools already performed.
- Is raising maxIterations ever the right fix?Only when the trace shows legitimate progress — new tool calls, new information each pass — and the task honestly needs more round-trips than the cap allows. If the trace shows the same node pair alternating or the same failing tool retried, raising the cap just makes the failure slower and more expensive. Fix the graph or the tool first.
- Why does Koog fail the run instead of returning what it has so far?Because a partial run is indistinguishable from a complete one once it is rendered as an answer. Hitting the cap means the graph never reached nodeFinish, so no value was ever produced for the strategy's output; presenting intermediate text would ship an unfinished, possibly wrong answer as a real result. A loud failure is the safer default.
- What state can be left behind when the cap trips mid-run?Whatever the tools already did. The cap stops iteration; it does not roll anything back. If the loop was retrying a tool that writes — sending mail, posting a record, charging something — those side effects have happened, possibly several times. That is the argument for idempotency keys on write tools rather than relying on the loop terminating cleanly.
saying these in an interview costs you the question
- Treats maxIterations as a token or cost budget
- Raises the cap as the first response to failures
- Assumes compiling proves the graph can terminate
- Thinks a cap hit returns a partial answer
- Ignores side effects tools performed before the abort