skip to content

A CrewAI agent keeps calling the same tool and never finishes — which Agent settings bound it?

level: seniorimportance: should knowfreq 50%

answer

  1. Three different units: steps, seconds, requests
  2. The iteration cap does not raise
  3. Rate limiting is not loop protection
  4. Per-step hook for visibility
  5. Caps contain, they do not cure

basics

~20 s

Three Agent-level caps bound it: max_iter limits reasoning iterations (default 20) and forces a best-effort final answer when hit, max_execution_time caps wall-clock seconds, and max_rpm throttles requests per minute. Tool caching and a step_callback let you see the loop happening.

solid answer

~50 s

On a CrewAI `Agent`, `max_iter` caps how many reasoning-and-tool iterations it may take — default 20 — and reaching it does not raise: the framework pushes the agent to return its best answer with what it has, which is why a looping agent can end in a confidently wrong result rather than an error. `max_execution_time` is the wall-clock bound in seconds and is the one that actually protects a service, because slow tools can burn a long time inside a small iteration count. `max_rpm` throttles that agent's requests per minute; it protects the provider's rate limit, not you from a loop. `cache` (on by default) short-circuits repeated identical tool calls, so a literal same-arguments loop gets cheaper but not shorter. To see the loop, attach a `step_callback` and run `verbose=True` — the callback fires per step and gives you the tool action to log. Caps are containment; the cure is usually the tool's output or the agent's stopping instructions.

code

python · 17 lines
python
from crewai import Agent


def log_step(step):
    print(f"[researcher] step: {step}")


researcher = Agent(
    role="Web Researcher",
    goal="Find three sources on the topic, or report that none exist",
    backstory="If a search returns nothing, you say so rather than searching again.",
    max_iter=6,
    max_execution_time=120,
    max_rpm=20,
    step_callback=log_step,
    verbose=True,
)

go deeper

for a junior

Know that max_iter limits how many steps an agent takes and that max_execution_time limits how long it runs. Be able to set both on an Agent.

for a middle

Explain the units precisely — iterations, seconds, requests per minute — and that hitting the iteration cap yields a forced best-effort answer rather than an exception. Mention tool caching and the per-step callback.

for a senior

Show a diagnosis, not a knob list: read the step trace, identify the ambiguous tool result or missing stop condition, bound the agent so failures stay cheap, then fix the cause. Set a time bound whenever the crew runs behind a caller with a deadline.

for a principal

Own the policy: default caps for every agent class in the system, an alert when agents reach them, and a rule that a run ending at the cap counts as a failed run in evaluation rather than a successful one.

## The failure being described An agent that repeats a tool call is almost always one that cannot tell it is finished. The tool returns something ambiguous — an empty result, an error string, a page of text that does not contain the answer — and the model, seeing no progress and no stop condition it recognises, tries again. Left alone this consumes tokens, wall-clock time and provider quota. CrewAI gives you four `Agent`-level knobs to contain it and one to observe it. ## max_iter `max_iter` is the maximum number of iterations the agent may perform, defaulting to 20. An iteration is one pass of the agent's loop: think, optionally call a tool, observe. When the cap is reached CrewAI does not throw; it forces the agent to give its best final answer with whatever it has. That behaviour is deliberate — a crew of dependent steps degrades better than it crashes — but it has an important operational consequence: **hitting the iteration cap is silent unless you look**. A task that returns a plausible paragraph may in fact be an agent that gave up. Lowering `max_iter` for a well-scoped agent (five or six is plenty for a single-tool specialist) turns a runaway into a fast, visibly-poor result you can catch in evaluation. ## max_execution_time `max_execution_time` bounds the agent's execution in seconds. This is the cap that matters when the problem is tool latency rather than iteration count: three calls to a slow scraper can outlast twenty calls to a fast lookup. In any crew that runs behind a request or a queue with its own deadline, set this — an unbounded agent inside a bounded caller is how a crew takes down the thing that invoked it. ## max_rpm `max_rpm` limits how many requests per minute that agent may issue. It is frequently misunderstood as a loop guard; it is not. It throttles the rate of the loop, so a runaway agent with a low `max_rpm` runs longer and just as far. Its real job is to keep an agent inside a provider's rate limit, especially when several agents share one API key. Note that it interacts with `max_execution_time`: throttling hard enough can make an agent hit its time bound before its iteration bound. ## cache Tool result caching is on by default (`cache=True`). Identical tool invocations reuse the previous result rather than calling out again. For a literal repeat loop this cuts the external cost sharply, and it makes the loop obvious in a trace — the same call, over and over, returning the same thing. It does not shorten the loop, because the model still spends a turn on each attempt. Individual tools can also declare their own caching policy, which matters when a tool's result is time-sensitive and must not be reused. ## step_callback `step_callback` on the `Agent` is a callable invoked after each step of that agent's execution. It receives the step output — a tool action when the agent decided to call a tool, and the finishing output at the end — which makes it the natural place to emit a structured log line per step, count iterations, push a metric, or detect that the same tool and arguments have appeared N times in a row. A crew-level callback exists too for blanket coverage; the agent-level one is what you reach for when a single agent is the suspect. Together with `verbose=True` during development, it turns "the crew hangs" into a readable timeline. ## respect_context_window Worth knowing in the same breath: `respect_context_window` (default on) keeps the agent's growing message history inside the model's context window rather than letting the request fail. A long loop is exactly what fills a context window, so this is often what stands between a runaway agent and a hard provider error — at the price of the agent quietly losing earlier detail. ## Containment is not a cure Caps buy you a bounded failure. They do not make the agent finish. The actual causes worth fixing are: a tool that returns an unhelpful empty or error result instead of a message the model can act on; a tool description that does not say what the tool cannot do; an agent whose goal has no observable stopping condition; and a missing instruction for the empty case ("if the search returns nothing, report that no results were found"). Fix those and the caps stop firing. ## How to answer this in an interview Name the three caps and be precise about what each bounds — iterations, seconds, request rate. Say explicitly that reaching `max_iter` produces a forced best answer rather than an exception, because that is the detail that shows you have watched it happen. Then close on the diagnosis: caps contain, tool output and stop conditions cure.

  • Why can an agent that hit max_iter still return a confident-looking answer?
    Because reaching the iteration cap is not an error path. CrewAI pushes the agent to produce its best final answer from whatever it has gathered, so the task completes and downstream steps consume the result normally. Detect it by logging steps through `step_callback` or by counting iterations, and treat a run that repeatedly reaches the cap as a quality regression rather than a tuning problem.
  • Why is max_rpm the wrong knob for stopping a runaway loop?
    `max_rpm` throttles the rate of requests, not their number. A looping agent under a tight rate limit performs the same number of iterations, just more slowly — often burning the wall-clock budget instead. Use `max_iter` for the number of steps and `max_execution_time` for latency; keep `max_rpm` for staying inside a provider's quota, especially when several agents share one key.
  • What does tool caching change about a repeated identical tool call?
    With `cache` enabled the second and subsequent identical invocations reuse the stored result instead of calling the tool again, which cuts external cost and makes the repetition obvious in a trace. It does not shorten the loop, since the model still spends a turn per attempt, and it is wrong for time-sensitive tools whose results must not be reused.
  • You bounded the agent and it now fails fast. What do you actually fix?
    The stopping condition. Usually the tool returns an empty or error result the model cannot act on, the tool description does not state its limits, or the agent's goal has no observable finish line. Give the tool a message the model can reason about, state the empty case explicitly in the backstory, and narrow the goal until "done" is checkable.

saying these in an interview costs you the question

  • Thinks max_iter raises an exception when reached
  • Uses max_rpm as protection against infinite loops
  • Leaves max_execution_time unset behind a request deadline
  • Believes caching shortens the loop rather than cheapening it
  • Raises the caps instead of fixing the tool output

context