skip to content

In CrewAI, what happens to a crew run when a tool's _run raises an exception?

level: seniorimportance: should knowfreq 46%

answer

  1. the run does not die
  2. failures become context, not crashes
  3. schema rejects bad arguments first
  4. error strings are prompt surface
  5. bound iterations, don't cache failures

basics

~20 s

CrewAI's tool-usage layer catches the exception and hands the error back to the agent as that tool call's observation, so the kickoff continues and the LLM can retry or change course. The cost is extra model turns, so bound it and return short, actionable error text yourself.

solid answer

~60 s

An exception inside `_run` does not normally abort `kickoff()`. CrewAI wraps the invocation, converts the failure into an error observation, and appends it to the agent's context — the LLM then decides what to do next, which usually means retrying with different arguments or trying another tool. Argument problems are caught earlier: CrewAI validates the model's arguments against the tool's `args_schema` before `_run` is entered, so a missing or mistyped field comes back as a validation message the model can correct. That recovery loop is a feature until it isn't: each failed call costs a full model round trip plus tokens, and an agent can burn its whole iteration budget re-trying a tool that will never succeed. So bound it with the agent's `max_iter`, and shape the errors: catch expected failures inside `_run` and return a short string that tells the model what to do ("no results for that ISO date; try a date within the last 90 days") rather than letting a stack trace or a raw HTTP body land in the prompt.

go deeper

for a junior

Know that a tool error does not crash the crew: CrewAI turns it into an observation the agent reads, and the agent can try again or take another route.

for a middle

Distinguish the two layers — args_schema validation rejecting bad arguments before _run, and exceptions inside _run being caught and surfaced — and explain why returning a short, shaped error string beats raising.

for a senior

Treat the error text as an interface: retry transients in code with bounded backoff, surface only decisions to the model, keep the message short because it is billed every turn, and bound the loop with the agent's iteration limit.

for a principal

Own the failure contract across the tool platform: a shared taxonomy of retryable versus model-visible failures, mandatory logging and per-class metrics because a green kickoff hides tool errors, and a policy that failures are never cached.

## The default behaviour CrewAI executes a tool through a usage layer, not by calling your function directly from the agent loop. That layer is what converts a Python exception into something an LLM can act on. In practice: the exception is caught, an error observation is produced, it is appended to the agent's message history as the result of that tool call, and the loop continues. The crew keeps running. This is deliberate — an agent framework whose runs died on the first 429 would be useless — but it means **your tool's failures become prompt content**, and that reframes how you should write them. ## Two layers of failure ### 1. Argument validation, before your code CrewAI validates the arguments the model produced against the tool's `args_schema`. A missing required field, a string where an int was declared, an unparseable date — these are rejected before `_run` executes, and the validation message goes back to the model. This layer is free error handling and it is the reason a tight Pydantic schema with per-field `Field(description=...)` is worth writing: it converts "the model guessed the format wrong" from a crash inside your code into a self-correcting turn. ### 2. Execution failures, inside your code Anything your `_run` body raises — a timeout, an auth failure, a KeyError on an unexpected response shape — is caught and surfaced to the model. What the model sees is derived from the exception, so an unhandled `KeyError: 'items'` becomes context the LLM has to interpret. It has no idea what your response schema is. It will often respond by retrying identically, which is the worst case. ## Why raising is the lazy option The error text is prompt engineering under a different name. Compare: - Raised, unhandled: a traceback fragment or `HTTPError: 500 Server Error for url ...`. The model learns almost nothing actionable and tends to repeat the call. - Caught and shaped: `"upstream search is unavailable; do not retry this call — proceed using the notes already gathered"`. The model has an instruction it can follow. So the discipline is: catch the failures you expect inside `_run`, classify them, and return a concise string that names the cause *and the next action*. Reserve raising for genuinely unexpected states. Also keep the text short — a 4 KB error body in context is billed on every subsequent turn of that agent's loop, and it crowds out the actual work. ## Retry, and who owns it Some failures should never reach the model at all. A connection reset, a 429 with a `Retry-After`, a transient DNS blip: retry those *inside* `_run` with a bounded backoff, because a code-level retry costs milliseconds while a model-level retry costs a full round trip and tokens. Failures that need a *decision* — no results, ambiguous input, permission denied — belong to the model, because the model is the thing that can choose a different approach. ## Bounding the loop An agent that cannot make progress will keep trying. The bounds available are the agent's iteration limit (`max_iter`) and its rate limit (`max_rpm`), and CrewAI additionally discourages an agent from repeating an identical failing call rather than letting it spin on the same input forever. Set `max_iter` deliberately for agents holding fragile tools; the default is generous, and "generous" times "a tool that is down" equals a large bill and a task that eventually returns an apology instead of an answer. ## Caching interacts badly here Crew-level tool caching is on by default and keys on tool plus arguments. If a failure result gets cached, the model's retry is answered from the cache and never re-executes — so the agent sees the same error no matter what, and concludes the capability is dead. Give any tool that can fail a `cache_function` that returns False for error and empty results. ## Observability Because failures are absorbed into the conversation, they are invisible in a naive "did the crew succeed" check: `kickoff()` returns a `CrewOutput` and the run looks green while a tool failed forty times. Log inside `_run` — before the return, on every path — and emit a metric per failure class. When you are debugging, `verbose=True` on the agent or crew shows the tool calls and their observations in the console, which is usually the fastest way to see that the model has been staring at the same error for six iterations. ## The interview answer Lead with "the run continues, the error becomes an observation". Then make the senior point: because the error is prompt content, error text is an interface you design; retry transients in code and escalate decisions to the model; bound iterations; and don't let a failure land in the cache.

  • Which tool failures should you retry inside _run, and which should the model see?
    Retry transients in code — connection resets, timeouts, 429s with a retry hint — because a code retry costs milliseconds while a model-visible retry costs a full round trip and tokens. Surface anything requiring a decision: no results, ambiguous input, permission denied, a resource that does not exist. Those are cases where the right next step is a different query or a different tool, and only the model can choose.
  • Why can a failing tool make a crew run look successful?
    Because CrewAI absorbs tool errors into the agent's context rather than propagating them, `kickoff()` still returns a `CrewOutput` and nothing raises. The agent simply produces a weaker answer. You need logging inside `_run` on every path and a metric per failure class; a green run is not evidence the tools worked.
  • How does default tool caching worsen a failing tool?
    Caching is on by default and keys on tool plus arguments, so a cached error result is replayed for every identical retry within the run — your code never executes again. The agent sees an unchanging failure and gives up. Assign a `cache_function` that returns False for error or empty results so retries stay real.

saying these in an interview costs you the question

  • Says an exception in a tool aborts the whole kickoff
  • Lets raw stack traces or full HTTP bodies reach the model's context
  • Retries every failure at the model level instead of in code
  • Leaves max_iter at the default for agents with fragile tools
  • Assumes a green kickoff means every tool call succeeded

context