When a tool raises during a Haystack Agent run, what happens by default?
answer
- the exception does not escape by default
- the model is told and tries again
- one flag turns that into an abort
- retries cost a model call each
- fluent answer, failed tool underneath
basics
~20 sBy default the Agent does not propagate the exception. The error is captured and returned to the model as a tool result message, so the model can correct its arguments or pick another tool. Setting raise_on_tool_invocation_failure=True makes the run abort instead.
solid answer
~50 s`Agent` constructs its internal `ToolInvoker` with `raise_on_failure` taken from its own `raise_on_tool_invocation_failure` flag, which defaults to `False`. So when a tool raises — a bad argument, a timeout, a 500 from an upstream API — the invoker catches it, turns the error into a tool result message, and the loop continues with the model seeing what went wrong. That is usually right for recoverable, model-caused failures: a malformed argument gets fixed on the next step. It is wrong for failures the model cannot fix, because each retry costs a full model call over a transcript that keeps growing, and the model may eventually give up and assert success anyway. Set `raise_on_tool_invocation_failure=True` when a tool failure is a system fault you want surfaced, keep `max_agent_steps` low so retry loops are bounded, and never rely on the model to notice a failure — check the transcript for error tool messages before treating a run as clean.
code
python · 15 linesfrom haystack.components.agents import Agent
agent = Agent(
chat_generator=chat_generator,
tools=[billing_tool],
raise_on_tool_invocation_failure=True, # system faults abort the run
max_agent_steps=6,
)
result = agent.run(messages=[user_message])
degraded = any(
msg.tool_call_result is not None and msg.tool_call_result.error
for msg in result["messages"]
)go deeper
Know that a tool error does not crash the agent by default — the error text is sent back to the model as a tool result and the loop continues.
Explain the flag that changes this and why the forgiving default exists: model-caused argument errors are usually fixable on the next step.
Demonstrate the production instinct: retries cost a model call each, failures hide behind a confident answer, and you must scan the transcript for error tool messages and bound the loop.
Own the policy split — classify failures by who can fix them, push non-retryable faults out as exceptions with a real retry and circuit-breaking policy, and set organization-wide expectations for degraded-run detection.
## The default is forgiving, and that is a deliberate trade Most tool failures in an LLM agent are the model's own fault: it passed a string where an int was expected, invented a filter key, or called the search tool with an empty query. Aborting the run for those would be brutal, because the model is perfectly capable of fixing them if it is told. So Haystack's `Agent` defaults `raise_on_tool_invocation_failure` to `False`: the exception is caught by the internal `ToolInvoker`, converted into a tool result message carrying the error text, appended to the conversation, and the loop continues. From the model's perspective this is indistinguishable in shape from a successful call — a tool message came back — which is exactly why it works and exactly why it is dangerous. ## What the failure costs Each retry is a full round trip: one model call whose prompt now includes the previous attempt *and* the error. Three failed attempts against a downstream service that is simply down means three model calls, a longer transcript each time, and no progress. With `max_agent_steps` at its default of 100, a persistently failing tool can burn a startling number of tokens before anything stops it. This is the concrete reason the step cap matters even when you never expect to reach it. ## The failure mode nobody catches in review The model reads "HTTP 502 Bad Gateway" and, under pressure to produce an answer, writes a confident summary anyway. The run completes, `last_message` is fluent prose, and every automated check passes because nothing raised. Silent-success-after-failed-tool is one of the most common production agent bugs, and it is a direct consequence of the forgiving default. Mitigations, in order of value: 1. **Inspect the transcript.** After a run, scan the returned `messages` for tool results that carry errors, and mark the run degraded regardless of how good the answer reads. 2. **Say it in the system prompt.** Instruct the agent explicitly to report a tool failure rather than answer around it. Cheap, imperfect, worth doing. 3. **Fail fast where recovery is impossible.** If a tool's failure is always a system fault, `raise_on_tool_invocation_failure=True` turns it into an exception the caller handles — with a retry policy you control, at the level where retries are cheap. 4. **Handle the recoverable inside the tool.** A tool that catches its own timeout and returns a structured `{"error": "upstream timeout, do not retry"}` gives the model something actionable and keeps the exception channel for genuine faults. ## Choosing the flag The useful distinction is *who can fix it*. - Model-fixable (bad arguments, wrong tool, missing required field) → leave the default; feeding the error back is the whole point. - Not model-fixable (auth expired, service down, quota exhausted, misconfiguration) → raise, or convert to a terminal structured result. Letting the model retry a 401 is pure waste. Because the flag is agent-wide, a mixed toolset is best handled by having individual tools catch and classify their own errors, reserving raised exceptions for the terminal class. ## Related knobs The agent forwards invoker settings through `tool_invoker_kwargs`, which is where invoker-level behaviour such as result serialization and concurrency for parallel tool calls is configured. Note also that when a model requests several tool calls in one turn, one failing call does not cancel the others — you get a mix of successful and error tool messages in the same step, and your transcript check has to look at all of them, not just the last. ## Testing it Agents are worth testing against a deliberately failing tool: assert that the run terminates within the step budget, that the failure is visible in the returned messages, and that your wrapper marks the run degraded. That test catches the silent-success bug long before production does.
- Why is a fluent final answer not proof that the tools worked?Because a failed tool still produces a tool message — carrying an error string — and the model will often summarize around it rather than admit the gap. Nothing raises, so the run looks clean. The only reliable check is to scan the returned `messages` for error results and mark the run degraded independently of how good `last_message` reads.
- When would you set raise_on_tool_invocation_failure=True?When failure is a system fault the model cannot fix: expired credentials, a service that is down, exhausted quota, misconfiguration. Letting the model retry those wastes a model call per attempt and risks a confident wrong answer. Raising hands the problem to the caller, where a real retry or circuit-breaking policy lives.
- How do you keep a failing tool from consuming the whole step budget?Bound the loop with a low `max_agent_steps`, and make the tool itself classify errors: return a structured, explicitly terminal result for non-retryable failures so the model stops asking, and reserve raised exceptions for faults that should abort. Relying on the model to stop retrying on its own is not a control.
- If the model requests three tool calls in one turn and one fails, what happens?The others still run and return their results; the failing one contributes an error tool message alongside them. The loop continues with a mixed set of results in the same step, so any check for failures must examine every tool message from that turn rather than only the last one.
saying these in an interview costs you the question
- Assuming a tool exception propagates out of Agent.run by default
- Treating a fluent last_message as evidence the tools succeeded
- Letting the model retry unrecoverable failures like expired auth
- Leaving max_agent_steps high while a tool retries in a loop
- Thinking one failed call in a turn cancels the other tool calls