In LangChain, what happens when a tool raises or the model sends invalid args?
answer
- two failure paths, only one reaches your code
- invalid calls live in their own list
- a hook decides whether the exception escapes
- an error string is a prompt
- not every failure deserves a retry
basics
~20 sTwo different failures. Arguments that do not match the schema never reach your code: they land in AIMessage.invalid_tool_calls with an error. A tool that runs and raises propagates unless you handle it and return a ToolMessage with status set to error so the model can retry.
solid answer
~60 sSeparate the two paths. **Schema failures** happen before execution: the model produced arguments that failed to parse or validate, so LangChain records the call in `AIMessage.invalid_tool_calls` with an `error` string rather than in `tool_calls`. Nothing of yours ran. The fix is schema design — tighter types, enums instead of free strings, argument descriptions — not exception handling. **Execution failures** happen inside `_run`. By default the exception propagates and kills the run. `BaseTool` exposes `handle_tool_error`, which accepts `True`, a fixed string, or a callable taking the `ToolException` and returning a string; the returned text becomes the tool's output instead of an exception. Combined with a `ToolMessage` whose `status` is `"error"`, that turns a failure into feedback the model can act on. The judgment call is which errors are worth feeding back. A validation error the model can correct — yes. A 500 from a downstream service, an auth failure, a timeout — no; those are your problem, and letting the model retry them burns turns and money without ever succeeding.
code
python · 22 linesfrom langchain_core.tools import ToolException, StructuredTool
def _handle(err: ToolException) -> str:
return (
f"Tool failed: {err}. "
"Valid regions are eu-west-1 and us-east-1; retry with one of those."
)
def fetch_usage(region: str) -> str:
if region not in {"eu-west-1", "us-east-1"}:
raise ToolException(f"unknown region {region!r}")
return f"usage for {region}"
fetch_usage_tool = StructuredTool.from_function(
func=fetch_usage,
name="fetch_usage",
description="Fetch usage totals for a billing region.",
handle_tool_error=_handle,
)go deeper
Know that a tool raising an exception normally aborts the agent run, and that LangChain provides a way to turn that into text the model sees instead.
Distinguish invalid_tool_calls (schema failure, your code never ran) from an exception inside the tool, and explain handle_tool_error's three accepted forms.
Show the classification judgment: model-correctable errors get fed back with repair guidance, transient ones are retried inside the tool with backoff, terminal ones stop the path — and account for the turns each retry costs.
Own the contract across a tool catalogue: a house error taxonomy, scrubbing rules so nothing internal reaches the transcript, retry budgets per tool, and metrics that separate schema quality from dependency health.
## Three failure classes, three owners **1. Invalid arguments (model's fault, caught before execution).** The model emitted a tool call whose arguments do not parse as JSON or do not satisfy the `args_schema`. LangChain surfaces these separately: the call appears in `AIMessage.invalid_tool_calls` — each entry carrying `name`, the raw `args`, an `id` and an `error` — rather than in `tool_calls`. Your tool body never runs, so no amount of try/except inside the tool helps. The remedies are upstream: make the schema narrower so fewer wrong shapes are expressible, use `Literal` or enums for closed vocabularies, add per-field descriptions via `parse_docstring=True` or `Field(description=...)`, and reduce the number of near-identical tools competing for the same call. Repeated invalid calls for one tool are a schema smell, not a model defect. **2. Tool raised (your code's fault, or the world's).** The tool ran and threw. Default behaviour is propagation: the exception escapes the agent run. That is the correct default — silent swallowing would let an agent confidently answer from a failed lookup. `BaseTool.handle_tool_error` changes it: - `True` — the exception message becomes the tool's returned content. - a string — that fixed string is returned, useful when the raw exception would leak internals. - a callable — receives the `ToolException` and returns the string to show the model, which is where you write a *repair instruction* rather than a stack trace. `ToolException` from `langchain_core.tools` is the type intended to be raised for "this is an expected, model-visible failure". There is a matching `handle_validation_error` for input-validation failures raised at invocation time. **3. Wrong-but-successful results (nobody's exception).** The tool returned 200 OK and an empty list, or a plausible answer for the wrong entity. No error exists anywhere, and the agent proceeds confidently. This is the most damaging class and the only defence is the tool's own output design: say "no orders found for ORD-1234" rather than returning `[]`, so the model can distinguish absence from failure. ## Making an error message useful to a model An error string is a prompt. Compare "KeyError: 'region'" with "Missing required argument 'region'. Valid values: eu-west-1, us-east-1." The second reliably produces a corrected retry; the first produces a guess. Good model-facing errors state what was wrong, what the valid space is, and whether retrying is worth it. They must also be scrubbed — connection strings, internal hostnames and stack traces in an error string end up in the conversation, in your logs, and potentially in the user-visible answer. ## Which errors deserve a retry Feeding every error back to the model is a common and expensive mistake. Classify: - **Model-correctable** — bad enum value, malformed identifier, missing argument. Feed back with guidance; the model fixes it in one turn. - **Transient infrastructure** — timeout, 429, connection reset. Retry *inside* the tool with backoff, where you control the policy. A model-driven retry has no jitter, no budget, and consumes a full turn each time. - **Terminal** — 401/403, feature disabled, resource genuinely absent. Tell the model plainly that this path is closed so it stops trying, or abort the run. The failure mode to avoid is an agent burning its entire step budget re-calling an endpoint that will never authorise. ## Interaction with the step budget Every fed-back error costs a full turn: one model call to see the error, one to retry. Combined with a graph recursion limit counting super-steps, a tool that fails three times can consume half a default budget before real work starts. Cap per-tool retries explicitly rather than relying on the global limit to eventually stop the bleeding. ## Parallel calls complicate it When one `AIMessage` carries several tool calls and one fails, the others still completed. Each still needs its own `ToolMessage` — providers validate that every call is answered — so the failed one gets an error-status result while the successes get real ones. The model then decides whether the partial set is enough. Silently dropping the failed call's result is the bug that produces a provider 400 rather than a graceful degradation. ## Observability Structure the signal: count invalid tool calls per tool name (schema quality), tool exceptions per tool name (dependency health), and turns-consumed-by-retry per run (budget waste). Those three series tell you whether to fix a schema, a dependency, or a prompt, and they are far more actionable than a single agent-error rate.
- Why is retrying a 429 by feeding it back to the model a bad idea?The model has no backoff, no jitter and no budget awareness, so it retries immediately and each attempt costs a full turn plus the tokens of the whole re-sent conversation. Rate limiting is a transport concern: handle it inside the tool with proper backoff and a bounded retry count, and only surface it to the model as terminal once your own policy is exhausted.
- A tool keeps receiving arguments that fail validation. Where do you look?At the schema and the description, not at error handling — invalid calls never reach the tool body. Narrow the types (Literal or enum instead of str), add per-argument descriptions so the model knows the valid space, and check whether a sibling tool has an overlapping description causing the model to aim at the wrong target. Track invalid-call counts per tool name as a schema-quality metric.
- One of three parallel tool calls fails. What must the agent still do?Produce a ToolMessage for every one of the three calls, including the failed one — marked with status error and a useful message. Providers validate that each tool call is answered, so dropping the failed result causes a hard request rejection rather than graceful degradation. The model then judges whether two successful results are enough to answer.
saying these in an interview costs you the question
- Try/except inside the tool to catch schema validation failures
- Feeding raw stack traces back into the conversation
- Retrying auth failures until the step budget is exhausted
- Returning an empty list where 'not found' was the real answer
- Dropping the result of a failed call in a parallel batch