skip to content

When a tool call fails, why return an is_error tool result rather than letting the exception propagate?

level: middleimportance: must knowfreq 68%

answer

  1. exception ends the turn
  2. error result keeps the loop alive
  3. flag says failure, not data
  4. the message decides the recovery
  5. harness owns what the model cannot fix

basics

~20 s

An exception ends the agent's turn; a tool result flagged as an error keeps the loop alive and hands the model a fact it can act on. The model can then fix an argument, pick another tool, or report honestly that the step failed.

solid answer

~50 s

The agent loop only continues while the conversation keeps growing, and the model can only react to things that appear in it. If your wrapper lets a tool's exception escape, the turn dies at the harness level: the model never learns the call failed, and whatever recovery you wanted has to be coded in the harness instead. Returning the failure as a tool result — with the error flag set, so the model is not misled into treating the text as data — keeps the call paired and gives the model the material to replan: bad argument, missing permission, no such record. The flag matters because a failure serialized as an ordinary success can be read as content; the text matters more, because the model can only recover from an error it understands. Errors it cannot possibly act on are your wrapper's problem, not the model's.

code

json · 6 lines
json
{
  "type": "tool_result",
  "tool_use_id": "call_7f2a",
  "is_error": true,
  "content": "search_logs failed: start_time must be an ISO-8601 timestamp, got \"yesterday 3pm\". No query was executed; retry with e.g. 2026-04-11T15:00:00Z."
}

go deeper

for a junior

Be able to say that a failed tool still owes the model a result, and that the result should be marked as an error and explain in plain words what went wrong.

for a middle

Explain why propagating the exception removes the model from the loop, what the error flag disambiguates, and which parts of an error message actually drive recovery.

for a senior

Show where you draw the line: model-actionable failures go back as errors, harness bugs and policy violations abort. Talk about redaction, retry guidance, and stating the state of half-completed side effects.

for a principal

Own error taxonomy across the tool fleet — consistent shapes, a policy for what is surfaced versus terminal, per-tool error-rate telemetry, and loop guards so accumulated failures do not silently degrade every long run.

## Two places a tool failure can go When the code behind a tool raises, the harness has exactly two choices. It can let the exception propagate — the agent run aborts, the caller gets a stack trace, and the conversation stops mid-turn with an unanswered tool call in it. Or it can catch the failure, serialize it into the tool result that the pending call is waiting for, and mark that result as an error so the model knows the content describes a failure rather than data. Those are not two styles of the same thing. The first removes the model from the loop; the second keeps it in. Everything else follows from that. ## Why keeping the model in the loop is usually right A large share of tool failures are things the model caused and can fix. It passed a date in the wrong format, referenced a record that does not exist, called the read-only tool for a write, filtered on a field the schema does not have. Given a clear error, the model corrects the argument and calls again — recovery that would otherwise have to be hand-coded in the harness for every tool and every failure shape. A second class of failure is not fixable but is still worth knowing about. The billing API is down, the permission was denied, the query timed out. The model cannot repair those, but it can route around them: try a different source, degrade to a partial answer, or tell the user plainly that the step failed. The alternative — a run that dies silently — produces the worst outcome in an agent product, which is a user who never finds out that half the work did not happen. ## What the error flag buys you Marking the result as an error is a small, load-bearing detail. Without it, the failure text sits in the transcript looking exactly like ordinary tool output, and models will sometimes read `Error: user 8831 not found` as a fact about user 8831 rather than as a failed lookup. The flag disambiguates that: this turn is a failure report about the call, not content returned by it. It is also the signal your own telemetry keys off — error rate per tool is one of the most useful agent metrics you can collect, and it is free if failures are tagged consistently. ## The message is what actually determines recovery A flagged result whose body is `Internal error` is almost as useless as a swallowed exception. Useful error text does three things. It names what was wrong in the model's own vocabulary — the argument, the value, the constraint. `start_date must be YYYY-MM-DD; got "last Tuesday"` produces a corrected call on the next turn. `ValidationError at line 214` does not. It says whether retrying could plausibly help. "The upstream service is unavailable; the wrapper already retried three times" stops the model from burning its remaining turns on the same call. Without that sentence, a model faced with an unexplained failure will very often just try again. It says what state the world is in, especially after a partial or side-effecting call. "The transfer request was submitted but the confirmation read timed out; status is unknown" is the difference between an agent that checks and one that duplicates the transfer. What error text must not do is leak internals the model has no business seeing — stack traces, SQL with credentials, internal hostnames, raw tokens. That material adds no recovery value and now lives in the transcript. ## When propagating really is correct Not every failure belongs in front of the model. Programmer errors in the harness — a tool wired up wrong, a serialization bug, a misconfigured client — should fail loudly at development time rather than being smoothed into a polite message the model apologizes about. Failures that violate a hard safety or policy boundary should terminate the run, not invite the model to find another route. And infrastructure faults that make the whole run meaningless — the model provider itself is down, the sandbox died — are the harness's to handle. The useful rule is that the model should see failures it could plausibly act on, and the harness should own everything else. "Could the model do something different because of this?" is the question that decides. ## Consequences downstream Because the failure lands in the conversation, it stays in the conversation. Repeated flagged errors accumulate and can crowd the window and bias the model toward pessimism; a run that has seen four failures in a row often gives up early. That argues for terse errors, for collapsing repeated identical failures, and for a loop-level guard that stops after N consecutive failed calls rather than letting the model rediscover the same wall on every turn.

  • What belongs in the error text, and what must never appear in it?
    Include the offending argument and constraint, whether a retry could help, and the state of any side effect that may have half-landed. Leave out stack traces, raw SQL, internal hostnames, credentials and provider tokens — they add nothing the model can act on and they persist in the transcript and in your logs.
  • When should a tool failure abort the run instead of going back to the model?
    When the model could not act differently on it, or should not be given the chance: harness bugs and misconfiguration that should fail loudly in development, violations of a hard safety or policy boundary, and infrastructure faults that invalidate the whole run. Everything the model could plausibly repair or route around goes back as an error result.
  • What happens when the same tool keeps failing turn after turn?
    The flagged errors accumulate in the window, crowding out signal and pushing the model toward premature give-up or repetitive retries. Keep error text terse, collapse repeated identical failures, and put a loop-level guard on consecutive failures so the harness stops rather than letting the model rediscover the same wall.
  • Does an empty or zero-match result count as an error?
    No. A query that ran correctly and matched nothing is a successful call with an informative result — say so in words, naming the filter and the zero count. Flagging it as an error tells the model the tool is unreliable and triggers retries or workarounds for a fact that is simply true.

saying these in an interview costs you the question

  • Says exceptions are fine because the framework will retry the whole run
  • Returns Internal error as the message and calls that error handling
  • Flags legitimate empty results as errors
  • Pastes stack traces and SQL into the result the model reads
  • Believes the model cannot recover from any tool failure

context