skip to content

Which tool failures should your wrapper retry in code, and which should the model see?

level: seniorimportance: should knowfreq 47%

answer

  1. who can act on this failure?
  2. transient and mechanical goes in code
  3. semantic and ambiguous goes to the model
  4. idempotency gates every retry
  5. tell the model what you already tried

basics

~20 s

Retry mechanical, transient failures in code — rate limits, connection resets, brief timeouts — with bounded backoff, and only when the call is safe to repeat. Surface anything the model could decide differently about: bad arguments, missing records, denied permissions, exhausted retries.

solid answer

~50 s

Split failures by who can act on them. A 429 from a paging provider or a dropped connection carries no information the model can use — it would just re-emit the identical call, spending a turn and its token budget to do what a loop in your wrapper does in milliseconds. Retry those in code with exponential backoff and jitter, bounded by attempts and by a deadline, and only for idempotent operations or ones you can make idempotent with a request key. Surface the failures the model can actually respond to: a malformed argument it can fix, a record that does not exist so it should search differently, a permission it must route around, or retries exhausted so it should stop trying this path. When you do surface, say what was already attempted — otherwise the model retries what your wrapper just retried three times.

code

python · 22 lines
python
import random, time

TRANSIENT = {429, 500, 502, 503, 504}

def call_tool(run, *, idempotent, attempts=3, deadline_s=8.0):
    started = time.monotonic()
    for attempt in range(attempts):
        status, body = run()
        if status == 200:
            return {"is_error": False, "content": body}
        if status not in TRANSIENT or not idempotent:
            return {"is_error": True, "content": f"failed with status {status}: {body}"}
        wait = min(2 ** attempt, 4) * (0.5 + random.random())
        if time.monotonic() - started + wait > deadline_s:
            break
        time.sleep(wait)
    return {"is_error": True,
            "content": f"upstream unavailable after {attempts} attempts; do not retry this call"}

calls = iter([(503, "restarting"), (200, "ok: 3 rows")])
print(call_tool(lambda: next(calls), idempotent=True))
print(call_tool(lambda: (400, "start_time must be ISO-8601"), idempotent=True))

go deeper

for a junior

Know that transient network and rate-limit failures are usually retried by the code around the tool, while mistakes in the arguments should come back to the model so it can fix them.

for a middle

Explain the split by who can act, and the mechanics of a safe retry: exponential backoff with jitter, bounded attempts, a deadline, and idempotency before repeating anything with a side effect.

for a senior

Demonstrate the operational judgment — total-time caps, circuit breakers, per-dependency concurrency limits, and timeout results that state whether a side effect may have landed so the agent verifies instead of duplicating.

for a principal

Own the policy across the tool fleet: one classification boundary per dependency, retry and deadline budgets that fit inside the agent's wall-clock and cost caps, and the failure-mode analysis for many agents retrying one struggling dependency at once.

## The question behind the question Every tool failure is handled somewhere. The only real decision is whether the handling happens in your code, silently, or in the model's next turn, visibly. The test that decides it: **could the model plausibly do something different if it knew?** If the answer is no, handling it in code is strictly better — it is faster, cheaper, deterministic, and it does not consume a turn or pollute the transcript. If the answer is yes, hiding it removes the agent's ability to adapt. ## What belongs in the wrapper Mechanical, transient faults. A rate-limit response from an upstream provider, a 503 from a service that is restarting, a connection reset, a DNS blip, a lock contention error. The model has no better move than to issue the same call again, so a retry loop in the wrapper does the job in milliseconds instead of a full round trip through the model. Do it properly: exponential backoff with jitter so a fleet of agents does not synchronize into a thundering herd; honour the wait hint the provider sends with a rate-limit response rather than guessing; cap both the number of attempts and total elapsed time, so retries cannot silently eat the agent's wall-clock budget. The hard constraint is idempotency. Retrying a read is free. Retrying "page the on-call engineer" or "issue a refund" can double the side effect, and a timeout is exactly the case where you do not know whether the first attempt landed. For side-effecting tools, either send a client-generated request key the upstream deduplicates on, or do not retry blindly — instead read back the state and decide, or surface the ambiguity to the model in words. ## What belongs in front of the model Semantic failures. The argument was the wrong shape or out of range; the identifier does not exist; the filter matched a field the schema does not have; the operation is not permitted for this principal; the requested resource is in a state that forbids the action. Each of these implies a different next move — reformat, search first, choose another tool, ask the user, escalate — and only the model has the task context to choose. Exhaustion also belongs here. When your bounded retry loop gives up, that is now a fact about the world: this dependency is unavailable right now. The model needs it to decide whether to degrade to a partial answer, try an alternative source, or stop. And it must be told what was already tried — "unavailable after 3 attempts over 8 seconds" — or it will spend its own turns re-running your retry policy by hand. ## Timeouts sit in both camps Every tool needs a deadline; a call with no timeout can hang an agent indefinitely, and the model sees nothing at all while it does — no result turn, no error, just an agent that appears stuck. So the wrapper enforces the deadline and always produces a result turn. Within the deadline, a slow-but-transient failure is a retry candidate like any other. Once the deadline fires, the timeout becomes information the model needs, and the wording matters more than for most errors, because a timeout is genuinely ambiguous: the work may have completed on the far side. A good timeout result says which tool timed out, after how long, and critically whether the side effect may have landed — "the escalation request was sent but the acknowledgement read timed out; status unknown". That sentence is what stops an agent from paging the same engineer twice. Where a status check exists, the better move is for the wrapper to perform it before reporting. Deadlines should also be tiered. A sub-second lookup and a multi-minute report build cannot share one timeout, and the per-tool deadline must fit inside the agent's overall wall-clock budget rather than exceeding it. ## Where retries actually hurt Retrying is not free even when it works. It consumes the agent's latency budget invisibly — a user watching a spinner cannot tell a retry loop from a hang — so long retry ladders belong behind a total-time cap, not just an attempt cap. Retries also multiply under concurrency: an agent making several parallel calls to the same struggling dependency turns a bounded policy into an unbounded load spike, which is why per-dependency concurrency limits and a circuit breaker matter more than the retry count. When the breaker is open, the correct behaviour is to fail fast to the model with an honest "this dependency is down", not to queue. ## A workable default Classify at the boundary. Map the upstream's status codes and exception types into two buckets — retryable-transient and terminal-semantic — once per tool, in one place, rather than scattering `except` blocks. Retry the first bucket up to a small attempt count under a total-time cap, with jitter and idempotency protection. Convert the second bucket, and any exhausted retry, into a flagged error result whose text names the cause, the attempts already made, and the state of any side effect. That gives you a fast path for noise and a visible path for judgment, which is the whole point of the split.

  • What changes when the failing tool has a side effect?
    Blind retries can duplicate it, and a timeout is the worst case because you do not know whether the first attempt landed. Use a client-generated request key the upstream deduplicates on, or read back state before retrying. If neither is possible, do not retry — return an error that states plainly that the side effect may or may not have occurred.
  • How should the result read when a tool exceeds its deadline?
    Name the tool, the deadline it exceeded, whether anything was retried, and above all whether the work may still have completed remotely. "submit_escalation timed out after 30s; the request was sent but unacknowledged, status unknown" lets the model verify before acting. A bare "timeout" invites a duplicate call.
  • Why cap total retry time and not just the attempt count?
    Attempts with exponential backoff can consume tens of seconds of the agent's wall clock while the user sees only a spinner, and retries across parallel calls compound. A total-time cap bounds the user-visible latency regardless of how the backoff ladder plays out, and pairs naturally with a circuit breaker that fails fast while a dependency is known to be down.
  • Is it ever right to let the model do the retrying?
    Yes, when the retry is not identical — a narrower query after a payload-too-large error, a different endpoint after a permission denial, a corrected argument after a validation failure. Those are decisions, not repetitions. Identical re-issues of the same call are the wrapper's job; anything requiring a changed call is the model's.

saying these in an interview costs you the question

  • Retries every failure, including invalid arguments
  • Retries a side-effecting call after a timeout with no idempotency key
  • Surfaces raw 429s so the model can decide to wait
  • Uses fixed-interval retries with no jitter or total-time cap
  • Reports exhaustion without saying anything was already retried

context