An agent repeats a failing tool call every turn — what should the error result say?
answer
- the error is a prompt, not a log line
- nothing to change means repeat the call
- say whether a side effect happened
- name the allowed values inline
- cap attempts outside the model
basics
~20 sAn error fed back to a model is a prompt. It must say what failed, what the current state actually is, whether a side effect happened, and what to do instead. A bare status code or stack trace gives the model nothing to change, so it repeats the identical call.
solid answer
~50 sRepetition is usually a symptom of an uninformative error payload rather than a stubborn model. `Error: 500` contains no signal about what to vary, so the highest-probability next action is the same call again. Write the result as instruction. Three classes need different wording: **fix-and-retry** ("reason 'EVENT_POSTPONED' is not allowed; use EVENT_CANCELLED, VENUE_CHANGE or GOODWILL"), **already-satisfied** (`{"error":"order 8812 already refunded on 2026-03-04; no action taken"}` — the goal is met, stop calling), and **terminal** ("the payment processor is unavailable; tell the customer we will follow up"). Always state whether a side effect occurred, because the model cannot infer it. Keep payloads short and free of stack traces, internal hostnames and SQL — they burn context and leak internals without helping. Finally, the error text is a nudge, not a control: cap attempts in the harness so a bad message cannot cost you an unbounded loop.
code
json · 6 lines{
"status": "already_applied",
"message": "order 8812 was already refunded 12000 cents on 2026-03-04; no action taken",
"retryable": false,
"next": "tell the customer the refund is already processed; do not call issue_refund again"
}go deeper
Know that whatever a tool returns on failure becomes part of the conversation the model reads next, so a message that does not say what was wrong will simply produce the same call again.
Explain the four things a failed result must carry — what failed, current state, whether a side effect happened, what to do next — and give an example of rewriting a generic failure into an actionable one.
Show that you separate fix-and-retry, already-satisfied and terminal failures, that you mark retryability explicitly, keep internals out of the payload, and back the wording with an attempt cap so a bad message cannot cost unbounded turns.
Treat error payloads as a contract across every tool in the platform: one shape, one retryability convention, recovery rate measured per failure class, and a standard for what may never appear in a model-visible payload.
## Why the loop happens A model chooses its next action from the conversation as written. When the previous turn ends with `Error: 500` or an exception class name, the conversation contains no information about what was wrong with the call — so the most probable continuation is the call that seemed right a moment ago, unchanged. The model is not being stubborn; it has been given nothing to condition a different choice on. Uninformative error payloads are one of the most common causes of the repeating-tool-call pathology, and one of the cheapest to fix. The reframe that fixes it: **an error result is a prompt**. It is text you are writing into the model's context to steer its next decision, and it should be judged the way you judge any other instruction — is it specific, is it actionable, does it say what to do next. ## The four things an error must carry **What failed, specifically.** Not "invalid request" but which argument, with the value that was rejected. Naming the offending value matters because the model has to locate it among several arguments it produced. **What the current state is.** The model's model of the world is stale. If the refund already went through yesterday, say so with the date and amount. State beats diagnosis: telling it the world is already how it wanted it is far more useful than telling it a constraint was violated. **Whether a side effect occurred.** This is the field teams forget. "Failed" is ambiguous between "nothing happened" and "something happened and then the report failed". The model cannot distinguish them and will guess, which is how duplicate actions get initiated. Say explicitly: *no action taken*, or *the refund was created but the confirmation email failed*. **What to do next.** The valid enum values. The tool to call first. Or the instruction to stop and tell the user. A model given an explicit next step follows it far more reliably than one left to infer a recovery path. ## Three classes, three shapes **Fix-and-retry.** The call was wrong and a corrected version will succeed. Echo the constraint and the legal values. This is the only class where you actually want another attempt, and it is where a well-written message pays off immediately — a rejected enum with the allowed set inlined is typically corrected on the very next turn. **Already-satisfied.** The goal is met; the call is redundant. The failure word is misleading here, because the right behaviour is to stop and report success to the user, not to retry. Say plainly that no action was taken and that the desired state already holds. **Terminal.** Nothing the model can do will make this call succeed — the processor is down, the record is frozen, the customer is not entitled. Say it is not retryable and give the model the user-facing move. Otherwise it will grind through alternative phrasings of the same doomed call until a budget stops it. Marking retryability explicitly — a boolean or a status word the model can read — is worth the few tokens, because "retryable" is precisely the judgment the model is bad at making from prose alone. ## What to keep out Stack traces, framework class names, internal hostnames, raw SQL and connection strings do not help the model and do actively hurt: they consume context, they can end up quoted verbatim to the user, and they teach nothing about what to change. Truncate long provider errors to the part that carries the actionable claim. Keep the payload to a couple of lines — an error is competing for the same context budget as the actual work, and a verbose one degrades everything downstream in a long session. ## Consistency helps more than eloquence Use one error shape across every tool: a machine-readable status, a short human-readable message, an explicit retryable flag, and a suggested next step. A model generalises across a consistent shape much better than across a dozen bespoke formats, and consistency also makes the payloads greppable in your traces, so you can count which failure classes actually occur. ## The error text is not a control Even a perfect message is a probabilistic nudge. Some fraction of the time the model will retry a terminal error anyway. So enforce the limit outside the model: cap attempts per tool per task, detect repeated identical calls and short-circuit them with a stronger message or a hard stop, and make sure exhaustion produces a useful partial result rather than silence. Good error text lowers the rate at which you hit the cap; it does not replace the cap. ## How to tell if yours are good Instrument recovery rate: of the calls that returned each error class, what fraction were followed by a *different, valid* call within one turn? A class with a low recovery rate is a message-writing bug, not a model failure. That metric turns error-payload design from taste into something you can iterate on with evidence.
- Why must a failed tool result say explicitly whether a side effect occurred?Because "failed" is ambiguous between nothing happened and something happened before the report failed, and the model has no other way to tell. Left to guess, it will often assume nothing happened and try again, which is exactly how duplicate refunds and duplicate emails originate. An explicit phrase — no action taken, or the charge was created but confirmation failed — removes the guess and lets the model choose between retrying and reporting.
- Which failure class is most often mislabelled, and what does that cost you?Already-satisfied. Teams surface it as a generic error, so the model reads it as something to work around and keeps trying variations, burning turns and budget on a goal that is already met. Labelling it as a state statement — the refund exists, no action taken — flips the model's behaviour from retrying to reporting success, which is what the user needed anyway.
- How would you measure whether your error payloads are actually working?Track recovery rate per error class: of the calls returning that class, what share were followed within one turn by a different, valid call, or by an appropriate stop. A class with a low recovery rate points at the wording, not the model. Pair it with a count of repeated-identical-call events, which tells you where a message carries no actionable signal at all.
saying these in an interview costs you the question
- Returns a raw stack trace or bare status code to the model
- Leaves it ambiguous whether the side effect already happened
- Labels an already-completed action as a generic failure
- Expects prompt wording alone to stop a retry loop
- Uses a different error shape for every tool