A refund tool times out after the payment succeeded — how do you stop a double refund?
answer
- timeout says nothing about the outcome
- the model retries too, not just the client
- code mints the key, never the model
- argument-only keys suppress real repeats
- persist the attempt before the call
basics
~20 sGive every side-effecting tool call an idempotency key minted by the harness, not by the model, and pass it downstream so a repeat is a no-op returning the original result. Then return an honest unknown-outcome result telling the model to check status before retrying.
solid answer
~60 sA timeout tells you nothing about whether the refund happened — an 8-second client deadline can expire long after the processor accepted. Two things then retry: your transport layer, and the model itself, which sees an inconclusive result on a later turn and re-issues the call. The second path is the one people miss, and it is why network-level retry logic alone does not protect you. The fix is that the tool wrapper mints a stable key when a call is admitted — derived from the session, the specific tool-call identifier and the canonicalised arguments — and passes it to the payment provider, so any repeat of *that* call collapses to a no-op that returns the original outcome. Never let the model generate the key: it produces a fresh value on the retry, defeating the mechanism. Scope the key to the tool-call identifier rather than the arguments alone, so a genuinely different second refund still goes through. Then make the timeout result explicit: outcome unknown, call the status tool before retrying.
code
python · 10 linesimport hashlib
import json
def idempotency_key(session_id, tool_call_id, tool_name, args):
payload = json.dumps(
{"tool": tool_name, "args": args}, sort_keys=True, separators=(",", ":")
)
digest = hashlib.sha256(payload.encode("utf-8")).hexdigest()[:16]
return f"{session_id}:{tool_call_id}:{digest}"go deeper
Be able to say that a timed-out call may still have succeeded, and that repeating it can charge or refund a customer twice unless the tool layer recognises the repeat.
Explain the mechanism: a key minted by your code and passed downstream so a repeat returns the original result, plus a timeout result that tells the model to check status rather than blindly retry.
Show that you cover both retry paths — transport retry and the model re-issuing the call turns later — scope the key so legitimate repeat actions still succeed, persist the attempt before the call for crash resume, and prefer reconciling on timeout over handing the model an ambiguity.
Own the classification: which tools are irreversible enough to require guarded wrappers and durable attempt logs, what your system defines as the same action, and how that definition is enforced so no tool reaches a downstream system outside the guard.
## A timeout is not a failure The most important thing to internalise is that a client-side timeout is a statement about your patience, not about the world. When an 8-second deadline expires on a refund call, the request may never have arrived, may be in flight, or may have been fully applied by the payment processor a moment before you gave up. The outcome is genuinely unknown, and any design that treats the timeout as "it did not happen" will eventually issue a second refund for a real customer. ## Two retry paths, not one Engineers reflexively reach for retry-with-backoff inside the tool. That handles the first path. But an agent has a second, stranger path: the *model* retries. The tool returns something inconclusive, the model reads it on the next turn, concludes that its goal is unmet, and emits the same call again — possibly several turns later, possibly with slightly different arguments, possibly after a compaction that dropped the earlier attempt from the visible history. From the downstream system's point of view, tool delivery in an agent is at-least-once, and you cannot fix that by making the HTTP client smarter. This is why the protection has to live at the boundary where side effects happen, not in a retry policy. ## Who mints the key The tool wrapper — the deterministic code that sits between the model's request and the downstream call — allocates the key at admission time and reuses it for every attempt at that logical action. Why not the model? Because a model asked to supply an idempotency key will generate a plausible new random-looking string each time it emits the call. On the retry it produces a *different* value, which is precisely the case the mechanism exists to defend against. It may also reuse a value it saw earlier for a genuinely different action, silently suppressing a legitimate second refund. Key generation is a job for code: it needs to be deterministic under repetition and distinct under genuine difference, which is a property models do not have. ## Scoping the key correctly The common error is deriving the key from the arguments alone. A hash of `{order_id, amount_cents, reason}` looks elegant and quietly breaks the legitimate case: two 6000-cent goodwill refunds on the same order, days apart, are different actions and must both go through. A key scoped only to arguments makes the second one a silent no-op — and the model will believe it succeeded. A workable scope is session plus the specific tool-call identifier the model's turn carries, plus a digest of the canonicalised arguments. The tool-call identifier makes each *decision to call* unique; the argument digest catches the case where a retry mutates arguments while claiming to be the same action. Canonicalise before hashing — sorted keys, no incidental whitespace — or key stability evaporates under trivial serialisation differences. Whatever scope you pick, write it down, because it is the exact statement of what your system considers "the same action", and everyone downstream is relying on it. ## What the timeout should return to the model Silence and false failure are both wrong. The result should say the outcome is unknown and give the recovery move: check the status tool for this order before calling the refund tool again. Better still, resolve the ambiguity inside the tool before returning — on timeout, reconcile against the provider by looking up the order's refund state, and return the real answer. Reconciliation is more work but it means the model is never handed an ambiguity to reason about, and reasoning about ambiguity is where agents behave worst. ## The durable attempt log Keys only help if they survive. Record the attempt — key, tool, arguments, and outcome once known — in durable storage before the downstream call goes out, not after. That log covers the case people forget: the harness process crashes mid-call and the session resumes from a checkpoint, replaying a turn whose side effect already fired. Without a persisted record, the resumed run mints a fresh key and repeats the action. With one, it finds the prior attempt and returns its outcome. ## What not to ask the model to do Do not rely on the model remembering that it already refunded this order. Context gets compacted, sessions get resumed, subagents get spawned with fresh windows. Any invariant that depends on the model's recollection of its own past actions will hold most of the time and fail exactly when the session is long and the stakes are high. Deduplication is a property of the tool layer. ## Where the boundary sits Not every tool needs this. Read-only calls are naturally safe to repeat, and cheap idempotent writes rarely justify the machinery. Reserve keys and durable attempt logs for the calls whose repetition is expensive or irreversible: money moving, messages sent to customers, tickets created in an external system, records deleted. Deciding which tools are in that set — and enforcing that every one of them goes through the guarded wrapper rather than calling out directly — is the part worth arguing about in a design review.
- Why can a hash of the arguments alone be the wrong idempotency key?Because two genuinely distinct actions can share arguments. A second 6000-cent goodwill refund on the same order, issued days later, is not a duplicate — but an argument-only key makes it a no-op, and the model is told it succeeded. Scope the key to the specific decision to call, typically session plus the turn's tool-call identifier plus an argument digest, so repetition of one call collapses while a new call does not.
- Your harness crashes and the session resumes from a checkpoint mid-tool-call. What prevents a repeat?A durable attempt record written before the downstream request goes out, holding the key, tool and arguments. On resume, the wrapper finds the prior attempt for that tool-call identifier and either returns its recorded outcome or reconciles against the provider. If the record is only written after a successful response, the resumed run mints a fresh key and fires the side effect a second time.
- Which tools actually need this machinery, and which do not?Reserve it for calls whose repetition is expensive or irreversible: money movement, customer-visible messages, external ticket creation, deletions. Read-only calls are safe to repeat by construction, and cheap idempotent writes rarely justify the storage and complexity. The design decision worth enforcing is that every tool in the guarded set goes through the wrapper rather than calling the downstream system directly.
It is the same discipline as a cash register that records the transaction before opening the drawer: if the power fails mid-sale, the record decides what already happened, not the cashier's memory.
saying these in an interview costs you the question
- Treats a timeout as proof the action did not happen
- Lets the model supply the idempotency key
- Relies on client retry logic while the model re-issues the call
- Derives the key from arguments alone, blocking legitimate repeats
- Assumes the model remembers it already performed the action