After a supervisor decides to remediate a timed-out step by retrying it, why must the underlying operation the agent performs typically be idempotent, and what happens if it isn't?
answer
- timeout ambiguous: never-sent vs lost-ack
- idempotency key/request ID dedup
- burden shifts to remote service
- duplicate charge/shipment/email as symptom
- hard to reproduce - timing dependent
basics
~20 sBecause a timeout doesn't guarantee the original call failed - it might have already succeeded. Retrying a non-idempotent operation can make it run twice, like charging a customer or sending an email a second time.
solid answer
~50 sThe supervisor can't distinguish 'the call never reached the remote service' from 'the call succeeded but the acknowledgment was lost' - both look identical from the outside, a missed deadline. So its default remediation, retry, has to be safe even when the original attempt actually succeeded. If the agent's operation is idempotent - applying it twice has the same effect as applying it once, typically enforced via a unique operation/idempotency key that the remote service deduplicates on - a retry is safe regardless of what really happened. If it isn't idempotent, a retry after a false-positive timeout produces a real duplicate side effect: a second charge, a duplicate shipment, a second email. Making an operation idempotent usually means the remote service accepts a client-generated request ID and returns the original result for a repeat of that ID instead of re-executing the underlying effect.
go deeper
Should be able to say in plain terms that retrying after a timeout can accidentally do something twice, using an example like double-charging.
Should explain that a timeout can't distinguish 'failed' from 'succeeded but ack lost', and that this is why retries need idempotency.
Should describe the idempotency-key mechanism concretely (client-generated ID, service-side dedup) and identify which side of a call should own that check.
Should discuss the operational cost of idempotency (storage, retention, hot-path lookups), what to do when a downstream service doesn't support it, and connect it to real incidents like duplicate charges/shipments.
## Why idempotency is structural here, not decorative Idempotency is not a nice-to-have decoration on this pattern, it's a structural requirement created directly by the ambiguity of timeout-based failure detection. When the supervisor sees a step's deadline pass with no terminal status, that single observable fact - 'no acknowledgment arrived in time' - is consistent with several very different underlying realities: - the remote call never reached the service at all (network partition before the request); - the service received it but crashed before responding; - the service processed it successfully but the response was lost on the way back; - or the service is simply still working on a slow request. The supervisor has no way to tell these apart from where it sits, and building a perfectly reliable distributed detector that could tell them apart is fundamentally hard (this is essentially the same uncertainty behind the **two generals' problem**). So the pattern's designers make a pragmatic choice: treat a timeout as 'unknown, possibly failed' and default to retrying, but push the burden of safety onto the operation itself rather than onto the detection mechanism. ## How it is actually implemented 1. **The key is created once.** Concretely, idempotency in this context is usually implemented via a client-generated idempotency key or request ID that's created once, when the scheduler first dispatches the step, and reused on every retry of that step. 2. **The service checks that key before acting.** The remote service, on receiving a request, checks whether it has already processed that key; if so, it returns the previously computed result without re-executing the side effect (charging money, decrementing inventory, sending a message) a second time. This shifts deduplication to the side that can actually know for certain whether the effect already happened - the service holding the state - rather than the caller, which can never be certain from a timeout alone. Payment processors are the canonical example: Stripe's API, for instance, accepts an idempotency key header specifically so that callers can safely retry a charge request after an ambiguous timeout without risking a double charge. ## What it costs The trade-off is that idempotency isn't free: - it requires the remote service to persist a record of processed request IDs (often with a retention window, since keeping them forever is its own storage cost); - and it requires every write path that a retry could hit to check that record before acting, adding a lookup to the hot path of every request. For read-only or naturally idempotent operations (setting a value to a fixed state, as opposed to incrementing a counter) this cost is often near zero, but for operations with genuine external side effects - sending an email, calling a third-party API you don't control, moving physical inventory - you either need the downstream system to support idempotency keys itself, or you need to add your own deduplication layer in front of it, which is extra engineering effort that some teams skip under time pressure. ## The failure mode when a team skips it The failure mode when a team skips this is exactly the 'invisible until it isn't' kind that's expensive in production: everything looks fine in testing because timeouts rarely fire in a controlled environment, then under real network conditions - a brief partition, a slow downstream dependency, a garbage-collection pause on the remote service - a step legitimately times out after actually succeeding, the supervisor retries it per policy, and the operation runs twice: - For a payment this shows up as a customer service ticket and a chargeback. - For an inventory reservation it shows up as overselling. - For a notification it shows up as an annoyed user getting the same email twice. These bugs are notoriously hard to reproduce because they require the exact timing window where the original call succeeded just after the deadline check ran, so they tend to surface only at production scale and get initially misdiagnosed as 'random' data corruption rather than traced back to the retry policy. ## Where you have already seen it A well-known real-world illustration is how orchestration engines like Temporal handle this: each activity invocation carries a deterministic identity tied to the workflow's execution history, and activities are documented as requiring idempotency specifically because the engine's own retry policy, driven by the same timeout ambiguity described here, may invoke the same activity code more than once for a single logical step. The pattern's designers don't try to solve the detection ambiguity - they just document that the burden shifts to the activity author to make the operation safe to repeat.
- Where should the idempotency key be generated - by the scheduler or freshly by the agent on each retry attempt?By the scheduler (or whatever party owns the step's identity), once, at first dispatch, and then reused unchanged on every retry attempt for that step. If the agent generated a new key on each retry, the remote service would see each retry as a distinct request and the deduplication would never trigger, defeating the purpose entirely.
- Is idempotency only relevant for the retry remediation, or does it matter for reassignment to a different agent too?It matters equally for reassignment, since from the remote service's point of view a retry from a different agent instance looks the same as a retry from the same one - it's just another request carrying the same operation identity. The deduplication has to be keyed on the operation, not on which agent instance sent it.
- What if the remote service you're calling doesn't support idempotency keys at all, like a legacy third-party API?You typically add your own deduplication layer in front of it - recording, before each call, that you're about to attempt operation X, and checking that record before retrying so you can skip the call if a prior attempt is believed to have already gone through. It's strictly weaker than the service doing it itself (you're inferring rather than being told), but it's the standard workaround when you don't control the downstream system.
Like re-mailing a letter because you never got a reply - if the post office keeps a log of every unique tracking number it's already delivered, resending with the same tracking number just returns 'already delivered, here's your receipt' instead of the recipient getting two copies.
saying these in an interview costs you the question
- Says retries are always safe without mentioning idempotency
- Thinks idempotency means 'the operation is fast to repeat' rather than 'repeating it has the same effect as doing it once'
- Doesn't realize the idempotency key must stay the same across retries of one step
- Assumes this is only a payment-specific concern rather than general to any side-effecting step
- Can't explain why timeout-triggered retries specifically create this need