In a saga, a compensating transaction for a shipping step (e.g., 'cancel shipment') times out and the orchestrator retries it. What could go wrong if the compensating transaction isn't idempotent, and what design measures make saga steps and their compensations safe to retry?
answer
- at-least-once delivery, not exactly-once
- idempotency key per step invocation
- timeouts to avoid stalling forever
- durable saga log for crash recovery
- dead-letter queue for stuck compensations
basics
~10 sIf a retried step or compensation isn't idempotent, running it twice can double-charge, double-refund, or double-cancel something. Saga steps need unique operation IDs so duplicates can be detected and safely ignored.
solid answer
~60 sSagas rely on retries at multiple levels - message delivery, orchestrator crash-recovery, network timeouts - and most messaging systems only guarantee at-least-once delivery, meaning a step or compensation can execute more than once. If 'cancel shipment' isn't idempotent, a retried call after a timeout could trigger duplicate side effects (e.g., double refund of a shipping fee, duplicate warehouse notifications). The standard fix is to attach a unique, stable idempotency key to each saga step invocation (often derived from the saga instance ID plus the step name), have the receiving service persist which keys it has already processed, and short-circuit duplicate requests by returning the previously recorded result instead of re-executing the side effect. Beyond idempotency, robust saga failure handling also needs: timeouts on each step so a hung participant doesn't stall the saga forever, a durable saga log so a crashed orchestrator can resume from its last known state, and an escalation path (dead-letter queue plus alerting) for compensations that repeatedly fail, since an un-compensated saga leaves real inconsistency that can't always be resolved automatically.
go deeper
Should recognize that things can fail and get retried, and have a rough sense that retrying isn't always harmless.
Should explain idempotency keys at a basic level and know that timeouts are needed to detect a stuck step.
Should describe the full mechanism (key generation/storage location, durable saga log, dead-letter escalation) and reason about realistic failure scenarios like duplicate delivery.
Should discuss the systemic implications - how idempotency-key storage itself needs a retention/cleanup policy, and how to design SLOs/alerting around stuck sagas at scale.
## What separates a whiteboard saga from a production one Idempotency and failure handling are what separate a saga that works on a whiteboard from one that survives real production traffic, because the network, the messaging layer, and the orchestrator itself are all unreliable, and a saga has to remain correct despite that. ## At-least-once delivery, not exactly-once The starting point is that most of the machinery driving a saga - message brokers, HTTP retries, orchestrator crash-recovery - offers **at-least-once** delivery rather than **exactly-once**. That's a deliberate, practical choice: guaranteeing exactly-once delivery across a network is prohibitively expensive or impossible in the general case, so systems instead guarantee 'at least once' and push the responsibility for handling duplicates onto the receiver. Concretely, this means every step in a saga, and every compensating transaction, must be prepared to receive the same command or event more than once and produce the same result each time, rather than repeating its side effect. This property is called **idempotency**, and it is not automatic - a naive 'charge $50' or 'release 1 unit of stock' handler is not idempotent by default; calling it twice really does charge twice or release twice. ## The idempotency key The standard mechanism for making a step idempotent is an **idempotency key**: a unique identifier for this specific invocation of this specific step, commonly composed from the saga instance's ID plus the step name (e.g., `saga-482-charge-payment`). The receiving service, before executing the side effect, checks whether it has already processed a request with this key. 1. **If yes**, it doesn't repeat the side effect - it returns the previously computed result (success or failure) so the caller sees a consistent outcome. 2. **If no**, it executes the operation and durably records the key alongside the outcome, typically in the same local transaction that performs the side effect, so the record-keeping and the effect can't get out of sync even if the process crashes between them. This pattern applies identically to compensating transactions - 'refund `saga-482-charge-payment`' needs its own idempotency key so a retried refund doesn't double-credit the customer. ## The other mechanisms a production saga needs Beyond idempotency, a production-grade saga needs a few more mechanisms to handle failure robustly. - **A timeout on every step.** If a participant doesn't respond within a bounded window, the saga (orchestrator or the choreography participant waiting on an event) must treat that as a failure and begin compensating rather than waiting indefinitely, since an indefinitely stalled saga holds the business operation (and often a customer-visible order or booking) in limbo. - **Durable persistence of saga progress.** The orchestrator (or, in choreography, each participant's own state) needs it - which steps have committed - so a crash doesn't lose track of where the saga was; on restart, the orchestrator reads this state and resumes rather than either re-running already-completed steps or abandoning the saga silently. - **An escalation path.** Because compensations themselves can fail (a refund API might be down, a third-party service might reject a cancellation), the system needs one: aggressive retry with exponential backoff first, and if that's exhausted, routing the failed compensation to a dead-letter queue with alerting so a human can intervene, because a saga that's 'stuck' half-compensated is a real data-inconsistency incident, not something that resolves itself. ## A failure mode that ties it together A concrete failure mode that ties all of this together: an orchestrator sends 'cancel shipment' to the Shipping service, the request is processed successfully, but the response is lost on the way back due to a network blip, so the orchestrator sees a timeout and retries. If the Shipping service's cancel handler isn't idempotent, this retry runs the cancellation logic a second time - which might, depending on implementation, issue a second refund of the shipping fee, send a duplicate 'your shipment was cancelled' email, or throw an unexpected error because the shipment is already in a 'cancelled' state that the handler didn't anticipate seeing twice. Any of these either directly costs money (double refund) or damages trust (confusing duplicate notifications) or breaks the saga's own failure-handling logic (an unexpected error on retry can itself cause the orchestrator to treat the compensation as failed and escalate incorrectly). This is why, in practice, teams treat 'is this step idempotent, and is that idempotency tested with a duplicate-request test case' as a required checklist item for every saga step and compensation before it ships - it's the single most common class of production bug in saga implementations, precisely because it's easy to build the happy path and easy to forget that at-least-once delivery is not a hypothetical.
- Where should the idempotency key be generated and checked?The caller (orchestrator, or the publishing participant in choreography) generates a stable key tied to the specific saga instance and step, and the receiving service is responsible for checking and storing it before performing the side effect - ideally in the same local transaction as the effect itself, so a crash between recording the key and applying the effect can't desynchronize the two.
- How do you choose a sensible timeout for a saga step?Base it on the participant's realistic p99 latency plus margin, not a guess, and make it explicit per step since different services (a fast internal reservation check vs. a third-party payment gateway) have very different latency profiles. Too short causes false-positive failures and unnecessary compensations; too long stalls the whole saga and any customer-facing status waiting on it.
- What should happen when a compensation is retried past its backoff limit and still fails?It should be routed to a dead-letter queue or equivalent holding area with alerting, rather than silently dropped or infinitely retried, because at that point the system has a genuine data inconsistency that needs human judgment - automated retries alone can't resolve a compensation that's failing for a structural reason.
It's like pressing an elevator call button that's already lit - a well-designed button (idempotent) recognizes the request is already registered and does nothing extra, while a poorly-designed one would dispatch a second elevator every time you pressed it again out of impatience.
saying these in an interview costs you the question
- Assumes message delivery is exactly-once by default
- Doesn't know what an idempotency key is or where it's checked
- Thinks retries are always safe regardless of the operation
- Has no answer for what happens when a compensation itself keeps failing
- Doesn't mention timeouts as necessary for detecting a stuck step