skip to content

When designing the compensating step for a saga stage — say, 'release reserved inventory' to undo an earlier 'reserve inventory' step — what makes some compensations straightforward and others genuinely hard or impossible to write correctly?

level: seniorimportance: should knowfreq 65%

answer

  1. compensation = new forward op, not rollback
  2. order steps: reversible first, irreversible last
  3. semantic lock for in-flight races
  4. idempotency key for redelivered compensations
  5. some effects truly can't be undone (email sent, non-refundable booking)

basics

~20 s

A compensation is a new action that semantically cancels an earlier step, not a database undo. It's easy when the earlier step is fully reversible (like un-reserving stock); it's hard or impossible when the earlier step already had a real-world, external, or irreversible effect, like an email that was already sent or money that already left the system.

solid answer

~50 s

A compensating transaction is a forward-moving business operation designed to semantically negate an earlier step's effect, not a literal database rollback — 'refund' rather than 'un-charge.' Compensations are straightforward when the original step's effect is entirely contained within the service's own data and reversible without external side effects — releasing a reservation, decrementing a counter back. They get hard when the original step had an irrevocable external effect: an email or SMS already sent to a customer, a non-refundable third-party API call already made, a physical shipment already dispatched. In those cases you can't undo the effect, only mitigate it (send a follow-up 'ignore the previous email' notice, initiate a return process instead of un-shipping). Compensations also need to be idempotent and safe under retries, must handle the case where the original step is still in flight when the compensation is triggered (requiring careful state checks), and for non-atomic steps you sometimes need semantic locks to prevent other transactions from acting on data mid-compensation.

go deeper

for a junior

Should understand that undoing a step means running a new opposite-ish action (like 'cancel' or 'release'), not that the database automatically reverses it.

for a middle

Should be able to give an example of an easy compensation (release a reservation) and note that compensations need to handle being called more than once.

for a senior

Should discuss ordering steps by reversibility, semantic locks for races between in-flight steps and their compensations, and idempotency keys for redelivered compensating commands.

for a principal

Should identify genuinely non-compensable effects and propose structural mitigations — reordering the saga, deferring irreversible actions, or falling back to human/manual processes — as a design-level decision, not just a coding detail.

## A new operation, not a rollback A compensating transaction is not a database rollback — there is no undo log spanning services the way there is inside a single ACID transaction. It's a deliberately authored, forward-moving business operation whose purpose is to semantically cancel the observable effect of an earlier saga step. This distinction matters because it changes how you have to think about designing one: instead of asking "how do I reverse this write," you have to ask "what new operation makes the world consistent again, given that the original operation already happened and other things may have happened since." ## The easy end of the spectrum The easy end of the spectrum is when a step's effect is entirely internal to the owning service's data and structurally reversible: - a "reserve 2 units of `SKU-123`" step decrements an `available_quantity` counter and inserts a reservation row; - its compensation, "release reservation," deletes the reservation row and increments the counter back. Nothing outside the service's database was touched, no third party was informed, and the operation is symmetric — you can write it, test it, and reason about it almost like a mathematical inverse. ## The hard end of the spectrum The hard end is when the original step had an effect that left the system's boundary: - money moved to a third party; - a physical action was triggered; - or a human was notified. Consider a step that calls a card network to capture a payment: the compensating step, "refund," is not truly symmetric — it's a new transaction that returns funds, but it may take days to settle, may incur different fees than the original charge, and shows up as two separate ledger entries rather than one that never happened, which matters for accounting and tax reporting. Now consider a step that sends a "your order has shipped" email or SMS: there is no API call that unsends a message a human has already read. The best you can do is compensate with a different kind of action entirely — send a follow-up correction, and design the original step so the irrevocable action (the send) happens as late as possible in the saga, ideally after all the steps that are more likely to fail have already succeeded, specifically to minimize the number of scenarios where you'd need to compensate for something unsendable. This is a general design principle: **order saga steps so that the truly hard-to-compensate steps run last, and the easily reversible steps run first**, so failures are statistically more likely to occur before you've done anything irreversible. ## Concurrency and timing A second class of difficulty is concurrency and timing. If a compensation for "reserve inventory" fires while the original reservation step is still technically in flight (e.g., the orchestrator issued the compensating command based on a timeout, but the reservation actually succeeds moments later), you can end up compensating something that hasn't happened yet, or racing the compensation against the original step's own retry. This is usually handled with a **semantic lock**: the reserved row is marked with a state flag (e.g., `PENDING_CONFIRMATION`) that other transactions — including a stray compensation — must check before acting, so a late-arriving compensation for a step that actually succeeded is either rejected or converted into the correct "release" action rather than silently corrupting state or running twice. ## Idempotency under redelivery A third difficulty is **idempotency under redelivery**: because saga coordination runs over unreliable messaging, the same compensating command can be delivered more than once (a retry after a timeout that actually succeeded, or a redelivery from an at-least-once broker). The compensation must be safe to execute twice — releasing an already-released reservation should be a no-op, not an error or a double-increment of the counter — which usually means keying the operation by an idempotency key (e.g., the original reservation ID) rather than by a raw "add N back" instruction. ## When nothing can compensate The genuinely impossible case is when the original step's effect can't be compensated by any business-level operation at all — most classically, a legally or physically final action, like a non-refundable, non-cancelable third-party booking, or data that has already been disclosed to another party in a way that can't be recalled. Here the saga design has to change: 1. either make that step the very last one in the sequence (so nothing after it can fail and require compensating it); 2. require manual/human intervention as the "compensation" (a support ticket, a customer-service credit); 3. or restructure the business process to avoid ever needing to undo it. A concrete real-world example: airline booking sagas typically place the actual seat-lock/ticket-issuance step as late as possible and put the truly irreversible fare-rule-locked purchase confirmation last, precisely because a ticket, once issued against certain fare classes, is either non-refundable or refundable only through a separate, slower, fee-bearing process rather than a clean compensating transaction.

  • Why does step ordering matter for how easy compensation is, given that all steps eventually need compensations written anyway?
    Every step needs a compensation defined, but ordering determines how often that compensation actually has to run against an irreversible effect. Running easily-reversible, high-failure-probability steps (like inventory checks) first and hard-to-reverse, low-failure-probability steps (like sending a confirmation email) last statistically minimizes the number of real-world irreversible actions you end up trying to compensate for.
  • What is a semantic lock and why does a saga need one instead of a normal database lock?
    A semantic lock is an application-level flag (like a status column set to PENDING) that other transactions check before acting on that data, used because a normal database row lock can't be held open across the whole multi-step, multi-service saga without blocking other work; the semantic lock achieves similar protection against races without requiring a long-held database transaction.
  • If a compensation truly can't undo an effect, what are the realistic options?
    Escalate to a human-driven process (support credit, manual correction, customer apology plus goodwill gesture), redesign the saga so that step runs last and nothing can fail after it, or restructure the business process to avoid the irreversible action being taken before consistency is confirmed — for example, holding a booking instead of confirming it until the whole saga completes.

Undoing a database transaction is like erasing a pencil mark — the paper looks like it never happened. A compensating transaction is like sending a follow-up letter to correct an earlier one you already mailed — the earlier letter still arrived and was read, so the best you can do is send a clear correction, and if the earlier letter announced something irreversible (like 'the package left the warehouse'), no follow-up letter can call the truck back.

saying these in an interview costs you the question

  • Describes a compensating transaction as a database rollback rather than a new business operation
  • Doesn't recognize that some effects (sent notifications, executed third-party payments to others, physical shipments) can't be cleanly undone
  • Writes a compensation that isn't idempotent, assuming it will only ever be called once
  • Ignores the race between an in-flight original step and a compensation triggered by a timeout
  • Places genuinely irreversible steps early in the saga without noticing the compensation burden that creates

context