skip to content

When a business transaction spans multiple services that each have their own local database, why can't a normal ACID transaction hold it together, and how does a saga's use of compensating actions address that?

level: principalimportance: should knowfreq 55%

answer

  1. 2PC blocks/locks across services -> avoided for availability
  2. saga = sequence of local ACID commits, no cross-service atomicity
  3. compensation = forward-moving semantic undo, not a DB rollback
  4. choreography (event chain, no coordinator) vs orchestration/process manager (explicit coordinator + state)
  5. pivot point = last compensable step

basics

~20 s

Each service's database can only guarantee atomicity for its own local changes, not across services over a network. A saga breaks the transaction into a sequence of local steps, and if a later step fails, it runs 'undo' actions (compensations) for the steps that already succeeded, instead of a true rollback.

solid answer

~50 s

A distributed ACID transaction across independently-owned databases would require a coordination protocol like two-phase commit, which forces every participant to hold locks and stay blocked until all participants vote, making the whole system only as available as its least-available participant and fundamentally at odds with services that are meant to fail and scale independently — so it's avoided in practice. A saga instead sequences the transaction as a series of local ACID transactions, one per service, each of which commits independently and immediately. If a later step fails, the saga doesn't roll back the earlier steps' database transactions (they're already committed and durable) — it instead runs a compensating action for each already-completed step, a separate forward-moving operation that semantically undoes the business effect (refund instead of un-charge, release-reservation instead of un-reserve). The trade-off is that the system is only ever eventually consistent, and compensations must be designed for every step up front, including handling the case where a compensation itself fails.

go deeper

for a junior

Should recognize that a transaction across multiple services' databases can't have one shared commit/rollback, and that 'undo' steps are used instead.

for a middle

Should describe the sequence-of-local-transactions mechanic and give a correct example of a compensating action paired with its forward step.

for a senior

Should compare choreography and orchestration, articulate why 2PC is avoided, and reason about designing compensations for a specific multi-step flow, including a pivot point.

for a principal

Should decide when a saga is warranted at all versus redesigning service boundaries to avoid needing one, design the orchestrator's own reliability (durable, resumable state) and failure-of-compensation handling, and make the choreography-vs-orchestration call for a system at scale with the operational-visibility trade-off in mind.

## Where the single-database guarantee ends A single-database ACID transaction gives you **atomicity for free**: either every statement in the transaction commits, or none of them do, enforced by the database engine. The moment a business transaction spans multiple services, each with its own database it does not share with the others, that guarantee disappears, because no single database engine can span the transaction — each service's local commit is genuinely final and independent the instant it happens. The classical answer to this problem in distributed systems is **two-phase commit** (2PC): a coordinator asks every participant to prepare (lock resources and vote yes/no), and only after every participant votes yes does it tell them all to commit. - 2PC does restore atomicity, but **at a severe cost**: every participant must hold locks on its resources for the entire duration of the protocol, meaning the whole transaction is only as fast, and only as available, as its slowest or least-reliable participant, and a coordinator crash mid-protocol can leave participants blocked indefinitely holding locks. - This is exactly the kind of **tight run-time coupling** that independently deployable, independently scalable services are built to avoid, which is why 2PC across service boundaries is rare in practice despite solving the problem 'correctly' on paper. ## What a saga does instead The saga pattern accepts that atomicity across services is off the table and replaces it with a different guarantee: **eventual consistency achieved through compensation**. A saga decomposes the business transaction into a sequence of steps, where each step is a normal local ACID transaction inside one service, committed and durable the instant it completes — there is no cross-service locking and no coordinator holding participants hostage while others are still deciding. 1. If every step succeeds, the saga completes normally. 2. If a step fails partway through, the saga does not (and cannot) undo the database commits of the steps that already succeeded — those are already true and durable facts. 3. Instead, for each already-completed step, the saga triggers a **compensating transaction**: a separate, deliberately designed operation whose job is to semantically reverse that step's business effect. Compensations are not database rollbacks; they are new forward-moving business actions: - `RefundPayment` compensates `ChargePayment`, - `ReleaseInventoryReservation` compensates `ReserveInventory`, - `CancelShipment` compensates `ScheduleShipment`. Because compensations are themselves separate operations against separate services, they inherit all the same delivery and idempotency concerns as any other command in the system — a compensation can itself fail or be retried, so it needs to be designed to be safely re-runnable. ## Two ways to coordinate one There are two structural styles for coordinating a saga. - **Choreography** has each service publish an event when it finishes its step, and the next service(s) react to that event to perform their own step and publish their own event — there's no central coordinator, just a chain of services each reacting to the last one's fact. This keeps individual services simple and loosely coupled, but the overall business process ends up implicit, scattered across the event-handling logic of every participating service, with no single place that shows the whole flow — which makes it hard to answer 'what state is order #789's saga in right now?' without piecing it together from many services' logs. - **Orchestration** (the process manager pattern) instead introduces an explicit coordinator that owns the sequence: it sends a command to the first service, waits for the result, decides the next command to send, and explicitly tracks the saga's state (which steps have completed, which compensations may be needed) in its own persistent store. This makes the business process visible and centrally controllable — there's one place to look for 'what's the state of this order's saga' — at the cost of that coordinator becoming a real dependency every participating service now indirectly relies on, and it must itself be built reliably (durable state, resumable after crash) since it's now the single place that knows how to drive the whole flow forward or trigger compensation. ## The trade-off The trade-off, in both styles, is **consistency model and design burden**. - You give up **strong consistency** — there is a real window, however short, during which some services have already reflected the transaction's effect and others have not yet, and any code that reads state during that window can observe a partially-completed saga. - In exchange you get **availability and independence**: no service is blocked holding locks waiting on another service's decision. - **The design burden** is that compensations aren't automatic or symmetric — every forward step that can leave a lasting effect needs a corresponding compensating action designed and tested up front, and not every action is cleanly compensable (an email that's already been sent, or a physical package that's already left the warehouse, can't be truly 'un-sent' or 'un-shipped,' only followed up with a corrective action like a cancellation notice). ## Failure modes Failure modes concentrate at the edges: 1. What happens when a compensation itself fails needs its own answer — usually retries with alerting, sometimes a manual/human recovery path, since you don't want an infinite compensation-failure loop; 2. **Pivot points** — steps chosen deliberately as the last steps that can still be compensated, after which the saga is committed forward no matter what (e.g., once a shipment physically leaves the warehouse, you don't compensate, you route to a different corrective process). ## Putting it together A concrete scenario: an order saga orchestrator sends `ReserveInventory`, then `ChargePayment`, then `ScheduleShipment`; if `ChargePayment` fails after `ReserveInventory` succeeded, the orchestrator issues `ReleaseInventoryReservation` as a compensation and marks the order failed, without ever needing a distributed lock spanning the inventory and payment services.

  • Why is two-phase commit generally avoided across independently owned services even though it does provide real atomicity?
    2PC requires every participant to hold locks on its resources for the full duration of the protocol and remain blocked until the coordinator's final decision, which ties every participant's availability to the slowest or least-reliable one and to the coordinator itself. That's the opposite of what independently deployable, independently scalable services are meant to achieve, so most systems accept eventual consistency via sagas instead.
  • What is a 'pivot point' in a saga, and why does it matter for compensation design?
    A pivot point is the step in the saga after which the outcome is no longer safely compensable — for example, once a shipment has physically left the warehouse, you can't truly un-ship it. Steps before the pivot point need real compensating actions designed; steps after it are treated as committed-forward, with any recovery handled as a corrective business process (like a return) rather than a compensation.
  • How does choreography's lack of a central coordinator create a specific operational difficulty compared to orchestration?
    Because the business process is scattered across each service's event-handling logic with no single owner, answering 'what state is this specific transaction in right now' requires reconstructing the flow from multiple services' independent event histories rather than checking one coordinator's state store. This makes monitoring, debugging, and manually intervening in a stuck transaction significantly harder as the number of participating services grows.

It's like booking a multi-leg trip yourself instead of through one travel agent with a single unified booking: you book the flight, then the hotel, then the rental car, each a separate confirmed transaction. If the rental car booking fails, you don't magically un-book the flight and hotel — you deliberately cancel them (and eat any cancellation fee), which is compensation, not rollback.

saying these in an interview costs you the question

  • proposes distributed two-phase commit across services as the default solution
  • describes a saga's failure handling as 'rolling back' the earlier services' database transactions
  • assumes compensating actions are automatically derived and don't need to be separately designed
  • can't distinguish choreography from orchestration
  • has no answer for what happens when a compensating action itself fails
  • assumes strong consistency is preserved throughout a saga

context