A team is designing a checkout flow that needs to debit a payment account, decrement inventory, and create a shipping order across three separate services. They're deciding between wrapping the whole thing in a two-phase-commit (2PC) distributed transaction versus using a Saga (a sequence of local transactions with compensating actions to undo earlier steps if a later one fails). What factors should drive that decision, and what does each approach give up?
answer
- atomic+blocking vs eventual+compensating
- compensations aren't true rollbacks
- Saga intermediate state is visible
- 2PC needs protocol-compatible participants
- trusted/low-latency boundary vs cross-org boundary
basics
~20 s2PC keeps everything locked until all three services agree, so it's always consistent but slow and fragile if one service is down. A Saga lets each step commit right away and fixes mistakes afterward with compensating actions, so it's faster and more resilient but briefly shows an inconsistent state.
solid answer
~60 s2PC gives strong, immediate atomicity — either all three writes happen or none do, visible to no one until commit — at the cost of holding locks across a network round trip to all three services and blocking if any one of them (or the coordinator) is slow or down; it also requires every participant to speak the same commit protocol, ruling out most external/legacy services. A Saga breaks the transaction into three independently-committed local transactions with compensating actions defined for rollback (refund the payment, restock the inventory) if a later step fails; it never locks across services and tolerates individual service slowness far better, but it only gives eventual consistency — there's a real window where payment is debited but inventory hasn't decremented yet — and compensations must be idempotent and handle side effects that already escaped (a shipping email already sent). The right choice depends on how tolerant the business is of transient inconsistency versus how much it needs true locking-based isolation and how many participants are trusted, low-latency, protocol-compatible systems.
go deeper
Should understand that 2PC keeps things locked and consistent but slow, while a Saga commits pieces one at a time and fixes mistakes afterward.
Should be able to name the cost on each side — blocking/latency for 2PC, temporary visible inconsistency for Saga — and know compensations aren't true rollbacks.
Should reason about concrete decision factors — trust boundary, protocol compatibility, latency tolerance, long-running workflows — and design mitigations like soft reservations to shrink a Saga's inconsistency window.
Should set organizational and architectural policy — when to allow 2PC at all in the system (e.g., only within one database engine's internal shards), and how to standardize Saga compensation and idempotency conventions across teams so cross-service transactions stay reliable at scale.
## Where the cost lands The choice between two-phase commit and a Saga is really a choice about where you're willing to pay a cost: | | 2PC | Saga | |---|---|---| | **Pays with** | availability and latency | a temporary, visible inconsistency window | | **To buy** | strict, immediate atomicity and isolation | availability, latency, and looser coupling between participants | ## How each one works **Mechanism recap for context.** - **2PC** coordinates all participants to vote and then commit or abort together — the defining property is that no partial state is ever externally visible: either the payment debit, inventory decrement, and shipping order all become visible atomically, or none do, and locks held during the prepare window prevent anyone from observing or corrupting the in-flight state. - **A Saga** instead commits each local step as its own independent transaction, immediately visible, and only reacts after the fact if a later step fails: it invokes compensating actions (a refund, a restock) to semantically — not literally — undo the already-committed earlier steps. **Compensation is not a rollback** in the database sense; it's a new forward-moving transaction that tries to cancel the effect of a previous one, and it has to be designed case by case (refunding a payment is a new debit/credit that reverses the net effect, not an undo). ## What each side gives up **The core trade-off: isolation and blocking versus visibility and resilience.** - Because 2PC holds locks across the network round trip to every participant, its latency is bounded by the slowest participant and network hop, and if any one participant (or the coordinator) is unavailable, the whole transaction stalls and already-prepared participants sit holding locks, which can freeze unrelated work on those systems too. - A Saga has no analogous blocking window: each step commits and moves on, so a failure late in the sequence doesn't lock earlier resources — it just triggers compensations. This makes Sagas dramatically more tolerant of one participant being slow or briefly unavailable, and is a major reason they're preferred for workflows spanning many services, especially ones owned by different teams or companies (a travel booking that reserves a flight, a hotel, and a rental car — you cannot 2PC-lock an airline's reservation system). ## What a Saga actually costs The real cost of a Saga is that intermediate states are genuinely visible, not just theoretically. Between the payment debit committing and the inventory decrement committing, another process reading the order or the inventory count can observe a state that will later turn out to have been transient. Handling this requires either: - designing consumers to tolerate transience (an order status of 'processing' rather than treating partial state as final), or - adding a semantic reservation step (inventory gets a soft 'reserved' hold rather than a hard decrement, itself a smaller-scoped local transaction). Compensations also have their own correctness burden: they must be **idempotent**, since retries happen, and they must account for side effects that already escaped the system boundary — you can refund a payment, but you can't un-send a 'your order shipped' email, so Saga design often sequences externally-visible, hard-to-compensate steps last, after every compensable step has succeeded. ## Choosing between them **When to pick which.** - **2PC** (or, more realistically, an XA-style transaction) is right when all participants are within a trusted, low-latency boundary (two tables in the same database, or a small number of internal systems supporting the same commit protocol) and the business genuinely cannot tolerate even a brief inconsistent window — financial ledger postings inside a single bank's core system are a classic case. - **A Saga** is right once you're crossing service or organizational boundaries, involve external/third-party systems that will never implement a 2PC-compatible prepare phase, need to tolerate individual participant slowness without freezing the whole workflow, or are dealing with a genuinely long-running process (multi-day shipping fulfillment) where holding database locks for the duration is a non-starter regardless of protocol. In practice, most modern microservice architectures default to Sagas for cross-service business transactions specifically because 2PC's blocking and all-participants-must-speak-the-protocol requirements don't survive contact with independently deployed, independently owned services — 2PC largely survives only within a single database engine's own internal multi-shard commits (e.g., Google Spanner, CockroachDB) where the 'participants' are all under one operator's control and one consistent protocol stack.
- Why can't compensating actions in a Saga simply be the reverse database operation of the original step?Because the original step's effects may have already propagated outside the database — a payment provider may have already notified a card network, or a downstream system may have already read the committed state — so 'reversing' has to be a new, semantically correct action (like issuing a refund) rather than literally deleting the original row. It also has to be idempotent, since Saga steps and compensations are commonly retried.
- What makes a service unsuitable as a 2PC participant even if it's technically reachable?If it doesn't implement a prepare/vote step that lets it durably promise to commit on request — most third-party APIs and many microservices simply execute and commit immediately without any such hook — then it can't participate honestly in the protocol's voting phase, so 2PC can't be layered on top of it without changing its API.
- In the checkout example, what's one concrete technique to reduce the visible inconsistency window in a Saga without resorting to 2PC?Use a soft reservation/hold on the scarce resource — e.g., inventory gets marked 'reserved' for this order rather than immediately decremented — so other processes see accurate availability sooner, and the hold is only converted to a hard decrement (or released back) once the rest of the saga resolves. This narrows the window where the system looks wrong without requiring cross-service locking.
2PC is like a group of friends agreeing to only pay for a shared gift once everyone has confirmed their share is ready, holding the wallet closed until then. A Saga is like each friend paying their share right away and, if the gift falls through, everyone individually requests a refund afterward — faster, but there's a window where money has moved and the gift still might not happen.
saying these in an interview costs you the question
- Thinks a Saga provides the same isolation guarantees as 2PC
- Believes compensating actions are just automatic database rollbacks
- Doesn't recognize that 2PC requires every participant to support the protocol
- Ignores the visible-intermediate-state cost of Sagas entirely
- Recommends 2PC across independently-deployed microservices as a default