skip to content

A checkout flow must debit a payment service and reserve inventory in two separate databases owned by two separate services. Compare coordinating this with a two-phase commit (2PC) protocol versus a Saga with compensating transactions, and say which you'd pick and why.

level: seniorimportance: must knowfreq 75%

answer

  1. 2PC: prepare then commit, coordinator-driven
  2. in-doubt participants block on coordinator crash
  3. XA barely supported outside RDBMS
  4. Saga: local commits + compensations
  5. choreography vs orchestration

basics

~20 s

2PC locks both services until both agree to commit — all-or-nothing, but it freezes resources and can hang if the coordinator dies. A Saga commits each step and undoes earlier ones if a later step fails — faster, briefly inconsistent.

solid answer

~50 s

2PC has a coordinator ask both the payment and inventory services to "prepare" (lock resources, promise they can commit), and only after both say yes does it tell them to commit, giving atomicity but requiring both to hold locks the whole round trip; if the coordinator crashes after prepare but before commit, participants are stuck blocked until it recovers, and most modern datastores and payment gateways don't even support the XA protocol 2PC needs. A Saga instead lets each local transaction commit independently, and if reserving inventory fails, runs a compensating transaction (refund payment) rather than a distributed rollback — trading strict atomicity for availability and no cross-service locking, at the cost of an observable intermediate state and compensations that must be idempotent and sometimes only approximately reversible. For checkout, most systems pick Saga, since payment gateways rarely support 2PC anyway and cross-service locking costs more than compensation complexity.

go deeper

for a junior

Should describe, in plain terms, that 2PC waits for everyone to agree before making anything final, while a Saga does each step now and 'undoes' earlier steps if something later fails.

for a middle

Should describe the prepare/commit phases of 2PC and the local-commit-plus-compensation structure of a Saga, and identify that Sagas don't give isolation.

for a senior

Should discuss 2PC's blocking failure mode on coordinator crash, XA support gaps, and the idempotency/ordering requirements a Saga's compensations must satisfy, and justify a concrete choice for the given scenario.

for a principal

Should weigh choreography vs. orchestration trade-offs for the Saga, discuss monitoring/observability needs for long-running sagas, and articulate the narrow conditions under which 2PC remains appropriate at organizational scale.

## How two-phase commit works Two-phase commit (2PC) coordinates a single atomic transaction across multiple independent databases by splitting commit into two round trips. 1. **Phase one ("prepare")** — a coordinator asks every participant — here, the payment service's database and the inventory service's database — to do everything needed to commit (validate the operation, write it to a durable log, acquire locks on the affected rows) and then reply "yes, I can commit" or "no, abort," without actually making the change visible yet. 2. **Phase two** — only once every participant votes yes does the coordinator move to phase two and tell everyone to actually commit; if any participant votes no, it tells everyone to abort instead. Because every participant has already durably promised it can commit before phase two starts, the protocol guarantees all-or-nothing: either both the debit and the reservation happen, or neither does. ## Why 2PC exists The reason 2PC exists is that it gives real atomicity across resource managers that don't share a transaction log — exactly the payment-database-versus-inventory-database situation here — without requiring you to merge them into one database. It's the same guarantee a single ACID transaction gives you inside one database, extended across a network boundary. ## What that guarantee costs That guarantee is expensive on both sides. - **The lock window.** During the window between "prepare" and "commit," every participant holds its locks — on the customer's payment row, on the inventory row — so no other transaction can touch that data, for as long as the slowest participant and the coordinator's round trip take; under load or across regions this directly throttles throughput and adds latency to every checkout. - **The blocking problem.** Worse is what happens if the coordinator crashes after collecting all the "yes" votes but before sending "commit": participants are contractually obligated to hold their locks and can't unilaterally decide to commit or abort on their own — they're "in doubt" — so they block, potentially for a long time, until the coordinator recovers and tells them what happened. This is 2PC's well-known blocking problem, and it turns a coordinator outage into an availability outage for every transaction it was midway through. - **XA outside relational databases.** On top of that, XA (the standard protocol most 2PC implementations use) is poorly supported outside traditional relational databases — most managed payment gateways, message queues, and NoSQL stores don't implement it at all, which rules 2PC out for a huge share of real microservice stacks regardless of the performance trade-off. ## How a Saga works instead A Saga takes the opposite approach: instead of one distributed transaction, it's a sequence of independent local transactions, each of which commits immediately and durably on its own database, with no cross-service locking at all. Debit payment commits and is visible right away; reserve inventory then runs as its own separate local transaction. If a later step fails — inventory is out of stock — the Saga doesn't roll anything back in the database sense; instead it runs a compensating transaction for each already-completed step, in reverse order (refund the payment), to semantically undo the effect. Sagas can be coordinated by: - **choreography** — each service publishes an event and reacts to the previous service's event, with no central coordinator; - **orchestration** — a central saga orchestrator explicitly calls each step and, on failure, explicitly calls the compensations. Orchestration is easier to reason about and debug at the cost of a new component to build and run. ## What a Saga gives up The trade-off is that Sagas give up atomicity and isolation. - There is a real window — between "payment debited" and "inventory reserved" — where the system is observably in an intermediate state; a concurrent read of the order could see a charged customer with no confirmed inventory. - Compensations also aren't true rollbacks: "refund the payment" is a new, forward-moving transaction, not an undo of the charge, and it can itself fail (the refund API times out), which pushes you into needing retries and idempotency keys on the compensations themselves, and sometimes a manual reconciliation path when compensation genuinely can't complete. - Steps also need to be designed so a partial failure never leaves unrecoverable state — e.g., reserving inventory rather than immediately decrementing it, so the compensation is a cheap "release the hold" rather than something irreversible. ## Which one to pick for this checkout For a checkout flow like this, most real systems pick the Saga, for two converging reasons: 1. payment gateways essentially never participate in XA/2PC, so 2PC is often not technically available as an option; 2. even where it is, the cross-service locking and blocking-on-coordinator-failure risk is a worse operational trade for a customer-facing, high-throughput path than the complexity of writing a "refund payment" compensation. 2PC remains the right tool in narrower cases — coordinating commits across a small number of resource managers that do support XA, within a single trusted operational boundary, where strict atomicity matters more than throughput, such as some internal financial ledger systems spanning two databases under the same team's control.

  • What makes a Saga's compensating transactions hard to get right in practice?
    Compensations must be idempotent, since retries after timeouts can invoke them more than once, and they must be safe even if invoked on a step that never actually completed, in case the failure detection itself was wrong. They also aren't always true inverses — a 'cancel shipment' compensation might not be possible once a package has physically left the warehouse, forcing a fallback to manual intervention rather than a clean automated undo.
  • How does an orchestrated Saga differ operationally from a choreographed one, and what does orchestration cost you?
    An orchestrated Saga has a central component explicitly invoking each step and its compensation, giving one place to read the whole business process and add retries, timeouts, or monitoring. Choreography instead has each service react to the prior service's event with no central coordinator, which avoids adding a new component but makes the end-to-end flow harder to trace, test, and reason about since it's implicit in a web of event subscriptions.
  • Under what circumstances would 2PC still be the right choice over a Saga?
    2PC still makes sense when all participants are within a small, trusted operational boundary and actually support XA, when strict atomicity and isolation genuinely outweigh throughput and availability concerns, and when the operations aren't naturally expressible as reversible compensations — some internal ledger or settlement systems fit this profile.

2PC is like a wedding officiant asking both the bride and groom to privately promise 'I do' first, and only pronouncing them married once both have promised — if the officiant collapses after both promises but before the pronouncement, both are stuck in limbo, unable to leave or marry anyone else, until the officiant recovers. A Saga is more like eloping now and, if the honeymoon booking then falls through, filing an annulment afterward rather than freezing the wedding until every downstream booking is confirmed.

saying these in an interview costs you the question

  • Proposes 2PC across services without checking whether the participants (e.g., a payment gateway) even support XA
  • Doesn't mention that 2PC holds locks across the network for the full round trip
  • Assumes Sagas provide isolation, i.e. that no other transaction can observe the intermediate state
  • Designs compensations without addressing idempotency or partial-compensation-failure
  • Calls a Saga's compensation a 'rollback' as if it undoes the original transaction rather than performing a new, forward-moving corrective transaction

context