skip to content

In a microservices order flow where Orders, Payments, and Inventory each own separate databases, why do teams typically reach for a saga instead of a two-phase commit (2PC) to keep the data consistent across them?

level: juniorimportance: must knowfreq 80%

answer

  1. 2PC = coordinator + locks across DBs
  2. saga = chain of local ACID txns
  3. compensating tx, not rollback
  4. eventual not immediate consistency
  5. database-per-service breaks XA

basics

~20 s

A saga splits one big transaction into small local ones, each service commits its own piece right away; if something later fails, earlier steps are undone with a follow-up 'undo' action instead of everyone waiting and locking together like 2PC.

solid answer

~40 s

2PC needs a coordinator that holds locks on every participant's database until all of them vote to commit, which ties the whole flow's availability to the slowest or least reliable participant, doesn't cross different database engines or brokers cleanly, and breaks each service's database-per-service autonomy. It also doesn't work at all over asynchronous messaging. A saga instead runs a chain of independent local ACID transactions — one per service — each committing immediately and triggering the next step via an event or command. If a later step fails, already-committed steps are undone by explicit compensating transactions (refund the payment, release the reserved stock) rather than a DB-level rollback. You trade strong atomicity and isolation for availability, service autonomy, and horizontal scalability, and you take on the burden of designing compensations and tolerating temporary inconsistency.

go deeper

for a junior

Should say a saga breaks the transaction into steps that each commit on their own, and that failures are undone with a follow-up action rather than an automatic rollback. Doesn't need to know 2PC internals in depth.

for a middle

Should name 2PC's prepare/commit phases and locking, and explain that database-per-service plus async messaging make 2PC impractical; should use the term 'compensating transaction' correctly.

for a senior

Should articulate the availability/coupling cost of 2PC concretely (blocking on the slowest participant, coordinator as SPOF) and connect the saga's eventual consistency to specific design consequences (idempotent compensations, visible interim states).

for a principal

Should be able to discuss when 2PC or a hybrid (e.g., XA within a single-vendor cluster, or a saga combined with a coordinating service) is still the right call, and articulate the org-level reason microservices default to sagas: independent deployability trumps strict consistency.

## How two-phase commit works **Two-phase commit (2PC)** is the classic way relational databases keep a transaction atomic across multiple resources: 1. A coordinator asks every participant to "prepare" (lock the rows and promise it can commit). 2. It waits for all of them to vote yes. 3. It then tells everyone to "commit". 4. If any participant votes no, or times out, the coordinator tells everyone to roll back. The mechanism works, but it requires every participant to hold locks for the entire duration of the vote-and-decide round trip, and it requires a single coordinator that all participants trust and can reach. ## Why the model breaks down in microservices In a microservices architecture that model breaks down for several concrete reasons. 1. First, **"database per service"** is a foundational rule — each service owns its schema and nobody else touches it directly — so there is no single database engine to run 2PC across; you'd need a distributed transaction coordinator (like XA) that every service's data store and message broker support, and most modern brokers (Kafka, SQS) and NoSQL stores don't implement XA at all. 2. Second, even where XA is available, holding locks across service boundaries for the round trip of a network call means the whole chain's throughput is capped by its slowest, least available participant — one service having a bad day (GC pause, deploy, disk pressure) blocks every other service's writes that touch the same rows. 3. Third, 2PC **couples deployment and failure domains**: a coordinator outage or partition can leave participants in "in-doubt" states, holding locks indefinitely until an operator intervenes. None of this fits the reason microservices exist — independent deployability and failure isolation. ## What a saga does instead A saga solves the same consistency problem with a different mechanism: instead of one distributed transaction, it's a sequence of **local transactions**. - Each one is fully ACID within its own service and database. - Each one commits immediately rather than waiting on anyone else. - After a step commits, the service publishes a domain event (**choreography**) or reports completion to an orchestrator (**orchestration**), which triggers the next step. Because each step commits independently, there's a window — sometimes seconds, sometimes longer — where the system is in an intermediate, not-yet-consistent state; for example, an order can exist as "pending payment" after Orders commits but before Payments has run. This is **"eventual consistency"**: the saga guarantees the system reaches a consistent end state (either all steps succeed, or the completed ones are undone), but not that it's consistent at every instant. ## Compensating transactions, not rollback The other half of the mechanism is what happens when a step fails partway through. Because there's no distributed rollback, the saga must run **compensating transactions** for every step that already committed, in reverse order: - cancel the reservation; - refund the charge; - restock the inventory. These are not database-level undos — they are new, deliberate business operations, and they must be written to make semantic sense (a refund is not literally "un-charging a card"; it's a separate ledger entry) and, critically, be safe to retry, because the messaging layer that delivers "please compensate" can redeliver on timeout. ## The trade-off The trade-off is explicit: | | 2PC | A saga | |---|---|---| | Gives you | atomicity and isolation for free | availability, service autonomy, and no cross-service locking | | At the cost of | availability, coupling, and throughput | having to hand-design every compensation, tolerate a visible in-between state, and reason carefully about ordering and idempotency | Teams pick sagas by default in microservices not because they're strictly "better" but because 2PC's assumptions (shared coordinator, XA-capable stores, tolerance for blocking) don't hold once services are independently deployed, independently owned, and communicate over unreliable networks and brokers. ## A concrete example A concrete, widely cited example is an e-commerce checkout: 1. **Orders** creates a pending order and commits. 2. **Payments** charges the card and commits. 3. **Inventory** decrements stock and commits. 4. **Shipping** schedules a shipment and commits. If Inventory finds it's out of stock after Payments already charged the card, the saga runs a compensating "refund payment" transaction and marks the order failed, rather than ever having attempted to hold a lock across all four services' databases at once.

  • Does a saga give you the same isolation guarantee as a single ACID transaction?
    No — between steps, other transactions can read the intermediate state (e.g., an order marked 'pending' before payment completes), which is called a lack of isolation or an 'anomaly.' Sagas are typically deployed with mitigations like semantic locks or by designing reads to tolerate the interim state, not by faking true isolation.
  • What happens if a compensating transaction itself fails?
    It must be retried until it succeeds, which is why compensations are designed to be idempotent and retriable; some systems escalate unrecoverable compensation failures to a dead-letter queue and human/ops intervention rather than silently giving up.
  • Could you use 2PC for just two services and consider a saga overkill?
    In principle yes if both participate in the same XA-capable transaction manager and you accept the coupling, but most teams still prefer a saga once services deploy independently, because the coupling and blocking cost shows up in outages even at small scale.

2PC is like a wedding where the officiant won't pronounce anyone married until every guest, caterer, and photographer simultaneously confirms they're ready — one late guest stalls the whole ceremony. A saga is like a relay race: each runner completes their leg and hands off immediately, and if a later runner drops the baton, the earlier runners' completed legs are 'undone' by a separate cleanup crew rather than the race rewinding itself.

saying these in an interview costs you the question

  • Says sagas provide the same atomicity/isolation as a real transaction
  • Doesn't mention compensating transactions when explaining how a saga undoes work
  • Thinks 2PC 'doesn't work' for some vague reason rather than naming the locking/coordinator coupling
  • Proposes 2PC across services with different database vendors without acknowledging XA/driver support is required
  • Assumes intermediate states are never visible to other transactions

context