skip to content

Distributed Transactions & Saga

Sagas keep multi-service operations consistent without a distributed transaction, using a sequence of local transactions and compensating actions. You will compare choreography with orchestration and learn the outbox pattern that makes the messages reliable.

part ofMicroservices architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In a microservices order flow where Orders, Payments, and Inventory each own separate databases, why do teams typically reach for a saga instead of a two-phase commit (2PC) to keep the data consistent across them?

level: juniorimportance: must knowfreq 80%

answer

  1. 2PC = coordinator + locks across DBs
  2. saga = chain of local ACID txns
  3. compensating tx, not rollback
  4. eventual not immediate consistency
  5. database-per-service breaks XA

basics

~20 s

A saga splits one big transaction into small local ones, each service commits its own piece right away; if something later fails, earlier steps are undone with a follow-up 'undo' action instead of everyone waiting and locking together like 2PC.

solid answer

~40 s

2PC needs a coordinator that holds locks on every participant's database until all of them vote to commit, which ties the whole flow's availability to the slowest or least reliable participant, doesn't cross different database engines or brokers cleanly, and breaks each service's database-per-service autonomy. It also doesn't work at all over asynchronous messaging. A saga instead runs a chain of independent local ACID transactions — one per service — each committing immediately and triggering the next step via an event or command. If a later step fails, already-committed steps are undone by explicit compensating transactions (refund the payment, release the reserved stock) rather than a DB-level rollback. You trade strong atomicity and isolation for availability, service autonomy, and horizontal scalability, and you take on the burden of designing compensations and tolerating temporary inconsistency.

go deeper

for a junior

Should say a saga breaks the transaction into steps that each commit on their own, and that failures are undone with a follow-up action rather than an automatic rollback. Doesn't need to know 2PC internals in depth.

for a middle

Should name 2PC's prepare/commit phases and locking, and explain that database-per-service plus async messaging make 2PC impractical; should use the term 'compensating transaction' correctly.

for a senior

Should articulate the availability/coupling cost of 2PC concretely (blocking on the slowest participant, coordinator as SPOF) and connect the saga's eventual consistency to specific design consequences (idempotent compensations, visible interim states).

for a principal

Should be able to discuss when 2PC or a hybrid (e.g., XA within a single-vendor cluster, or a saga combined with a coordinating service) is still the right call, and articulate the org-level reason microservices default to sagas: independent deployability trumps strict consistency.

## How two-phase commit works **Two-phase commit (2PC)** is the classic way relational databases keep a transaction atomic across multiple resources: 1. A coordinator asks every participant to "prepare" (lock the rows and promise it can commit). 2. It waits for all of them to vote yes. 3. It then tells everyone to "commit". 4. If any participant votes no, or times out, the coordinator tells everyone to roll back. The mechanism works, but it requires every participant to hold locks for the entire duration of the vote-and-decide round trip, and it requires a single coordinator that all participants trust and can reach. ## Why the model breaks down in microservices In a microservices architecture that model breaks down for several concrete reasons. 1. First, **"database per service"** is a foundational rule — each service owns its schema and nobody else touches it directly — so there is no single database engine to run 2PC across; you'd need a distributed transaction coordinator (like XA) that every service's data store and message broker support, and most modern brokers (Kafka, SQS) and NoSQL stores don't implement XA at all. 2. Second, even where XA is available, holding locks across service boundaries for the round trip of a network call means the whole chain's throughput is capped by its slowest, least available participant — one service having a bad day (GC pause, deploy, disk pressure) blocks every other service's writes that touch the same rows. 3. Third, 2PC **couples deployment and failure domains**: a coordinator outage or partition can leave participants in "in-doubt" states, holding locks indefinitely until an operator intervenes. None of this fits the reason microservices exist — independent deployability and failure isolation. ## What a saga does instead A saga solves the same consistency problem with a different mechanism: instead of one distributed transaction, it's a sequence of **local transactions**. - Each one is fully ACID within its own service and database. - Each one commits immediately rather than waiting on anyone else. - After a step commits, the service publishes a domain event (**choreography**) or reports completion to an orchestrator (**orchestration**), which triggers the next step. Because each step commits independently, there's a window — sometimes seconds, sometimes longer — where the system is in an intermediate, not-yet-consistent state; for example, an order can exist as "pending payment" after Orders commits but before Payments has run. This is **"eventual consistency"**: the saga guarantees the system reaches a consistent end state (either all steps succeed, or the completed ones are undone), but not that it's consistent at every instant. ## Compensating transactions, not rollback The other half of the mechanism is what happens when a step fails partway through. Because there's no distributed rollback, the saga must run **compensating transactions** for every step that already committed, in reverse order: - cancel the reservation; - refund the charge; - restock the inventory. These are not database-level undos — they are new, deliberate business operations, and they must be written to make semantic sense (a refund is not literally "un-charging a card"; it's a separate ledger entry) and, critically, be safe to retry, because the messaging layer that delivers "please compensate" can redeliver on timeout. ## The trade-off The trade-off is explicit: | | 2PC | A saga | |---|---|---| | Gives you | atomicity and isolation for free | availability, service autonomy, and no cross-service locking | | At the cost of | availability, coupling, and throughput | having to hand-design every compensation, tolerate a visible in-between state, and reason carefully about ordering and idempotency | Teams pick sagas by default in microservices not because they're strictly "better" but because 2PC's assumptions (shared coordinator, XA-capable stores, tolerance for blocking) don't hold once services are independently deployed, independently owned, and communicate over unreliable networks and brokers. ## A concrete example A concrete, widely cited example is an e-commerce checkout: 1. **Orders** creates a pending order and commits. 2. **Payments** charges the card and commits. 3. **Inventory** decrements stock and commits. 4. **Shipping** schedules a shipment and commits. If Inventory finds it's out of stock after Payments already charged the card, the saga runs a compensating "refund payment" transaction and marks the order failed, rather than ever having attempted to hold a lock across all four services' databases at once.

  • Does a saga give you the same isolation guarantee as a single ACID transaction?
    No — between steps, other transactions can read the intermediate state (e.g., an order marked 'pending' before payment completes), which is called a lack of isolation or an 'anomaly.' Sagas are typically deployed with mitigations like semantic locks or by designing reads to tolerate the interim state, not by faking true isolation.
  • What happens if a compensating transaction itself fails?
    It must be retried until it succeeds, which is why compensations are designed to be idempotent and retriable; some systems escalate unrecoverable compensation failures to a dead-letter queue and human/ops intervention rather than silently giving up.
  • Could you use 2PC for just two services and consider a saga overkill?
    In principle yes if both participate in the same XA-capable transaction manager and you accept the coupling, but most teams still prefer a saga once services deploy independently, because the coupling and blocking cost shows up in outages even at small scale.

2PC is like a wedding where the officiant won't pronounce anyone married until every guest, caterer, and photographer simultaneously confirms they're ready — one late guest stalls the whole ceremony. A saga is like a relay race: each runner completes their leg and hands off immediately, and if a later runner drops the baton, the earlier runners' completed legs are 'undone' by a separate cleanup crew rather than the race rewinding itself.

saying these in an interview costs you the question

  • Says sagas provide the same atomicity/isolation as a real transaction
  • Doesn't mention compensating transactions when explaining how a saga undoes work
  • Thinks 2PC 'doesn't work' for some vague reason rather than naming the locking/coordinator coupling
  • Proposes 2PC across services with different database vendors without acknowledging XA/driver support is required
  • Assumes intermediate states are never visible to other transactions

context

open as a page

When implementing a saga across Order, Payment, and Shipping services, what's the difference between a choreography-based saga and an orchestration-based saga, and what would push you to pick one over the other?

level: middleimportance: must knowfreq 85%

basics

~10 s

Choreography: each service listens for events and reacts on its own, no one's in charge. Orchestration: one central coordinator tells each service what to do next, step by step.

open as a page

A service updates its own database and then needs to publish an event so the next saga step can run. What is the 'dual write' problem this creates, and how does the transactional outbox pattern solve it?

level: middleimportance: must knowfreq 75%

basics

~20 s

If a service saves to its database and separately sends a message, one can succeed while the other fails, leaving things out of sync. The outbox pattern saves the message in the same database transaction as the data change, then a separate process reliably delivers it, so both always happen together or neither does.

open as a page

When designing the compensating step for a saga stage — say, 'release reserved inventory' to undo an earlier 'reserve inventory' step — what makes some compensations straightforward and others genuinely hard or impossible to write correctly?

level: seniorimportance: should knowfreq 65%

basics

~20 s

A compensation is a new action that semantically cancels an earlier step, not a database undo. It's easy when the earlier step is fully reversible (like un-reserving stock); it's hard or impossible when the earlier step already had a real-world, external, or irreversible effect, like an email that was already sent or money that already left the system.

open as a page

A saga step's message handler receives the same 'PaymentCharged' event twice because the broker redelivered it after a timeout. What is the inbox pattern / idempotent consumer approach, and how does it prevent the handler from applying the event's effect twice?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The consumer keeps a record of which message IDs it already processed, in the same database transaction as the work it does. If the same message shows up again, it checks that record first and skips redoing the work, so a duplicate delivery has no extra effect.

open as a page

Because a saga is a sequence of independently committing local transactions rather than one ACID transaction, other operations can observe and act on data mid-saga. What isolation anomalies does this create, and what techniques mitigate them without resorting to a distributed lock?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

While a saga is still in progress, other users or processes can see and act on the half-finished data, which can cause weird bugs like double-spending a discount or acting on data that later gets reversed. Teams fix this with tricks like marking records 'pending' so others know to be careful with them, instead of using a database-wide lock.

open as a page