An order-processing flow spans three separate services (inventory, payment, shipping), each with its own database. Why can't you just wrap the whole flow in one ACID database transaction, and what pattern is typically used instead?
answer
- local transactions, not one big transaction
- compensate, don't rollback
- no cross-service locks
- trades atomicity/isolation for availability
basics
~20 sEach service owns its own database, so one transaction can't span them all. A saga runs a sequence of local transactions instead, one per service; if a later step fails, it runs compensating actions to undo the already-completed earlier steps.
solid answer
~50 sA saga is a sequence of local transactions, each owned by one service, coordinated so a business operation either completes fully or is undone. In microservices, wrapping cross-service ACID transactions (e.g., two-phase commit) means holding database locks across network calls and coupling services' uptime together, which doesn't scale. A saga trades strict atomicity for availability: each step commits immediately in its own service's database. If a later step fails, the saga doesn't roll back via the database - it runs compensating transactions, semantically inverse operations (e.g., 'cancel reservation', 'refund payment'), to undo the effects of already-completed steps. Coordination is either choreography (services react to each other's events) or orchestration (a central coordinator issues commands). The cost: intermediate, partially-applied states are visible to the rest of the system before the saga finishes, so isolation is weaker than in a single ACID transaction.
go deeper
Should recognize that microservices with separate databases can't use one big transaction, and give the one-line idea: sequence of local transactions plus compensations on failure.
Should describe the compensation mechanism concretely with an example (charge -> refund) and know the two coordination styles exist, even without deep trade-off analysis.
Should articulate the atomicity/isolation trade-off explicitly, discuss idempotency of compensations, and reason about which steps are and aren't compensable.
Should connect the pattern choice to system-wide architecture decisions - when sagas are the wrong tool entirely, and how to design domain boundaries to minimize saga complexity.
## The problem it solves The **saga pattern** solves a problem that appears the moment you split a monolith with one database into services that each own their own database: how do you keep a multi-step business operation consistent when no single database transaction can span all the steps? A classic example is an e-commerce checkout that must: - **reserve inventory**, - **charge a payment**, - **schedule shipping**, where each of those actions lives in a different service with its own datastore. In a monolith you'd wrap all three in one ACID transaction and either commit everything or roll back everything. Across services, that option effectively disappears (or becomes prohibitively expensive), so the saga pattern replaces one big transaction with a sequence of small **local transactions**, `T1`, `T2`, ..., `Tn`, each of which is a normal ACID transaction inside a single service's own database. ## How a saga actually runs Mechanically, a saga executes these local transactions in order. 1. Each one **commits durably and immediately** when it runs - there's no distributed lock held between steps, and no waiting for a global 'prepare' phase. 2. If every step succeeds, the saga is done and the system is in a consistent, desired end state. 3. If step k fails (say, payment processing fails after inventory was already reserved), the saga does not roll the database back the way a single transaction would; instead, it runs **compensating transactions** `C(k-1)`, `C(k-2)`, ..., `C(1)` for each of the steps that already committed, in reverse order, to semantically undo their effects. 'Semantically undo' is a deliberate choice of words: a compensating transaction is not a literal database rollback, because the original transaction already committed and other parts of the system may have already observed and acted on that committed state. So: - 'reserve 1 unit of inventory' is compensated by 'release 1 unit of inventory' (increment a counter back), not by physically erasing a row; - 'charge $50' is compensated by 'refund $50', a new, separate transaction with its own record, not an undo of the original charge. ## Why the pattern exists The reason this pattern exists is **availability and autonomy**. Microservices are valuable specifically because each service can deploy, scale, and fail independently, with its own database technology and schema. A distributed ACID transaction across service boundaries breaks that independence: it typically needs a transaction coordinator, and every participant has to hold locks on its rows while waiting for every other participant to be ready to commit. If one service is slow or down, all the others sit blocked holding locks - exactly the kind of cascading unavailability microservices are meant to avoid. Sagas sidestep this by never requiring a resource to be locked across a network call - each local transaction is fast and self-contained, and the 'undo' work happens through additional, equally fast local transactions rather than through held locks. ## What the pattern gives up The trade-off is that a saga gives up two things ACID transactions normally give you for free: - **atomicity** - there is a window, however short, where step 1 has committed and step 2 hasn't yet, so the overall operation is 'in progress' rather than atomic; - **isolation** - during that window, another part of the system (a report, another workflow, a customer checking their order status) can read the partially-completed state, e.g., inventory reserved but payment not yet confirmed. Handling this requires either accepting the anomaly, designing the domain to tolerate it (e.g., marking the order 'pending' until the saga completes), or applying countermeasures like **semantic locks**. Sagas also push real complexity onto the team: - every step needs a compensating action; - compensations must be safe to retry (**idempotent**) because failures during the saga itself can force a compensation to be re-run; - and not every action is compensable (an email that's already been sent, or a payment already spent by the recipient, can't be truly undone, only mitigated). ## A failure mode in production A concrete failure mode in production: a shipping-cancellation compensating transaction gets sent, times out because the shipping service is briefly overloaded, and the saga orchestrator retries it. If cancellation isn't idempotent, a second retry might, for example, double-refund a shipping fee or double-notify a warehouse. This is why saga implementations in practice - workflow engines like AWS Step Functions or Temporal orchestrating multi-step business processes, or e-commerce platforms with checkout sagas - invest heavily in **idempotency keys** and durable **saga logs**, so that every step and every compensation can be safely retried without corrupting state, and so a crashed coordinator can resume exactly where it left off by replaying the log.
- Does every step in a saga need a compensating transaction?Not the very last step, since once the saga's final action succeeds there's nothing left to undo. Read-only or non-side-effecting steps also don't need one. But every step that mutates state and could be followed by a later failure needs a compensation, and steps that are truly irreversible (like sending a physical package) need to be sequenced last or handled with a separate mitigation strategy rather than a saga.
- What happens if a compensating transaction itself fails?This is the hardest operational problem with sagas: compensations are expected to always eventually succeed, so systems retry them (requiring idempotency) and often keep retrying with backoff or escalate to a dead-letter queue with manual intervention. Some designs make compensations simpler and more reliable than the original action specifically so they're less likely to fail - e.g., a refund is usually easier to guarantee than the original charge.
- Can a saga guarantee the same consistency as a single ACID transaction?No - it guarantees eventual consistency, not immediate atomicity or isolation. The system will end up in a consistent state (either all steps effectively applied or all compensated), but there's a window where the state is partially applied and visible to others, which a true ACID transaction never exposes.
Booking a multi-leg trip yourself (flight, hotel, rental car) instead of using one travel agency that holds all three on hold simultaneously: you book each one separately, and if the rental car falls through, you cancel the flight and hotel yourself rather than the whole trip vanishing atomically.
saying these in an interview costs you the question
- Says a saga rolls back the database like a transaction abort
- Doesn't mention compensating transactions at all
- Thinks a saga gives full ACID isolation across services
- Assumes compensations are just the original operation run backwards with no special design
- Can't explain why cross-service ACID transactions are avoided in microservices