An order-placement flow in a CQRS/event-driven system needs to reserve inventory, charge payment, and update a read-model 'order status' projection, each owned by a different service with its own write model. No single ACID transaction spans all three. What coordination pattern keeps this consistent, and what happens if the payment step fails after inventory has already been reserved?
answer
- local transactions plus compensations
- choreography vs orchestration
- process manager holds durable saga state
- compensation is business logic, not a DB rollback
- transactional outbox avoids the dual-write gap
basics
~20 sUse a saga: a sequence of local transactions where each step tells the next one to proceed, coordinated either by a central orchestrator or by services reacting to each other's events. If payment fails after inventory was reserved, the saga runs a compensating action - releasing the reserved inventory - to undo that earlier step.
solid answer
~40 sA saga breaks a cross-service business transaction into a series of local transactions, each committed independently, with a compensating transaction defined for each step to undo its effect if a later step fails. Two styles: choreography (each service publishes an event and the next reacts, no central brain, good for short flows) and orchestration (a process manager explicitly issues commands and tracks saga state, better for complex or branching flows and gives one place to observe progress). Here: reserve-inventory, then charge-payment, then update the read-model status. If charge-payment fails, the orchestrator invokes the compensation for the already-committed step - release-inventory - and marks the order's read-model status 'failed' rather than 'confirmed'. This restores business-level consistency even though there was never a single atomic transaction spanning all steps.
go deeper
Not generally expected to design sagas; should recognize that 'no single transaction spans multiple services' is a real constraint in this kind of system.
Should be able to describe what a saga is and give a simple choreography example, even without yet designing orchestration state machines.
Should design the saga - steps, compensations, orchestration vs choreography choice - for a real multi-step flow, and reason about idempotency and the dual-write/outbox problem.
Should set standards for which coordination pattern (saga, orchestration engine, or avoiding cross-service transactions via redesign) fits which class of business flow across the org, and evaluate build-vs-buy on workflow engines.
## What a saga is A **saga** decomposes a business transaction that spans multiple services or modules into a sequence of local transactions, each of which commits independently within its own owning service, plus a **compensating transaction** defined for every step that has a real, externally-visible effect. In this order-placement example, the sequence is: 1. `reserve-inventory` 2. `charge-payment` 3. and finally a projection update to the order-status read model Someone — either a central process manager (**orchestration**) or the participating services reacting to each other's published events (**choreography**) — has to track which steps have completed and drive the flow forward or backward. ## Orchestration against choreography | Style | How it drives the flow | Where it fits | |---|---|---| | Orchestration | explicitly models the saga as a state machine: a durable process manager holds the current step, issues the next command, and on failure looks up and issues the compensations for whatever already succeeded | tends to win for anything with branching logic, timeouts, or more than two or three steps, because it gives one place to see and debug the flow | | Choreography | instead has each service publish a domain event when it finishes its step, and the next service in line subscribes and reacts, with no single component holding the overall picture | stays simpler for very short chains but gets hard to reason about as steps multiply, since no one component owns the 'is this saga stuck?' question | ## Why the pattern exists This pattern exists because cross-service ACID transactions — classically implemented via **two-phase commit** — don't work well at the scale and topology CQRS and microservices operate at. - Two-phase commit requires every participant to hold locks and stay available until a coordinator finalizes the outcome, which is fundamentally at odds with independently-deployed services, asynchronous messaging, and horizontal scaling; a coordinator outage or a slow participant blocks everyone. - Business processes, however, still need multi-step consistency — an order genuinely shouldn't end up charged with no inventory reserved, or reserved with no charge — so sagas provide a way to get eventual business-level consistency across boundaries that will never share a real transaction. ## The trade-off The trade-off is that a saga gives up true atomicity for availability and loose coupling. - **Intermediate states become visible to the outside world** during the flow — inventory is decremented against available stock while payment is still pending, which a naive read of 'stock remaining' could observe. - **Compensations are not database rollbacks;** they are business logic that must be explicitly designed as the semantic opposite of the original step, and they can have residual effects of their own (releasing a hotel hold might trigger a notification email that can't be un-sent, for instance). - **Building and maintaining a saga's state machine and its compensations is real, ongoing engineering work** per business flow, not something that comes for free from choosing CQRS. ## Failure modes Several failure modes are specific to sagas. - **A compensation can itself fail** — releasing inventory might hit a transient error — which typically needs its own retry policy and, ultimately, a dead-letter queue plus manual intervention if retries are exhausted. - **An orchestrator process crashing mid-flow** needs durable, persisted saga state (not just in-memory) so it can resume from the last recorded step on restart, and every step handler needs to be idempotent so a resumed or retried step doesn't double-charge or double-reserve. - **Choreography is additionally vulnerable to ordering issues** — events processed out of sequence can leave participants in an inconsistent view of where the saga actually is — which is usually mitigated with correlation ids and explicit step sequencing. - **The 'dual write' problem** is a subtler, very common bug: a step commits its local database change but then fails to publish the event that should trigger the next step, silently stalling the saga; the standard fix is a **transactional outbox**, writing the event to an outbox table in the same local transaction as the state change, then reliably publishing from that outbox afterward. ## Where the pattern comes from This is a well-documented, named pattern — Chris Richardson's microservices.io catalog describes it under exactly this order-processing example (reserve inventory, charge card, ship), and the underlying idea traces back to the 'process manager' pattern in Hohpe and Woolf's Enterprise Integration Patterns. In practice, teams implement the orchestration style using durable workflow engines such as Temporal, AWS Step Functions, or Camunda, specifically because they provide the durable-state, retry, and idempotency machinery a hand-rolled saga orchestrator would otherwise need to reinvent.
- Why not just use a distributed two-phase-commit transaction across all three services instead of a saga?Two-phase commit requires all participants to hold locks and stay available until the coordinator finalizes, which doesn't work well with message brokers or independently-deployed services, doesn't scale across network boundaries, and creates a coordinator single point of blocking. Sagas trade strict atomicity for availability and loose coupling appropriate to that kind of architecture.
- What's the difference between a saga's compensating transaction and just rolling back a database transaction?A database rollback undoes an uncommitted transaction with no external side effects. A compensating transaction semantically reverses an already-committed, externally-visible action - 'release the reservation', 'refund the charge' - and that reversal is itself business logic that may have side effects, delays, or partial failures of its own; it's not automatic.
- How do you make saga steps safe to retry after a crash mid-flow?Each step handler must be idempotent, keyed by a saga or step id so replays don't double-charge or double-reserve, and the orchestrator's own progress must be persisted durably - typically in a saga-state table or a durable workflow engine - so it can resume from the last recorded step instead of restarting the whole flow from scratch.
Like booking a multi-leg trip yourself - you first hold the hotel room, then book the flight, then reserve the rental car; if the flight booking fails, there's no automatic magic undo, you have to explicitly cancel the hotel hold you already made. That manual 'undo my earlier booking' step is the compensating transaction.
saying these in an interview costs you the question
- proposes a single distributed ACID transaction across services as the default fix
- calls a compensating transaction a 'rollback' with no acknowledgment of visible intermediate side effects
- no mention of what happens to already-completed steps when a later step fails
- doesn't distinguish orchestration from choreography
- assumes saga steps never need to be idempotent or retried