skip to content

In an event-sourced CQRS system, a single business workflow such as placing an order needs to update several independent aggregates -- say Order, Inventory, and Payment -- each with its own consistency boundary. What role does a process manager (saga) play here, and how does it drive the workflow forward?

level: middleimportance: should knowfreq 70%

answer

  1. saga = stateful workflow reacting to events, emitting the next command
  2. each aggregate keeps its own transaction boundary
  3. orchestration = central coordinator, choreography = no coordinator
  4. compensation is a new business action, not an undo
  5. saga state must be durable and idempotent to survive a crash

basics

~20 s

A saga listens for events, and each time something happens it decides the next command to send to the next aggregate. It's like a conductor reacting step by step, instead of one big all-or-nothing transaction across everything.

solid answer

~40 s

A process manager, or saga, is a stateful component that subscribes to events from multiple aggregates and drives a multi-step business workflow by issuing the next command each time it observes a relevant event, and by issuing compensating commands to undo earlier steps if a later step fails. Because each aggregate stays its own transactional, single-append consistency boundary, cross-aggregate business rules can't rely on a shared database transaction; the saga replaces that with a sequence of local transactions plus compensations. Sagas come in two flavors: orchestration, where a central process manager owns the workflow state and issues every command, and choreography, where aggregates react directly to each other's events with no central coordinator.

go deeper

for a junior

Can describe a saga as something that reacts to events and sends the next command in a sequence.

for a middle

Can name choreography versus orchestration and give a concrete compensation example.

for a senior

Can design a concrete saga state machine including timeout handling and idempotency keys for its commands.

for a principal

Decides at the org level when a saga is warranted versus redesigning aggregate boundaries to avoid needing one, and sets standards for saga durability, observability, and tracing across teams.

## The problem it solves A **saga** (or **process manager**, the terms are used near-interchangeably in this context) exists to solve a specific structural problem: Event Sourcing paired with CQRS typically keeps each aggregate small, with one aggregate instance mapping to one event stream and one optimistic-concurrency-checked append per command, so that writes stay fast and uncontended. That design deliberately rules out wrapping a business workflow that spans several aggregates -- reserve inventory, then charge payment, then confirm the order -- inside one ACID database transaction, because each aggregate is its own transaction boundary. The saga pattern, originating from Garcia-Molina and Salem's 1987 paper on long-lived transactions, replaces that single distributed transaction with a sequence of local transactions, each against one aggregate, coordinated by explicit logic that reacts to outcomes and, when something downstream fails, issues compensating actions rather than an automatic rollback. ## How it drives the workflow forward Concretely, a process manager subscribes to the domain events published by the participating aggregates and maintains its own durable state describing where a given workflow instance currently is. - On `OrderPlaced` it issues a `ReserveInventory` command to the Inventory aggregate; - on `InventoryReserved` it issues a `ChargePayment` command to the Payment aggregate; - on `PaymentCharged` it issues `ConfirmOrder` back to Order; - but if it instead observes `PaymentFailed`, it issues a compensating `ReleaseInventory` command and a `CancelOrder` command rather than trying to reverse anything at the database level, because there is no shared transaction to roll back. ## Orchestration versus choreography There are two structurally different ways to implement this. | Style | How it plays out | |---|---| | **Orchestration** | A dedicated process-manager component owns the entire workflow's state machine and is the only thing that issues commands to participants, which makes the workflow's logic and current progress easy to find, test, and monitor in one place, at the cost of that component needing knowledge of every participant and becoming a central piece every workflow change must go through. | | **Choreography** | There is no dedicated coordinator: each aggregate's handler reacts directly to events published by the others, which keeps components decoupled and avoids a central bottleneck, but makes the overall workflow harder to see or reason about as a whole, since the logic is scattered across many independent handlers and typically requires correlation ids and distributed tracing to debug. | ## The trade-off The trade-off, either way, is complexity moved rather than removed: instead of one aggregate boundary and one transaction, the team now owns - **saga state persistence** (the saga's own progress must survive a crash, so it is itself often event-sourced or at least durably checkpointed), - **idempotent command handling** on every participant (so a saga that crashes and resumes doesn't double-issue a command like 'charge payment' a second time), - and **timeout handling** for steps that never acknowledge. Compensations also aren't a true undo: once a payment has actually been captured, the compensating action is a refund, a distinct business operation with its own delay and its own possible failure, not a mirror-image rollback of the original charge. ## Where it shows up A widely referenced real-world shape of this pattern is the classic order/inventory/payment saga described in the microservices literature (for example Chris Richardson's Saga pattern write-ups), and it appears concretely in domains like travel booking, where reserving a flight, a hotel, and a rental car are each independent aggregates or services, and a failure in any one step triggers cancellation compensations in the steps that already succeeded, rather than the system attempting one atomic cross-service transaction across three different providers. ## Timeouts Timeouts deserve their own mention because they're the failure mode teams most often forget until it bites them in production: a saga that issues `ReserveInventory` and waits for `InventoryReserved` has to decide what happens if that event never arrives at all, whether because the message was lost, the Inventory service is down, or it simply takes far longer than expected. Without an explicit timeout, the workflow instance sits stuck indefinitely, invisible until a customer complains their order never confirmed. A robust saga schedules a timer alongside issuing the command, and if the expected response event hasn't arrived by the deadline, treats that exactly like a failure response, triggering the same compensation path, sometimes after a small number of retries with backoff if the failure looks transient rather than permanent. ## Correlation Correlation is the other detail that separates a working saga implementation from a broken one: every command the saga issues and every event it reacts to needs to carry a shared **correlation id**, usually the saga instance's own id, so that when `InventoryReserved` comes back, the saga can find the correct in-flight workflow instance it belongs to, especially when many instances of the same workflow are running concurrently for different orders. Some teams build sagas by hand on top of a message bus with this bookkeeping done explicitly; others use a dedicated orchestration engine, such as **Camunda** or **AWS Step Functions**, that provides durable workflow state, timeout handling, and correlation as built-in features rather than something each team reimplements per saga.

  • What's the difference between choreography and orchestration sagas, and when would you pick one over the other?
    Choreography has each aggregate react directly to other aggregates' events with no central coordinator; it suits simple, short workflows and keeps components loosely coupled. Orchestration has a dedicated process manager own the workflow's state and issue every command; it suits longer, more complex workflows where you want one place to see progress, change logic, and monitor the whole thing, at the cost of that component knowing about every participant.
  • How do you make sure a saga doesn't issue a duplicate 'charge payment' command if it crashes and replays events after restart?
    The saga's own progress must be persisted durably, typically itself event-sourced with a marker for which step it last completed, and every downstream command handler should be idempotent keyed on a saga or correlation id, so that if the saga re-sends a command it already issued before crashing, the handler recognizes it as already applied and takes no further action.
  • Why can't compensations always be a true undo of the original action?
    Some effects, like sending a confirmation email or capturing a card payment, cannot literally be un-happened. The compensating action becomes its own distinct business operation, such as a refund or a cancellation notice, with its own timing, its own possible failure, and sometimes its own manual review, rather than a symmetric reversal of the original step.

Like a wedding planner coordinating separate vendors -- caterer, venue, band -- who each run their own independent business: the planner reacts to each vendor's confirmation by booking the next one, and if the band cancels, calls the caterer to adjust rather than being able to roll back the whole wedding in one atomic step.

saying these in an interview costs you the question

  • proposes wrapping all the participating aggregates in one shared database transaction
  • calls saga compensation a 'rollback' as though it erases history
  • no mention of idempotency or duplicate-command risk if the saga restarts
  • doesn't distinguish the saga's own internal consistency from the eventual consistency of the overall cross-aggregate workflow
  • treats sagas as tied to one specific product or framework rather than as a general pattern

context