skip to content

You're implementing a saga for an e-commerce checkout that touches Order, Payment, and Shipping services. Compare a choreography-based saga (services react to each other's events) with an orchestration-based saga (a central coordinator issues commands) for this flow - how does each work mechanically, and what are the trade-offs as the number of participating services grows?

level: middleimportance: must knowfreq 85%

answer

  1. events vs commands
  2. no coordinator vs central coordinator
  3. decoupled but hard to trace vs visible but extra component
  4. grows harder to debug as steps increase (choreography)

basics

~20 s

In choreography, each service listens for events and decides what to do next itself - no central boss. In orchestration, one coordinator service tells each step what to do and waits for the result, like a conductor directing an orchestra.

solid answer

~50 s

Choreography-based sagas have each participant publish domain events after its local transaction commits, and other participants subscribe to relevant events and react by running their own local transaction (and possibly publishing further events). There's no central coordinator - the saga's logic is distributed across all the services' event handlers. Orchestration-based sagas have a dedicated orchestrator that knows the whole saga's steps, sends explicit commands to each participant, and reacts to their responses to decide the next command, including triggering compensations on failure. Choreography scales better initially, has no single point of coordination, and keeps services decoupled, but as the number of steps grows it becomes hard to see the overall flow, debug, and avoid cyclic event dependencies. Orchestration centralizes the flow into one place - easy to visualize, test, and modify - at the cost of an extra component and a service that needs to know about every participant, which can become a bottleneck or a hidden coupling point.

go deeper

for a junior

Should be able to describe the basic difference: events vs a central coordinator, with a simple example of each.

for a middle

Should explain the mechanism of both (event subscription chains vs command/response) and name at least one concrete trade-off each way.

for a senior

Should reason about when to choose each based on saga size/complexity, and know how orchestrators survive crashes (durable state).

for a principal

Should discuss hybrid approaches, the organizational/team-ownership implications, and how the choice affects long-term evolvability of the system.

## Two ways to coordinate a saga Choreography and orchestration are the two ways to coordinate the sequence of local transactions and compensations that make up a saga, and the choice mostly comes down to how much you want the coordination logic **centralized** versus **distributed** across the participating services. ## Choreography: everyone reacts to events In a choreography-based saga, there's no dedicated coordinator process. Each service completes its own local transaction and then publishes a **domain event** describing what happened - for example, the Order service commits an order in 'pending' state and publishes `OrderCreated`. Other services subscribe to the events relevant to them: 1. the Inventory service listens for `OrderCreated`, reserves stock, commits its own local transaction, and publishes `InventoryReserved` (or `InventoryReservationFailed`); 2. the Payment service listens for `InventoryReserved`, attempts to charge the customer, and publishes `PaymentCompleted` or `PaymentFailed`. The saga's overall logic - the sequence of which step follows which - is nowhere written down as a single artifact; it emerges from the sum of all the individual event subscriptions across every service. Compensation works the same way in reverse: if `PaymentFailed` is published, the Inventory service (which subscribed to that event too) runs its own compensating transaction to release the reserved stock. ## Orchestration: one process issues commands In an orchestration-based saga, one process - the **orchestrator** - explicitly owns the saga's definition. It sends a command to the first participant ('reserve inventory'), waits for a response, and based on that response sends the next command ('charge payment'), continuing step by step. If a step reports failure, the orchestrator is responsible for issuing the compensating commands to every participant that already succeeded, in the right order. The orchestrator typically persists its own state (which step it's on) so it can resume after a crash, often implemented with a durable state machine or workflow engine. ## Why both styles exist The reason both styles exist is that they optimize for different things. | | Choreography | Orchestration | |---|---|---| | Saga definition | emerges from event subscriptions | one process explicitly owns it | | Coupling | maximally decoupled | knows about every participant | | Debugging | trace event subscriptions across many codebases | flag the step it is waiting on | Choreography keeps services maximally decoupled: a new service can join the saga just by subscribing to existing events, without any existing service needing to change or even know the new service exists. This is attractive early in a system's life, when the flow is short (2-3 steps) and the team wants to avoid introducing a new central component. But this benefit inverts as the saga grows: - with five or more participants and several failure branches, there is no single place to read and understand 'what does checkout actually do,' which makes onboarding, debugging, and testing significantly harder - you have to trace event subscriptions across many codebases to reconstruct the flow; - it's also easy to accidentally create cyclic or ambiguous event dependencies, where two services each wait on an event the other is supposed to produce first. Orchestration trades that decoupling for explicit visibility and control: the entire saga's happy path and every compensation branch live in one place, are testable as a unit, and are straightforward to modify (add a step, reorder steps) without touching the participant services' internals beyond their command handlers. The cost is that the orchestrator becomes a piece of infrastructure that must itself be highly available and correctly handle its own crashes and retries (usually solved by persisting saga state to a durable store and resuming from there), and it necessarily knows about every participant, which can become a coupling and scaling bottleneck if it also has to encode business rules that arguably belong to the participants. ## A failure mode that shows the difference A concrete failure mode that shows the difference: in a choreography saga, if a new team adds a service that subscribes to `InventoryReserved` but forgets to publish a completion or failure event, the whole saga silently stalls with no owner noticing, because no component is 'watching' the entire flow - discovering this in production often requires distributed tracing across services. In an orchestration saga, the same kind of bug (a participant not responding) is easier to detect, because the orchestrator has an explicit timeout and step it's waiting on; it can flag 'stuck on step 3' directly. ## What real systems settle on In practice, many real systems use a **hybrid**: - **choreography** for short, stable, low-branching subflows (e.g., publishing an `OrderPlaced` event that a few independent services react to for notifications or analytics, where failure doesn't need compensation), - **orchestration** for the core business-critical saga with several steps and compensation branches (e.g., checkout, loan approval, travel booking), where correctness and observability of the full flow matter more than minimizing coupling. Once a workflow has several steps each with its own compensation, an explicit, inspectable orchestrator tends to win over an implicit, event-driven web of subscriptions.

  • At what point does a team typically switch from choreography to orchestration for a saga?
    Usually when the number of participating services or branching failure paths grows past roughly 3-4 steps, or when debugging 'what happened to this order' starts requiring tracing events across multiple codebases. The tipping point is really about cognitive load and observability, not a hard number - if nobody can explain the whole flow without reading five services' event handlers, orchestration usually pays for itself.
  • How does an orchestrator survive its own crash mid-saga?
    It persists the saga's current state (which step completed, which is pending) to durable storage before and after each command, so on restart it can read that state and resume from the last known point rather than starting over or losing track. Workflow engines build this durability in as a core feature rather than something each team re-implements.
  • Can choreography and orchestration be mixed in the same system?
    Yes, and it's common - some subflows use choreography (independent, low-stakes fan-out like notifications or analytics) while the core business transaction uses orchestration for its steps and compensations. The two aren't mutually exclusive at the system level, only within a single saga's coordination logic.

Choreography is like a group of dancers who each know their own cues and react to the music and each other's movements with no conductor; orchestration is a conductor calling out each section's entrance explicitly from a single score.

saying these in an interview costs you the question

  • Claims choreography always scales better than orchestration
  • Can't explain what the orchestrator actually persists or why
  • Thinks orchestration means the coordinator holds a distributed lock
  • Believes choreography requires no event ordering or subscription design
  • Doesn't recognize that compensation logic needs a home in either style

context