When implementing a saga across Order, Payment, and Shipping services, what's the difference between a choreography-based saga and an orchestration-based saga, and what would push you to pick one over the other?
answer
- choreography = events, no boss
- orchestration = central coordinator/state machine
- choreography decouples but hides the big picture
- orchestrator needs its own durable state
- switch to orchestration past ~4-5 steps
basics
~10 sChoreography: each service listens for events and reacts on its own, no one's in charge. Orchestration: one central coordinator tells each service what to do next, step by step.
solid answer
~50 sIn choreography, each service publishes a domain event when it finishes its step (OrderCreated, PaymentCharged) and subscribes to the events it cares about, deciding for itself what to do next — there's no central authority, just a web of event producers and consumers. In orchestration, a dedicated orchestrator (a state machine or workflow engine) explicitly calls each service in sequence, tracks saga state, and issues compensating commands on failure. Choreography scales well for a small number of steps because it avoids a coordinator bottleneck and keeps services decoupled, but as the number of steps grows, the flow logic gets smeared across every service, making the overall business process hard to see, test, or change in one place, and it's easy to create implicit cyclic dependencies between event consumers. Orchestration centralizes that complexity into one visible workflow definition, is easier to monitor and version, but introduces a coordinator that must itself be highly available and becomes a point of coupling all participants depend on.
go deeper
Should be able to describe the basic difference — events vs a central coordinator — with a simple example, without needing to weigh trade-offs deeply.
Should explain the readability/observability trade-off and know that choreography tends to get messy as steps increase, prompting a switch to orchestration.
Should discuss concrete coupling implications (command contracts vs event contracts), how each style handles compensation ordering, and how to add observability to a choreographed saga.
Should be able to recommend an org-appropriate default (e.g., orchestrate within a team, choreograph across teams), discuss orchestrator HA/durability design, and know when a hybrid approach reduces both blast radius and cognitive load.
## Where the know-how lives A saga is a sequence of local transactions with compensations, and the coordination style — **choreography** or **orchestration** — is about how the next step gets triggered and where the "know-how" of the overall business process lives. ## Choreography Choreography is fully decentralized. Each service, after committing its own local transaction, publishes a domain event describing what happened — `OrderCreated`, `PaymentCharged`, `PaymentDeclined`, `InventoryReserved` — usually onto a message broker like Kafka or RabbitMQ. Other services subscribe to the events relevant to them and decide independently what to do in response. In an order flow: - the Payment service subscribes to `OrderCreated` and, on receiving it, attempts a charge, then emits `PaymentCharged` or `PaymentDeclined`; - the Inventory service subscribes to `PaymentCharged` and attempts to reserve stock, emitting `InventoryReserved` or `InventoryOutOfStock`; - and so on down the chain. Nobody holds a global view of "the saga" as an object — the saga's behavior emerges from each service's local event-handling logic. This keeps services maximally decoupled (Payment doesn't need to know Inventory exists, only that it reacts to events that happen to originate there) and avoids any single component becoming a bottleneck or SPOF for the whole flow. ## Orchestration Orchestration flips that: a dedicated component — often called a **saga orchestrator**, workflow engine, or process manager (tools like Camunda, Temporal, AWS Step Functions, or a hand-rolled state machine) — owns an explicit definition of the steps and their order. 1. It sends a command to Payment ("charge $50"), waits for a reply or a timeout. 2. It then sends a command to Inventory ("reserve 2 units"), and so on, tracking the saga's current state itself. 3. On a failure, the orchestrator knows exactly which steps have completed and issues the compensating commands in reverse order. The participating services are unaware of each other; they only implement command handlers the orchestrator calls. ## The core trade-off The core trade-off is where complexity lives and how coupling is shaped. | Style | What it buys | What it costs | |---|---|---| | Choreography | avoids a central coordinator and its availability/scaling concerns, and each service change is local | as the number of steps grows past three or four, the overall process is no longer readable from any one place; a new engineer has to trace event subscriptions across five service codebases to understand what happens when a payment is declined, and it's easy to accidentally create cyclic or fan-out event chains that are hard to test end-to-end | | Orchestration | makes the business process a first-class, versionable artifact you can read, test, and monitor in one place — you get saga-level observability (which step is a given order stuck on right now) essentially for free | the orchestrator itself needs to be made highly available and durable (its own state must survive crashes, usually via an event-sourced or persisted state machine), and every participating service now has an implicit dependency on the orchestrator's command contract, which is a form of coupling even if the services don't call each other directly | ## In practice In practice, teams commonly start with choreography for two or three services because it's cheap and decoupled, and migrate to orchestration once the saga grows past roughly four or five steps or once operators start asking "why is this order stuck" and there's no single place to look. It's also common to mix styles: a large system might choreograph high-level domain events between bounded contexts while using a lightweight orchestrator internally within one team's saga. ## The failure mode that shows the difference A concrete failure mode that shows the difference: - **In a choreographed saga**, if the Inventory service is redeployed and temporarily stops consuming `PaymentCharged` events, orders silently stall in a "paid but not reserved" state with no single component even aware anything is stuck — you find out from a customer complaint or a reconciliation job. - **In an orchestrated saga**, the orchestrator's own state store shows the exact order sitting at the "waiting for inventory reservation" step with a timestamp, so the stall is directly queryable and alertable. Orchestration engines like Uber's Cadence/Temporal for trip and payment flows are commonly cited real-world examples adopted specifically because choreographed event chains became too hard to reason about at scale; conversely, many event-driven e-commerce platforms deliberately keep small sagas (like cart-abandonment cleanup) choreographed because a dedicated orchestrator for a two-step flow is not worth the operational overhead.
- How do you monitor 'where is this saga stuck' in a choreography-based design, given there's no central orchestrator?You typically build a separate saga-tracking read model that subscribes to all the relevant events and materializes progress per saga instance, or you rely on distributed tracing (correlation IDs propagated through events) plus dashboards/alerts on services that stop consuming. Some teams add a lightweight 'saga log' service purely for observability without giving it control authority, which is effectively a partial orchestrator bolted on for visibility only.
- Does orchestration make the orchestrator a single point of failure for in-flight sagas?It can, unless the orchestrator persists its state durably (e.g., event-sourced state machine, or a workflow engine like Temporal that checkpoints progress) so that a crash and restart resumes exactly where it left off rather than losing track of in-flight sagas; the orchestrator process being stateless and reconstructing state from a durable log is the standard mitigation.
- Can choreography and orchestration be combined in the same system?Yes — a common pattern is orchestrating steps within one team's bounded context while choreographing high-level domain events between bounded contexts, so each team gets local visibility without forcing a single global orchestrator across organizational boundaries.
Choreography is like a group of dancers who each know their own cues from the music and react to each other without a director — beautiful when small, chaotic to debug when the troupe grows. Orchestration is like a conductor calling out each section's entrance from a single score — you can always point to the score to see what should happen next, but the whole performance depends on the conductor showing up.
saying these in an interview costs you the question
- Claims orchestration means services call each other directly rather than via commands from the orchestrator
- Thinks choreography has no coupling at all between services
- Can't explain how you'd debug a stuck saga in whichever style they claim to prefer
- Says one style is always strictly better regardless of number of steps
- Forgets the orchestrator itself needs durable state to survive a crash