You're designing the communication for a saga spanning five services with compensating actions on failure. Under what conditions would you choose choreography versus orchestration, and what does each cost you in observability and evolvability as the system grows?
answer
- failure-logic density, not service count, drives the choice
- choreography: cheap participant growth, expensive sequence changes
- orchestration: cheap sequence changes, expensive participant growth
- hybrid: orchestrate the transactional core, choreograph the fan-out
basics
~20 sIf the steps are simple and mostly independent, letting each service react on its own (choreography) keeps things flexible as you add more services. If the steps are tightly sequenced with real error handling and rollback logic, having one coordinator (orchestration) keeps the whole process understandable as it grows, even though that coordinator becomes something everyone has to touch.
solid answer
~1 minThe deciding factor isn't service count alone but process complexity: how many steps have failure/compensation logic, how often does that logic change, and how many teams need to modify the sequence. Choreography scales well when steps are largely independent side effects with simple or no compensation, since new participants can subscribe without touching existing services, but as compensation logic grows, choreography starts requiring a 'saga per event' mental model spread across every participant, and observability degrades because no service holds the end-to-end state, forcing teams to build separate tooling (event-sourced process trackers, distributed tracing correlation) just to answer 'where is this saga right now.' Orchestration scales well when the process has real branching, compensation, and needs a single source of truth for state, since that logic stays in one reviewable place, but the orchestrator becomes a shared-change bottleneck once many unrelated teams need to add steps, and it becomes a single component whose outage or bug can halt the entire process rather than just one step. In practice, mature systems often end up hybrid: orchestration for the core transactional saga (payment, inventory, shipping) with real compensation, and choreography for independent, non-transactional fan-out (analytics, notifications, recommendations) hanging off events the core saga publishes.
go deeper
Not typically expected to reason at this level; a junior doing well here would at least recall that adding services is easier in choreography and that orchestration keeps logic in one place.
Should recall the basic scaling directions (choreography scales participant growth, orchestration scales sequence complexity) even without deep justification.
Should explain why observability tooling for choreography (tracing, saga trackers) has to be deliberately built rather than assumed, and should be able to justify picking one pattern for a given saga's failure-logic density.
Should propose and justify a hybrid architecture (orchestrate the transactional core, choreograph independent fan-out), reason about organizational bottleneck symptoms and mitigations, and discuss how the right boundary shifts as the system and org grow rather than treating the choice as fixed at design time.
## The axis that actually decides it The choice between choreography and orchestration doesn't stay fixed as a system grows; the right answer at five services with two failure paths is often the wrong answer at fifteen services with a dozen failure paths, and reasoning about that trajectory is the actual principal-level question here, more than picking one pattern as universally correct. The axis that matters most is not the number of services involved but the **density and volatility of the failure and compensation logic**: - how many of the steps in the saga need a defined "undo" action; - how often does the sequence or its error handling change; - how many separate teams own pieces of that sequence. ## What choreography scales gracefully Choreography scales gracefully along one specific dimension: **adding participants** that react to existing events costs nothing to the services already in the flow. If five services already form a saga via events and a sixth team wants to add a new side effect triggered by one of those existing events, no existing service's code changes at all. This is a genuine, durable advantage as an organization grows, because it means the core saga's owning teams never become a bottleneck for other teams wanting to hook into the process. But this advantage specifically applies to independent side effects; it does not extend to changes in the core sequence or compensation logic itself. If the failure handling for the saga grows more complex, say a new rule that "if shipping fails after inventory is reserved, only partially refund and issue a store credit for the difference," that logic in choreography has to be implemented as new event subscriptions and handlers distributed across whichever services are involved in that specific failure path, and there is no single place to write, review, or test that rule as a coherent unit; it exists as an emergent property of several services' independently-deployed event handlers agreeing (hopefully) on the right behavior. ## What orchestration scales gracefully Orchestration scales gracefully along the opposite dimension: **complex, changing, branching failure and compensation logic**. Because the whole sequence, happy path and every compensation path, lives in one component, adding a new conditional branch to the saga, like the partial-refund rule above, means writing it once in the orchestrator's state machine, where it can be code-reviewed, unit-tested, and reasoned about as a whole. This becomes increasingly valuable as the number of failure paths grows, since the alternative, choreography, would require that same growing complexity to be correctly distributed and kept consistent across multiple independently-deployed services. What doesn't scale gracefully in orchestration is **participant growth from unrelated teams**: every team that wants their service to be part of the saga's sequence, even for something that doesn't need real compensation logic, has to get a change into the shared orchestrator, which means code review, deploy coordination, and shared on-call risk for a component that now sits in the critical path of a growing number of unrelated features. ## How observability degrades on each side Observability degrades differently under the two patterns as they grow. | Pattern | What growth does to it | |---|---| | **Choreographed** | Systems at five services are often still debuggable by reading five services' worth of event handlers, but at fifteen services, understanding "what is supposed to happen when an order is placed" stops being something any one engineer can hold in their head, and teams typically respond by building dedicated tooling: an event-sourced saga tracker that subscribes to everything purely to reconstruct state, or heavy investment in distributed tracing with consistent correlation IDs propagated through every event, essentially reconstructing, at significant tooling cost, the single-source-of-truth view that orchestration gives you for free. | | **Orchestrated** | Systems keep a clear state machine view as they grow, but the orchestrator's own internal complexity grows instead, its state machine can become a sprawling, deeply-nested set of conditional branches that's readable in principle but increasingly hard to safely modify, and its blast radius as a single component grows too: an orchestrator outage or bug doesn't degrade one step of the process, it can halt every in-flight saga across the whole system, which is a materially larger failure domain than any single choreography participant going down. | ## The hybrid mature organizations converge on What mature organizations converge on, and this is the practical resolution most principal engineers land on, is a **hybrid**: - **orchestration** for the genuinely transactional core of the saga, the steps that need real compensation, ordering guarantees, and a single source of truth for "is this order in a consistent state"; - **choreography** for everything hanging off that core that's a true independent side effect with no compensation needs, like notifications, analytics, and recommendation-model updates. Concretely, an orchestrator might own payment, inventory, and shipping, including their compensating actions, and simply publish `OrderFulfilled` or `OrderFailed` events at the end, which any number of unrelated services can then choreograph off of without ever needing to touch or even be known to the orchestrator. This gets the observability and correctness benefits of orchestration exactly where compensation complexity actually lives, while preserving choreography's frictionless extensibility exactly where it's genuinely low-risk to do so, and it's the pattern you'll see described in most serious writing on saga design (including the standard saga-pattern literature that treats choreography and orchestration as implementation strategies for the same underlying pattern, not competing architectures) once a system has grown past the point where one approach cleanly fits the whole workflow.
- How would you decide, concretely, whether a new participant being added to an existing saga belongs in the orchestrator or should just subscribe to an event choreography-style?Ask whether the new participant's success or failure needs to affect the outcome of the core transaction, meaning does a failure in this new step require compensating any of the already-completed steps. If yes, meaning it's genuinely part of the all-or-nothing unit, it belongs in the orchestrator so its failure path is handled with the same rigor as the rest. If the new participant is a pure side effect whose own failure shouldn't roll back payment or inventory, like a recommendation-engine update, it belongs as an independent event subscriber outside the orchestrator's scope.
- What's a concrete sign that an orchestrator has grown into an organizational bottleneck, and what's a typical mitigation short of a full rewrite?A concrete sign is deploy frequency or lead time for the orchestrator's codebase visibly dropping compared to other services, along with multiple unrelated teams routinely blocked on the same PR queue or the same on-call team's review bandwidth. A typical mitigation short of a full rewrite is auditing which of the orchestrator's current steps actually need transactional compensation versus which were added there out of convenience and can be extracted into independent choreographed subscribers, shrinking the orchestrator back down to its genuinely transactional core.
- Why doesn't simply adding a distributed tracing system fully solve choreography's observability gap?Distributed tracing gives you visibility into individual request/event flows after the fact, which helps debug a specific stuck saga instance if you know to look for it, but it doesn't give you the aggregate, queryable view an orchestrator's explicit state naturally provides, like 'show me every saga currently stuck between payment and inventory across the last hour.' Getting that aggregate view out of tracing data alone typically still requires building a dedicated saga-state projection on top of the trace or event data, which is effectively reconstructing an orchestrator's state tracking as separate tooling.
It's like the difference between a jazz ensemble improvising off each other's cues (choreography: adding a new musician who listens and reacts is easy, but changing the song's actual structure mid-performance means everyone has to somehow agree) versus a film crew following a shot list from one director (orchestration: changing the shot list is one edit in one document, but every new crew member needs the director's sign-off to join the shoot).
saying these in an interview costs you the question
- Picks one pattern as universally correct regardless of failure-logic complexity or team topology
- Doesn't distinguish 'adding a participant' cost from 'changing the sequence/compensation logic' cost between the two patterns
- Assumes observability tooling for choreography is free or automatic
- Can't describe a hybrid approach or thinks the choice must be all-or-nothing for the whole saga
- Doesn't mention orchestrator blast radius (single component that can halt all in-flight sagas) as a real cost