skip to content

Explain choreography versus orchestration for coordinating a multi-service business process (e.g. an order saga) over Kafka. What are the trade-offs and failure-handling implications of each?

level: seniorimportance: must knowfreq 68%

answer

  1. Choreography = react to events, no central brain
  2. Orchestration = central coordinator sends commands, tracks state
  3. Saga = local txns + compensating actions
  4. Choreography: decoupled but implicit/hard to trace; cyclic risk
  5. Orchestration: visible/central but coupling hub + SPOF
  6. Both need outbox + idempotent consumers

basics

~20 s

In choreography each service reacts to events and emits its own, with no central coordinator — logic is distributed. In orchestration a central orchestrator tells each service what to do and tracks progress. Choreography is decoupled but hard to trace; orchestration is centralized, explicit, and easier to monitor.

solid answer

~50 s

**Choreography**: services collaborate by reacting to each other's domain events on Kafka topics. `OrderPlaced` triggers Payment, which emits `PaymentCaptured`, which triggers Shipping, etc. There is no central brain — the workflow is emergent from local event reactions. It's highly decoupled and easy to extend, but the end-to-end flow is implicit, hard to visualize, and cyclic dependencies/'event storms' can creep in. Compensation (saga rollback) is distributed: each service must listen for failure events and undo its work. **Orchestration**: a central orchestrator (e.g. a saga coordinator, or a workflow engine) issues commands to each service and consumes their replies, holding the explicit state machine. It's easy to monitor, reason about, and change the flow in one place, and compensation is centrally driven. Costs: the orchestrator becomes a coupling hub and potential bottleneck/SPOF, and services know less autonomy. Over Kafka, choreography uses event topics; orchestration uses command topics + reply topics with correlation ids. Many systems mix both — choreography between bounded contexts, orchestration within one complex saga.

go deeper

for a junior

Recall that choreography has no central coordinator and orchestration does.

for a middle

Contrast decoupling vs visibility and map each to event vs command topics.

for a senior

Reason about saga compensation, outbox + idempotency, correlation-id tracing, and when to mix styles.

for a principal

Drive org strategy: choose per-flow, set tracing/observability standards, and govern orchestrator SPOF and contract versioning across teams.

## The problem: multi-step business processes Many business operations span several services. Example **order saga**: reserve inventory → capture payment → create shipment → notify customer. Each step lives in a different service with its own database. You need a way to coordinate them **and** to undo earlier steps if a later one fails (a **saga** is a sequence of local transactions with **compensating** actions, since there's no distributed ACID transaction). Two coordination styles: ### Choreography (event-driven, decentralized) No coordinator. Each service **subscribes to events** and **emits new events**: - Order service emits `OrderPlaced`. - Inventory service reacts, reserves stock, emits `InventoryReserved` (or `InventoryReservationFailed`). - Payment service reacts to `InventoryReserved`, emits `PaymentCaptured` / `PaymentFailed`. - Shipping reacts to `PaymentCaptured`... **Pros**: maximal decoupling and autonomy; adding a participant means subscribing to an existing event without touching others; no central bottleneck. **Cons**: the workflow is **implicit and emergent** — no single place shows the whole flow, making it hard to understand, test, and debug. Risk of **cyclic event chains** and accidental coupling. **Compensation is distributed**: each service must itself listen for downstream failure events (`PaymentFailed` → Inventory releases its reservation). As steps grow, this 'who-reacts-to-what' web becomes brittle. ### Orchestration (command-driven, centralized) A dedicated **orchestrator** owns the process state machine. It **sends commands** ('reserve inventory') and **awaits replies** ('reserved'), then decides the next step. On failure it issues **compensating commands** ('release inventory', 'refund payment') in reverse. **Pros**: the flow is **explicit and centralized** — one place to read, change, version, and monitor; built-in visibility into where each saga instance is; compensation logic is centralized and easier to get right; timeouts/retries are managed in one component. Tools like workflow engines persist the state machine durably. **Cons**: the orchestrator is a **coupling hub** (it must know every participant) and a potential **bottleneck/SPOF**; participant services are more anemic; you've reintroduced a degree of central control. ## How each maps to Kafka primitives - **Choreography** = **event topics**. Services are independent consumer groups reacting to facts. Ordering of a saga per entity is preserved by **keying** events on the saga/aggregate id so they land on the same partition. - **Orchestration** = **command topics** (orchestrator → service) + **reply/event topics** (service → orchestrator), tied together by a **correlation id**. The orchestrator's own state can itself be event-sourced to a Kafka topic for durability. - Either way you need the **outbox pattern** so that a local DB write and the Kafka publish happen atomically (no dual-write inconsistency), and **idempotent consumers** because Kafka gives at-least-once delivery. ## Failure handling & edge cases - **Partial failure**: a step succeeds but the next never runs (consumer crash). Both styles rely on at-least-once redelivery + idempotency to recover; orchestration additionally uses timeouts to detect stuck steps. - **Compensation ordering**: undo in reverse; compensations must themselves be idempotent and may fail (need retries / dead-letter / human escalation). - **Visibility**: choreography needs distributed tracing (correlation/trace ids propagated through event headers) to reconstruct flow; orchestration gets it 'for free' from the central state. ## Choosing / mixing Use **choreography** for simple, stable, loosely-related steps and to keep contexts autonomous. Use **orchestration** for complex, evolving, business-critical flows that demand visibility and central control. Real systems often **mix**: choreography across bounded contexts, orchestration inside one complex saga. The honest senior answer is 'it depends on flow complexity, the need for visibility, and how much autonomy each service must keep.'

  • How do you maintain end-to-end visibility of a saga implemented with choreography?
    Propagate a correlation/trace id through Kafka record headers across every event, and use distributed tracing (e.g. OpenTelemetry) to stitch the flow back together. Without it, choreography's emergent flow is hard to debug.
  • Why do both choreography and orchestration over Kafka need the outbox pattern and idempotent consumers?
    Kafka delivery is at-least-once and you can't atomically write your DB and publish to Kafka. The outbox pattern makes the DB write + event emission atomic, and idempotent consumers tolerate redelivered events so duplicates don't corrupt the saga.
  • What makes an orchestrator a risk, and how do you mitigate it?
    It's a coupling hub (knows all participants) and a potential bottleneck/SPOF. Mitigate by making it stateless-per-request with durable persisted state (e.g. event-sourced or a workflow engine), running it HA, and keeping command/reply contracts versioned.

saying these in an interview costs you the question

  • Saying orchestration is always better, or choreography is always better
  • Claiming choreography needs no failure/compensation handling
  • Ignoring that both require idempotency under at-least-once delivery
  • Confusing choreography (events) with orchestration (commands) terminology
  • Forgetting the saga = compensating-actions model (no distributed ACID)

context