skip to content

In a choreographed checkout flow (order placed, then payment charged, then inventory reserved, then shipment created - each step reacting to the previous step's event), an order is stuck: payment shows charged but no shipment was ever created, and no service logged an error. Why is finding the root cause structurally harder here than in an orchestrator-driven version of the same flow, and what would you build to make it tractable?

level: seniorimportance: must knowfreq 65%

answer

  1. state is implicit, scattered across services
  2. silence not failure - no error to page on
  3. correlation ID + causation ID
  4. distributed tracing across the broker
  5. process-tracker read model = visibility without control

basics

~20 s

Nobody owns the whole flow, so there's no single place that knows 'step 4 never happened.' You have to reconstruct the story by piecing together logs from every service using a shared trace ID, and build dashboards specifically for tracking process state across services.

solid answer

~40 s

In choreography, process state exists only implicitly, as the sum of events each service happened to publish - there's no single durable record saying 'this order is at step 3 of 4, waiting on shipment.' A stuck order doesn't fail loudly; the next event just never fires, and every involved service looks locally healthy in isolation. Diagnosing it means correlating events across services by a shared identifier (a correlation ID propagated through every message), usually via distributed tracing, and often means building an external read model that projects the scattered events back into an explicit 'where is this order' state machine purely for observability. An orchestrator doesn't have this gap because it already persists that explicit state as a queryable record with a clear owner.

go deeper

for a junior

Can say that it's hard to tell where things stand because no one service knows the whole picture.

for a middle

Names correlation IDs and cross-service log tracing as the basic tool for reconstructing what happened.

for a senior

Explains why the failure is a silent absence rather than a logged error, and proposes a dedicated process-tracker read model plus distributed tracing as the fix, distinguishing it clearly from turning the tracker into an orchestrator.

for a principal

Weighs this observability cost against choreography's decoupling benefit as a deliberate architectural trade-off, and can articulate the hybrid pattern (choreographed execution + read-only tracking) as a way to get both.

## Nothing owns the process The root cause of this pain is structural: choreography deliberately has no component whose job is to track "the process." Each service only knows its own local slice — the payment service knows it charged the card and published `PaymentCharged`; the shipment service knows nothing at all, because it never received (or never successfully processed) that event. From any single service's point of view, everything looks fine: no errors were thrown, no retries exhausted, no alert fired. The "stuckness" only exists at the level of the whole business process, and no component's health check is scoped to the whole business process. ## The symptom is silence, not failure This is different in kind from a service being down. The classic symptom is silence, not failure: the message either - **was never published** (a bug in the payment service's event-publishing code, a transaction that committed the charge but rolled back before the outbox write), - **was published but never delivered** (a broker partition issue, a consumer group that fell behind or was misconfigured), - **or was delivered but silently dropped** (a deserialization error swallowed by a catch-all handler, or routed to a dead-letter queue nobody was monitoring). None of these produce an error in the sense of "a service logged a stack trace" — they produce an absence, and absences don't page anyone by default. ## The two things a cross-service timeline needs Tracking this down requires reconstructing a timeline that spans services, which requires two things that don't exist automatically. 1. **First, a shared correlation ID** (sometimes paired with a causation ID identifying which specific prior event triggered a given one) must be generated once, at the start of the process, and propagated through every subsequent event's headers, so that a query for that ID across every service's logs and message store returns the full, ordered story. If services weren't disciplined about propagating this ID from day one, retrofitting it means the correlation is simply impossible to reconstruct for historical incidents. 2. **Second, distributed tracing infrastructure** (commonly OpenTelemetry-based) needs trace context to flow through the message broker the same way it flows through HTTP headers in a synchronous call chain, which many teams only wire up after their first painful multi-service outage. ## Build the process tracker Even with both in place, choreography still lacks a natural place to ask "show me every order currently stuck for more than 10 minutes." The standard fix is to build a dedicated **read model** — sometimes called a process tracker or saga log — that subscribes to every event in the flow purely to project them into an explicit state per order ("placed → charged → [missing: reserved] → [missing: shipped]"), without that component ever issuing commands or making decisions (which would make it an orchestrator in disguise). This tracker becomes the dashboard on-call engineers query, and it's what turns "silent stall" into "alertable condition" by flagging orders that haven't advanced within an expected SLA. ## Why an orchestrator never has this gap An orchestrator sidesteps this whole class of problem by construction: because it must persist "which step this instance is on" in order to function at all (it needs that state to decide what to call next), the same state that drives the workflow is also, for free, the answer to "where is order 123 right now" — a queryable row rather than a reconstruction exercise. This is the core observability trade-off: choreography's diffuse execution avoids a central bottleneck but requires teams to deliberately build the correlation-ID discipline, tracing, and read-model tooling that orchestration gets as a side effect of its own control-flow bookkeeping. A concrete, widely-cited pattern here is pairing choreographed event flows with an explicit process-manager or saga-tracking read model precisely so teams get orchestration-like visibility while keeping choreography's decoupled execution — a hybrid that trades implementation effort for both benefits at once.

  • What's the difference between a correlation ID and a causation ID in this context?
    A correlation ID identifies the whole process instance (e.g. one specific order's journey end to end) and stays the same across every event in that flow. A causation ID identifies the single specific event that directly triggered the current one, letting you reconstruct the exact chain of cause-and-effect rather than just knowing which events belong to the same overall process.
  • Why can't you just add error logging to each service and rely on log aggregation to catch this?
    Because there's no error to log - each service completed its own work successfully and the failure is an absence at the process level (a next step that should have happened but didn't), not an exception any single service caught. Log aggregation only surfaces what was logged; you need an active read model comparing 'expected next event' against 'events actually observed' to detect a stall.
  • Does a process-tracker read model turn choreography back into orchestration?
    No, as long as it only observes and reports and never issues commands or makes the participants' decisions for them - it's read-only visibility bolted onto still-independent, still-decoupled participants. The moment that tracker starts telling services what to do next, it has become an orchestrator in disguise.

It's like a relay race with no scoreboard: if the third runner never picks up the baton, nobody at the finish line knows why - you'd have to interview every runner individually to figure out where the handoff broke, versus a race with a live scoreboard that shows exactly which leg the baton is stuck on.

saying these in an interview costs you the question

  • Assumes a stuck choreographed flow would show up as an error somewhere
  • Has no concept of correlation/causation IDs propagated through event headers
  • Thinks adding more logging per-service solves a cross-service state-tracking problem
  • Doesn't recognize that an orchestrator gets this visibility for free as a side effect of its own bookkeeping

context