Your team coordinates a 3-step signup flow via choreography and it works well. The flow later grows to 9 steps with conditional branches - skip a KYC check for low-risk users, retry an email-verification step up to 3 times, escalate to manual review after a timeout. What signal tells you the flow has outgrown choreography, and what would you consider moving to instead?
answer
- branching + timeouts + retries = process logic needing a home
- logic leaks into participants or an ad-hoc uncontrolled orchestrator
- audit/compliance needs queryable state
- workflow engines: Temporal, Step Functions, Camunda
- decision is per-flow, not system-wide
basics
~20 sWhen you can no longer hold the whole flow's logic in your head just from the events - lots of branches, retries, timeouts, ordering rules - that's the signal. Move that control-flow logic into a central orchestrator so the branching lives in one place instead of scattered across many services' event handlers.
solid answer
~40 sThe signal is rising control-flow complexity: conditional branching, cross-step timeouts, retry/backoff policies, and ordering constraints that don't map cleanly to 'one service reacts to one event.' In choreography that logic has to live somewhere, and it typically leaks into participants that now need to know about steps they don't own (e.g. 'did KYC run or was it skipped?'), or into ad-hoc coordinator services that emerge organically without being designed for the job. Once the team is spending more effort tracing implicit event chains than writing actual business logic, it's time to introduce an explicit orchestrator, or a workflow engine, that owns branching, timeout, and retry logic as first-class, testable code - accepting that this concentrates coupling and adds a component whose availability the whole process now depends on.
go deeper
Can say that more complicated flows are harder to manage with just reactive events.
Names branching and timeouts as concrete complexity signals that push toward orchestration.
Explains what happens if the signal is ignored (logic leaks into participants or an unplanned ad-hoc orchestrator) and names real workflow-engine options.
Frames it as a per-flow, reversible architectural decision made against concrete signals, not a one-time system-wide commitment, and can justify keeping simple flows choreographed even in an orchestration-heavy platform.
## What choreography is actually good at A useful way to frame this decision is that choreography is well-suited to fan-out with independent, order-insensitive reactions, and gets progressively worse as the process accumulates control-flow features that don't naturally belong to any single participant. The 3-step signup flow started as a good fit: account-created triggers a welcome email and a CRM sync, two independent, parallel reactions with no ordering dependency between them and no branching. Growing to 9 steps with conditional skips, retries, and a timeout-triggered escalation changes the shape of the problem entirely — now there's real process logic (not just reactive fan-out) that someone has to own. ## The concrete signals to watch for The concrete signals to watch for, each of which is a symptom of control-flow logic that has nowhere natural to live in pure choreography: | # | Signal | Why it needs a home | |---|---|---| | 1 | **Conditional branching** | "do step D only if condition X from step B was true" forces either the step-B service to embed knowledge of step D's existence (violating its own boundaries) or a new component to be introduced that inspects multiple services' events to decide | | 2 | **Cross-step timeouts and escalation** | "if step C hasn't completed within 24 hours, notify a human" requires something to be watching elapsed time across the whole process, which no single reactive event handler naturally does | | 3 | **Centrally-defined retry policy** | if step E needs "retry 3 times with backoff, then escalate" as a process-level rule rather than a pure per-call resilience policy, that's process logic, not domain logic | | 4 | **Audit/compliance requirements** | regulators or support teams needing to answer "exactly what state is application 456 in right now, and why" pushes hard toward an explicit, queryable process state (this compounds the observability gap covered separately) | When two or more of these appear together, choreography's diffuse execution model starts fighting the team rather than helping it. ## What happens if you ignore the signal What typically happens if a team ignores the signal and keeps stretching choreography is that the missing control-flow home doesn't disappear — it gets built anyway, just badly. Either - the branching logic gets duplicated into whichever participant happens to touch it first (so "skip KYC for low-risk users" ends up half-encoded in the KYC service and half in the account service, inconsistently), - or someone builds an ad-hoc coordinator service under time pressure that ends up being an unplanned, undocumented, poorly-tested orchestrator — all the coupling costs of orchestration with none of its deliberate design benefits (no clear ownership, no workflow-engine durability guarantees, no visibility tooling built for it). ## The recommended move The recommended move at that point is to introduce a proper, explicit orchestrator — often backed by a workflow engine (Temporal, AWS Step Functions, and Camunda-style BPM engines are common real-world choices) that provides durable state, built-in retry/timeout primitives, and a visual or queryable representation of where each instance stands — and let it own exactly the process-level concerns (sequencing, branching conditions, timeouts, retries) while still delegating actual domain decisions to the owning services via calls (the orchestrator asks the risk service whether to skip KYC rather than encoding the threshold itself, keeping the orchestrator thin as covered in the coupling-concentration trade-off). ## The signals that argue for staying put The decision isn't one-directional, though — the same signals in reverse argue for staying with, or reverting to, choreography: - a small, fixed number of steps; - purely independent, order-insensitive reactions (parallel notifications, analytics, audit logging); - a strong desire to let teams add new listeners without ever touching the producer's code or deploy schedule; - and no regulatory need for a single source of truth on process state. Forcing orchestration onto a genuinely simple, branch-free fan-out adds an unnecessary single point of failure and an unnecessary bottleneck team for no real benefit — the mirror image of the mistake of stretching choreography past its control-flow limits. The judgment call, in practice, is made per-process, not once for the whole system: it's entirely normal, and often correct, for one platform to run some flows choreographed and others orchestrated, chosen by which of these signals each specific flow exhibits.
- What happens if a team ignores the signal and just keeps adding branches to a choreographed flow instead of switching?The control-flow logic doesn't disappear, it gets built anyway, badly - typically duplicated inconsistently across whichever services happen to touch each branch first, or an ad-hoc coordinator service gets built under deadline pressure without the deliberate design, durability, and visibility tooling a real orchestrator or workflow engine would have provided.
- Is it reasonable for one platform to run some flows as choreography and others as orchestration?Yes, this is normal and often the right call - the decision is best made per-flow based on that flow's own branching, timeout, and audit needs rather than as a single system-wide architectural mandate, since a simple independent-fan-out flow and a complex conditional workflow have genuinely different requirements.
A flash mob works great for a simple synchronized dance, but once you need 'if it starts raining, everyone moves indoors and waits for a signal before resuming' you need someone directing in real time - the same crowd, but the coordination needs have outgrown what cue-reacting alone can handle.
saying these in an interview costs you the question
- Treats 'choreography vs orchestration' as a one-time, system-wide decision rather than per-flow
- Doesn't recognize conditional branching and cross-step timeouts as the concrete signals to watch for
- Assumes ignoring the signal has no cost - doesn't predict where the missing logic will end up
- Would switch a simple, branch-free, order-insensitive flow to orchestration for no clear reason