skip to content

Event-Driven Design

Designing systems around events rather than requests: what an event carries, how flows are coordinated, and how publication is made reliable. Interviewers probe when events are the wrong choice.

part ofEvent-driven architecture & messagingoverview, primer and where to startread it →
on this pageshow

questions

24

In a system where several services each own one step of fulfilling an order (reserve inventory, charge payment, arrange shipping), what is the fundamental difference between coordinating those steps with choreography versus orchestration?

level: juniorimportance: must knowfreq 85%

answer

  1. dance troupe vs conductor
  2. implicit vs explicit process state
  3. events (facts) vs commands (orders)
  4. no coordinator vs single coordinator

basics

~20 s

Choreography: each service reacts to events from other services on its own, like dancers who each know their part with no director. Orchestration: one central 'conductor' service tells every other service exactly what to do and when.

solid answer

~40 s

Choreography has no central coordinator - each service publishes domain events (e.g. OrderPlaced, PaymentCharged) and other services subscribe to the events they care about, deciding independently what to do next. The workflow is implicit, emerging from the sum of these reactions. Orchestration puts a single component - the orchestrator - in charge: it holds the process definition explicitly and calls each participant directly, usually via command messages (ChargePayment, ReserveInventory), waiting for replies and deciding the next step itself. Choreography trades central control for service autonomy and loose coupling; orchestration trades some autonomy for a single place to see and change the whole flow.

go deeper

for a junior

Can state the one-line distinction (no coordinator vs. central coordinator) and give a simple example of each.

for a middle

Can also say why each style exists (avoiding distributed transactions) and name events vs. commands as the message shapes used by each.

for a senior

Connects the mechanism to concrete coupling and observability consequences, not just the definition.

for a principal

Uses the distinction to reason about where to draw process boundaries across teams, not just within one flow.

## The same problem, solved in opposite directions Both patterns solve the same problem: getting several independently-deployed services to complete a multi-step business process without a shared database transaction spanning all of them. They solve it in structurally opposite ways. ## Choreography — peer-to-peer and implicit In choreography, coordination is peer-to-peer and implicit. - **Service A** finishes its work and publishes a fact about the past — an event like `OrderPlaced` — onto a broker (Kafka, RabbitMQ, SNS/SQS, etc). Service A does not know or care who, if anyone, is listening. - **Service B** subscribes to that event because it needs to react (e.g. the payment service reacts to `OrderPlaced` by charging the customer, then itself publishes `PaymentCharged`). - **Service C** subscribes to `PaymentCharged` and reserves inventory. No component holds a map of the whole process; the process only exists as the sum of these independent reactions, discoverable by grepping who-subscribes-to-what across the codebase. ## Orchestration — centralized and explicit In orchestration, coordination is centralized and explicit. One component — **the orchestrator** — encodes the process as an actual artifact: a state machine, a sequence of steps, sometimes literally a BPMN diagram or a workflow-engine definition (Temporal, AWS Step Functions, Camunda are common implementations). The orchestrator issues **commands** — imperative, addressed messages like `ChargePayment` or `ReserveInventory` — to each participant, and for every one of them it: 1. waits for a response or a completion signal, 2. updates its own persisted state, 3. decides the next step. The participants generally don't know about each other at all; they only know the orchestrator. ## Why both patterns exist at all The reason both patterns exist is that a synchronous, distributed ACID transaction across order, payment, inventory, and shipping services is impractical — it would require locking rows across service and network boundaries, killing availability and scalability. Both choreography and orchestration replace that with asynchronous, eventually-consistent coordination; they differ only in where the "who does what next" knowledge lives. ## The trade-off: coupling versus visibility The trade-off is coupling versus visibility. | Style | What it buys | What it costs | |---|---|---| | **Choreography** | keeps services decoupled from each other's existence — the payment service never calls the inventory service directly, so either can be deployed, scaled, or replaced without touching the other's code, and adding a new consumer (e.g. a fraud-check service that also listens to `OrderPlaced`) requires zero changes to the producer | nobody has a single view of process state: "is order 123 done yet?" requires reconstructing history from scattered events | | **Orchestration** | gives you that single view for free — the orchestrator's state is the answer — and makes branching, timeouts, and retries easy to express as ordinary code | it does so by making every participant reachable by, and therefore coupled to, the orchestrator's command contract, and by turning the orchestrator into a component the whole process depends on | ## How each one fails - In production, choreographed flows tend to fail silently: an event is published but nobody realizes a listener died or was never wired up, and the process just stalls with no error anywhere. - Orchestrated flows tend to fail loudly but centrally: if the orchestrator is down or buggy, no workflow of that type makes progress, and the orchestrator's owning team becomes a bottleneck for every change to step ordering. ## Where each one earns its place A concrete illustration: many e-commerce platforms let "peripheral" concerns — sending a receipt email, updating a recommendation model, awarding loyalty points — happen via pure choreography off an `OrderPlaced` event, because those consumers are independent and it's fine if one is slow or briefly down. The same platforms often use an explicit orchestrator or workflow engine for the "critical path" — payment, inventory reservation, shipment creation — where ordering, timeouts, and retries actually matter and where a stalled step needs to be visible to an on-call engineer within minutes, not discovered days later by a customer complaint.

  • In choreography, are the messages passed between services usually events or commands, and why does that matter?
    They're events - past-tense facts like OrderPlaced that describe something that already happened, published without a specific addressee. This matters because the publisher makes no assumption about who will act on it, which is what gives choreography its loose coupling; a command like ChargePayment, by contrast, is addressed to a specific service and implies the sender knows that service exists.
  • Does orchestration mean the orchestrator does all the actual work itself?
    No - the orchestrator only sequences and tracks the process; the actual work (charging a card, reserving stock) still happens inside the domain services it calls. The orchestrator's job is purely coordination: deciding what happens next and recording where the process currently stands.

Choreography is a flash mob: everyone learned their own part in advance and reacts to cues from those around them, with no director on site. Orchestration is a conductor leading an orchestra: one person holds the score and points at each section exactly when it should play.

saying these in an interview costs you the question

  • Says orchestration means one giant service does everything itself
  • Can't name what kind of message choreography uses (events) versus orchestration (commands)
  • Thinks choreography has some hidden central registry tracking the workflow
  • Confuses choreography with simple pub/sub having nothing to do with a business process

context

open as a page

In distributed systems, what is the core difference between a request-driven (synchronous request/response) interaction and an event-driven (asynchronous event) interaction between two services?

level: juniorimportance: must knowfreq 75%

basics

~20 s

In request-driven, one service asks another and waits for an answer right away, like a phone call. In event-driven, a service announces "this happened" and moves on, and other services react whenever they get around to it, like a text message.

open as a page

A shipping service publishes a message saying only 'Order 4521 was placed', with no order details attached, and any subscriber that needs more information must call back to the order service's API to fetch it. What is this messaging style called, and how does it differ from a style where the message itself carries the full order payload (items, price, address)?

level: juniorimportance: must knowfreq 70%

basics

~10 s

It's called event notification: a thin 'something happened' alert with no data, so listeners must ask for details afterward. The other style, event-carried state transfer, puts all the needed details inside the message itself.

open as a page

A service needs to save an order to its database and publish an OrderPlaced event to a message broker. Why is calling db.save(order) followed by broker.publish(event) as two separate steps risky, and what could go wrong?

level: juniorimportance: must knowfreq 75%

basics

~20 s

If the app writes to the database and then sends a message as two separate steps, one can succeed while the other fails - the app crashing between them means the database and other services disagree about what happened.

open as a page

A team praises their choreographed order-fulfillment flow as 'loosely coupled' because services only communicate via published events with no direct calls between them. Six months later, changing the shape of the OrderPlaced event breaks three downstream services nobody remembered were listening. What went wrong with the 'loosely coupled' claim?

level: middleimportance: must knowfreq 70%

basics

~20 s

Choreography removes direct service-to-service calls, but every subscriber still depends on the exact shape of the events it listens to. That's still coupling - to a shared event contract instead of an API - and it's invisible because nobody 'calls' anywhere, so nobody tracks who depends on what.

open as a page

You're deciding whether a 'user updated their shipping address' change should trigger a synchronous API call to the shipping service or publish an AddressUpdated event instead. What's the actual trade-off you're making, beyond 'events are more scalable'?

level: middleimportance: must knowfreq 80%

basics

~20 s

If you call directly, you know right away it worked, but you depend on that service being up. If you publish an event, you don't depend on it being up right now, but you don't know immediately whether it processed the change, and it might be out of sync for a bit.

open as a page

A team designing an event-driven checkout flow is choosing between publishing thin 'OrderPlaced' notifications that consumers must query back for details, versus publishing events that carry the full order state. What are the concrete coupling and data-freshness trade-offs between these two approaches?

level: middleimportance: must knowfreq 75%

basics

~20 s

Thin events keep the producer as the single source of truth so data is always current, but every consumer has to call back, tying them to the producer being online. Fat events remove that call-back dependency but consumers can end up working with slightly outdated data.

open as a page

Walk through, step by step, how the transactional outbox pattern uses an outbox table plus a relay/poller process to publish an OrderPlaced event reliably after an order is saved.

level: middleimportance: must knowfreq 80%

basics

~20 s

The service writes both the order row and a row describing the event into the same database transaction. A separate background process then reads the new outbox rows, sends them to the message broker, and marks them as sent - so the event is only ever 'in flight' if the original database write actually succeeded.

open as a page

In a choreographed checkout flow (order placed, then payment charged, then inventory reserved, then shipment created - each step reacting to the previous step's event), an order is stuck: payment shows charged but no shipment was ever created, and no service logged an error. Why is finding the root cause structurally harder here than in an orchestrator-driven version of the same flow, and what would you build to make it tractable?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Nobody owns the whole flow, so there's no single place that knows 'step 4 never happened.' You have to reconstruct the story by piecing together logs from every service using a shared trace ID, and build dashboards specifically for tracking process state across services.

open as a page

Why does tracing the root cause of an incorrect order status typically take longer in an event-driven fan-out architecture than in a synchronous request/response chain, and what practices reduce that cost?

level: seniorimportance: must knowfreq 65%

basics

~20 s

In a normal call chain, you can follow one thread step by step in the logs. In an event system, many independent services react on their own schedule, so there's no single thread to follow, and you need extra tools like tracking IDs and dashboards to piece the story back together.

open as a page

You're designing the integration between an order service and five downstream consumers (inventory, shipping, analytics, fraud, and notifications), each needing different amounts of order detail. Walk through how you would decide, per consumer, between event notification and event-carried state transfer, and when you'd introduce event sourcing instead of either.

level: seniorimportance: must knowfreq 65%

basics

~20 s

Pick the style per consumer based on how much data they need, how current it must be, and how many consumers are asking - big or many-and-varied needs favor a fat event, few-and-current needs favor callbacks. Event sourcing is a different, bigger decision: making events the actual system of record, not just a way to tell other services something happened.

open as a page

A team's outbox relay publishes events out of order across different aggregates, and occasionally a consumer receives the same OrderPlaced event twice. What guarantees does the transactional outbox pattern actually provide here, and what must consumers/producers do to handle it correctly?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The outbox pattern guarantees an event eventually gets published if the database write succeeded, but not that it's published exactly once or perfectly ordered across everything - so consumers must be built to safely handle duplicate or occasionally reordered messages.

open as a page

A team replaces a tangle of event listeners with a single orchestrator service that explicitly calls each participant of an onboarding workflow. Why do critics call the orchestrator 'the new coupling point,' and what operational cost does that concentration create?

level: middleimportance: should knowfreq 55%

basics

~20 s

The orchestrator now has to know about and call every participant directly, so all the workflow logic and every change request funnels through one service. That service becomes a single point of failure and a bottleneck for both scaling and code changes.

open as a page

A checkout service publishes an OrderPlaced event that an inventory service consumes asynchronously to decrement stock. Immediately after checkout, the customer refreshes the product page and still sees the item as 'in stock' with the same quantity as before their purchase. What's happening, and what design choices reduce how often customers see this?

level: middleimportance: should knowfreq 60%

basics

~20 s

The inventory count hasn't caught up yet because the update happens a little after the purchase, not instantly. You can shrink that gap or accept it as normal, but some delay is expected with this design.

open as a page

A customer-profile service publishes event-carried state transfer messages ('CustomerUpdated') that include the customer's full current profile in each payload, delivered over a message broker that does not guarantee ordering across partitions. A downstream analytics service applies each event's payload directly onto its local copy as soon as it arrives. What can go wrong, and what technique fixes it?

level: middleimportance: should knowfreq 55%

basics

~20 s

If an older update arrives after a newer one, the analytics copy can be overwritten with stale data. Adding a version number or timestamp to each event and only applying it if it's newer fixes this.

open as a page

A team implements the outbox pattern using a database polling relay that queries SELECT * FROM outbox WHERE published = false every second. A colleague suggests replacing it with Debezium reading the database's write-ahead log via change data capture instead. What changes, and what are the trade-offs?

level: middleimportance: should knowfreq 60%

basics

~20 s

Polling repeatedly asks the database 'anything new?' which adds load and delay. CDC tools like Debezium instead tap directly into the database's internal change log and stream new rows out in near real time, without hammering the table with queries - but it needs more setup and access to that low-level log.

open as a page

Your team coordinates a 3-step signup flow via choreography and it works well. The flow later grows to 9 steps with conditional branches - skip a KYC check for low-risk users, retry an email-verification step up to 3 times, escalate to manual review after a timeout. What signal tells you the flow has outgrown choreography, and what would you consider moving to instead?

level: seniorimportance: should knowfreq 60%

basics

~20 s

When you can no longer hold the whole flow's logic in your head just from the events - lots of branches, retries, timeouts, ordering rules - that's the signal. Move that control-flow logic into a central orchestrator so the branching lives in one place instead of scattered across many services' event handlers.

open as a page

When a downstream service (say, an email/notification service) goes completely offline for 30 minutes, how does the blast radius differ between a request-driven design where the upstream service calls it synchronously versus an event-driven design where the upstream service publishes an event for it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

If the call is direct and synchronous, the upstream service can get stuck or start failing too while the other service is down. If it's an event, the upstream service is fine — the messages just pile up and get processed once the other service comes back.

open as a page

A team has been running an event-carried state transfer integration in production for two years. Consumers have drifted out of sync with producers a few times due to missed events, and the event schema has changed twice, breaking older consumers that hadn't upgraded. What concrete techniques would you put in place to make this integration more resilient to both staleness and schema drift going forward?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Add version numbers so consumers can tell old data from new, periodically resend full snapshots so gaps self-heal, and evolve the event schema by adding fields instead of breaking old ones. Together these catch and fix drift instead of letting it accumulate silently.

open as a page

Instead of writing to an outbox table, a service publishes an OrderPlaced event to Kafka first, and a consumer inside that same service subscribes to its own topic to then update its local 'orders' read table. What is this approach called, and how does it avoid the dual-write problem?

level: seniorimportance: should knowfreq 30%

basics

~20 s

This is called 'listen to yourself' - the service treats publishing the event as the single source of truth, and only updates its own database in reaction to consuming that event back, so there's just one write path instead of two independent ones.

open as a page

A platform team is under pressure to 'go event-driven' across the board after a successful migration of one workflow. As the architect, when would you push back and keep a workflow request-driven instead?

level: principalimportance: should knowfreq 45%

basics

~20 s

Not every part of the system benefits from going async. If people need an instant yes/no answer, if two things must happen together perfectly, or if the team can't yet trace bugs across services, staying with direct calls is often safer.

open as a page

A platform team wants the loose coupling of choreography for straightforward event fan-out, but the auditability and explicit control of orchestration for the parts of a process with strict ordering, SLAs, and branching. What hybrid approach lets them get both within the same business process, and what new problem does the hybrid itself introduce?

level: principalimportance: nice to knowfreq 40%

basics

~20 s

Split the process: keep simple, independent side-effects (like sending analytics or a receipt email) as choreographed event subscribers, and wrap the complex, ordered, time-bound core steps in a small orchestrator. The catch: now there are two coordination styles in one flow, and without a clear, enforced boundary, teams gradually blur the line and recreate the same confusion inside the hybrid.

open as a page

A platform team wants to give internal consumers the low-callback benefits of event-carried state transfer for order events, but the order aggregate is large (megabytes, with attachments) and some fields are sensitive (PII, payment tokens) that shouldn't be broadcast to every subscriber. Design an approach that gets most of the coupling and freshness benefits of carried-state events without shipping the entire aggregate to everyone, and explain how it relates to event sourcing if the order service internally uses an event-sourced write model.

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Send a medium-sized event with the commonly needed fields plus a reference (like a claim check) that lets consumers who truly need the big or sensitive parts fetch them separately, rather than sending everything to everyone. Event sourcing, if used internally, is a separate decision about how the order service stores its own history - it doesn't have to leak into what gets published externally.

open as a page

Under what circumstances would a senior engineer argue AGAINST introducing the transactional outbox pattern for a service that currently does a direct, unguarded dual write (save to DB, then publish), and what would they propose instead?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

If the event being published isn't critical - losing or duplicating it occasionally causes no real harm, or the same information can be recovered another way - then adding an outbox table and a relay process might be more complexity than the problem is worth.

open as a page