When does an event-driven architecture beat synchronous request/response, and what does that choice cost you?
answer
- Event = past-tense fact, no expectation
- Queue absorbs bursts; fan-out is free
- At-least-once → idempotent consumers
- Order only within a partition key
- DLQ + correlation IDs or no debugging
basics
~20 sChoose event-driven when a producer should not wait for, or even know about, its consumers — bursty load, many independent reactions to one fact, long-running work. You pay with eventual consistency, harder debugging, duplicate and out-of-order messages, and error handling that is no longer a simple exception.
solid answer
~60 sEvent-driven architecture has components publish **events** — immutable facts about something that happened — which a broker delivers to any number of subscribers. The producer does not know who reacts, so new consumers can be added without touching it, and load spikes are absorbed by the queue instead of by the producer's threads. Pick it when: traffic is bursty and you need elasticity; one business fact triggers many independent reactions; work is long-running and the caller must not block; or you need loose coupling between teams' services. The costs are real. Consistency becomes eventual, so read-your-writes breaks and users see stale data. Delivery is typically at-least-once, so consumers must be idempotent, and ordering only holds within a partition/key. Errors have no caller to return to, requiring retries, dead-letter queues, and reconciliation. End-to-end flows exist only in traces, so debugging needs correlation IDs and distributed tracing. Broker topology (broker vs mediator) then decides whether workflow control is distributed or centralized. Use it where the drivers demand it, not everywhere — hybrid systems with synchronous queries and asynchronous side effects are the common, sane result.
go deeper
Explain that the producer publishes a fact and does not wait, so work happens in the background and load spikes are buffered; mention that data may be briefly out of date.
Name concrete drivers (bursty load, fan-out, long-running work) and the main costs — eventual consistency, retries and duplicates, harder debugging — plus the need for idempotent consumers.
Cover broker vs mediator, delivery semantics and partition-level ordering, dead-letter queues and replay, correlation IDs and tracing, event schema evolution, and the events-as-commands anti-pattern.
Frame it as a system-wide consistency and contract strategy: where the async boundaries go, saga/compensation design, event ownership and registry governance, the organizational coupling implications, and how you'd measure whether the async complexity is earning its keep.
### Core vocabulary - **Event** — an immutable statement that something *has happened*, named in the past tense (`OrderPlaced`, `PaymentCaptured`). The producer asserts a fact and expresses no expectation. - **Command / request** — an instruction to do something (`PlaceOrder`), directed at one specific handler, usually with a response expected. - **Message broker / event bus** — infrastructure that accepts events and delivers them to subscribers (a queue or log such as a topic-partitioned commit log). - **Producer / consumer (subscriber)** — the emitter, and the components that react. - **Asynchronous** — the producer continues without waiting for the consumer's work to finish; **fire-and-forget** — it never learns the outcome at all. ### The two classic topologies **Broker topology** — events are published to channels; each interested processor consumes and, if relevant, publishes further events. No central coordinator. Highest decoupling, best extensibility and performance; but there is no single place that knows the state of the overall workflow, error recovery is distributed, and cascades are hard to reason about ("who else fires when this fires?"). **Mediator topology** — an orchestrator (workflow engine, saga coordinator) receives an initiating event and dispatches steps in a known order. You gain visibility, explicit error handling, and controllable sequencing; you pay with coupling to the mediator, which becomes a bottleneck and a modeling burden for highly dynamic flows. The underlying tension is **choreography vs orchestration**, and it is one of the sharpest questions in distributed design. A practical rule: choreograph inside a bounded context where flows are simple; orchestrate cross-context business processes with money, compliance, or compensations attached. ### When event-driven genuinely wins 1. **Bursty load / elasticity.** The queue is a shock absorber. Consumers are scaled independently and the producer's latency stays flat under a 40× spike. 2. **Fan-out.** `OrderPlaced` triggers email, inventory reservation, loyalty points, analytics, fraud scoring. Synchronously, the order service accumulates five dependencies and its availability becomes the product of theirs; with events, adding a sixth reaction touches no existing code. 3. **Long-running work.** Video transcoding, report generation, batch settlement — nothing should hold an HTTP connection for minutes. 4. **Temporal decoupling.** A consumer can be down for maintenance and catch up afterward. With synchronous calls, its downtime is the caller's outage unless you build fallbacks. 5. **Team autonomy.** Consumers evolve on their own schedule as long as the event contract holds. ### What it costs - **Eventual consistency.** After the write returns, derived state is not yet updated. "Read your writes" breaks; UIs need optimistic updates, polling, or push. Business rules that need a global invariant checked atomically become hard, pushing you toward compensating actions (sagas) instead of transactions. - **Delivery semantics.** Exactly-once end-to-end is effectively unattainable across systems; brokers give at-least-once (occasionally at-most-once). Therefore **consumers must be idempotent** — deduplicate on an event ID, or make the effect naturally repeatable. Charging a card twice because a retry was not deduplicated is the canonical production incident. - **Ordering.** Global ordering does not exist at scale. Logs preserve order within a partition, so you partition by an entity key (order ID) to keep per-entity order, and design for out-of-order arrival elsewhere (version numbers, last-writer-wins, or state machines that reject impossible transitions). - **Error handling.** There is no caller to throw to. You need retry with backoff, a **dead-letter queue** for poison messages, alerting on DLQ depth, and a human or automated replay path. Silent DLQ accumulation is a classic outage cause. - **Observability.** No stack trace spans the flow. You need correlation/causation IDs propagated on every event and distributed tracing, or debugging becomes archaeology across logs. - **Schema evolution.** Events are a public contract consumed by systems you may not know about. You need a registry, additive-only changes, and versioning discipline; a removed field can break a downstream team silently. - **Testing and cognitive load.** Reasoning about interleavings, replays, and duplicates is harder than reading a call stack. This is a permanent tax on every future maintainer. ### Anti-patterns - **Events as disguised commands.** Publishing `SendEmailRequested` aimed at exactly one known consumer is a request/response call with extra latency and no reply channel. If you know and require the single handler, call it. - **Distributed monolith by event.** Services that cannot proceed until specific other services have reacted, with a request/reply pattern over the broker everywhere, get async complexity plus synchronous coupling. - **Fat events vs thin events.** Thin events (ID only) force consumers to call back to the producer, restoring runtime coupling; fat events (full payload) risk staleness and leak the producer's model. Common middle ground: an event carrying the meaningful business fields plus an ID for detail retrieval — decide consciously. - **Async everywhere.** Synchronous calls are cheaper, simpler, and immediately consistent. Use async where it buys a ranked quality attribute. ### How to answer Define event vs command, give two or three concrete drivers, then be explicit about consistency, idempotency, ordering, DLQ, and tracing. Naming broker vs mediator (choreography vs orchestration) and stating when you'd choose each signals real experience. Concluding that most systems end up hybrid — synchronous queries, asynchronous side effects — is the mature position.
- How do you make a consumer idempotent in practice?Either store processed event IDs and skip duplicates (a dedup table with the ID as a unique key, written in the same transaction as the effect), or design the effect to be naturally repeatable — an upsert to a known key, or a state transition guarded by a version so a replay is a no-op. Non-idempotent effects such as 'increment balance' or 'charge card' need an explicit idempotency key.
- Broker (choreography) or mediator (orchestration) — how do you choose?Choreography for simple flows inside one bounded context where extensibility matters and no single owner needs the whole picture. Orchestration for cross-context business processes with compensations, compliance, or money involved, where you need one place that knows the state of a workflow and how to recover it. Many systems use both at different scopes.
- A user submits a form and immediately sees stale data because the update is processed asynchronously. How do you handle it?Options in rough order of preference: update the UI optimistically from the command result; return a resource version or job ID the client polls or subscribes to; push the completion over websockets/SSE; or make that one read path synchronous. The real point is that eventual consistency is a UX design problem, not only a backend one, and it must be designed for deliberately.
Synchronous calls are a phone call: you wait on the line, get an answer immediately, and are stuck if nobody picks up. Events are posting a notice on a public board: you carry on instantly, anyone interested can react, nobody blocks — but you don't know when or whether they acted, and two people may act on the same notice.