Event-driven architecture & messaging
Integrating systems through events and messages instead of direct calls: brokers, delivery guarantees, ordering, and models built from a log of what happened. Interviewers probe the subtle failures.
on this pageshowhide
guide
overview
~1 minEvent-driven architecture and messaging is the study of how services cooperate without calling each other directly: one side records or announces that something happened, a broker or a log carries it, and other sides act on it in their own time. Interviewers like the subject because the happy path is easy to draw and the failure paths are not. They want to hear what happens when a consumer crashes halfway through a message, when the same message arrives twice, when two arrive in the wrong order, and when a producer changes the shape of what it sends. A candidate who can sketch the boxes but cannot say which guarantee each arrow carries usually struggles with the follow-ups. The subject splits into four sections. [Messaging foundations](/topics/found-eda-messaging-foundations) covers the machinery every system shares: producers and consumers, brokers, queues versus logs, delivery guarantees, idempotent consumers, ordering, dead-letter handling, contract governance and stream processing. [Event-driven design](/topics/found-eda-event-driven-design) raises the architectural questions: what an event should carry, who coordinates a multi-step flow, how to publish reliably alongside a database write, and when a plain request is the better choice. [CQRS](/topics/found-cqrs) separates the model that accepts changes from the models that answer queries. [Event sourcing](/topics/found-event-sourcing) goes furthest, making the history of events the system of record and deriving current state from it. Start with the foundations, because every later section assumes you already accept that delivery is usually at-least-once and that consumers must cope with it. Design comes next, then CQRS, then event sourcing, which leans on both. Junior questions ask what a term means; senior and principal questions ask what breaks, how you would notice, and whether the pattern was worth adopting at all.
primer
Most of this hub rests on a few ideas. Hold them firmly and the questions below read as consequences rather than separate facts to memorise. - **Asynchrony trades immediacy for independence.** A producer that publishes and moves on no longer needs the consumer to be up, fast or even known. What it gives up is an immediate answer: it cannot tell in the same moment whether the work happened, and other parts of the system learn about the change later. - **Plan for duplicates as the normal case.** Losing messages is rarely acceptable, and avoiding loss across crashes means redelivering anything whose acknowledgement is in doubt. That makes at-least-once delivery the practical default and moves the burden of correctness onto the consumer. Exactly-once claims hold inside a defined boundary; name the boundary whenever you use the phrase. - **Ordering is local.** Order guarantees hold within a partition, a stream or a single consumer, and they weaken once work is spread across parallel consumers or retried. Designs pick a key, usually an entity's identifier, whose messages must stay in sequence, and accept that messages for different keys interleave. - **A queue and a log are different things.** A queue hands out work items and forgets the ones that are done; a log keeps an ordered history that each reader walks from its own position. The log is what makes replay, late-joining consumers and rebuilt read models possible; the queue makes spreading individual tasks across workers simple. - **An event is a fact; a command is a request.** A command can be refused. An event records something already decided and can only be reacted to. Naming, ownership and who may reject what all follow from which of the two you are sending. - **Events are a contract.** Removing direct calls does not remove coupling; it moves it into the shape and meaning of the messages. A publisher owns a schema that readers it may not know about depend on, and in an event-sourced system that schema must stay readable for as long as the history is kept. - **Derived views lag their source.** Read models, projections and downstream copies are updated after the write, so a reader can see an older state. Designing for that window, rather than hoping it stays small, underlies most CQRS and event sourcing questions.
- Message broker
- A server that accepts messages from producers and delivers them to consumers, so the two sides need not be running, reachable or known to each other at the same time.
- Point-to-point queue
- A channel in which each message is handled by one consumer out of a pool, which spreads work across instances; the basis of the competing-consumers pattern.
- Publish-subscribe
- A distribution style in which every subscription receives each message independently, so one event reaches many unrelated consumers without the producer knowing who they are.
- Event log
- An append-only, ordered record of messages that is retained after being read, so each consumer keeps its own position and can re-read history.
- Offset
- A consumer's position in a log; committing it records how far processing has got and decides what is read again after a restart.
- Partition
- A slice of a log that keeps its own order and is read by one member of a consumer group at a time; the unit of both parallelism and ordering.
- At-least-once delivery
- A guarantee that no message is lost, bought by redelivering whenever an acknowledgement is uncertain; duplicates are the price.
- Idempotent consumer
- A consumer whose effect is the same whether a message is processed once or several times, usually through tracked message identifiers or naturally repeatable updates.
- Dead-letter queue
- A holding place for messages that keep failing, moved aside after a retry limit so the main flow continues and someone can inspect them.
- Transactional outbox
- Recording an outgoing event in the same database transaction as the state change, then publishing it separately, so the change and the event cannot disagree.
- Choreography
- Coordination in which each service reacts to other services' events with no central controller; the overall flow exists only as the sum of the reactions.
- Orchestration
- Coordination in which one component directs each step of a process, holds its state and decides what happens next, including retries and compensation.
- Projection
- A read-optimised view built by applying events in order; because it is derived, it can be thrown away and rebuilt from the history.
- Eventual consistency
- A guarantee that derived copies converge on the source once updates stop, with no promise about how quickly.
- Aggregate
- A cluster of domain objects treated as one consistency boundary; in event sourcing each one usually owns a single event stream.
- Upcasting
- Converting an older stored event shape into the current one at read time, leaving the stored history untouched.
The four sections form a stack, and each assumes the one beneath it. **Foundations are the substrate.** Every pattern higher up runs on a broker or a log with at-least-once delivery, ordering that holds only per key, and consumers that can die mid-message. **Design decides what crosses service boundaries.** Whether an event carries full state or only a reference sets how much consumers depend on the producer staying available and how current their data is. Choreography versus orchestration decides where the logic of a multi-step process lives, and whether anyone can tell you where a stuck instance stopped. The [outbox and change data capture](/topics/found-eda-outbox-pattern) solve a problem the foundations create: a database and a broker are two systems, and no single local transaction spans both. **CQRS is a choice inside one service; events are one way to feed it.** Splitting the write model from the read models needs neither messaging nor event sourcing, but in practice the read side is often updated from events. That is where the foundations come back: projection updaters must tolerate duplicates, and users experience lag as a screen that has not caught up. **Event sourcing makes the log the source of truth.** Instead of storing state and publishing events on the side, the events are the state and everything else is derived. It pairs naturally with CQRS because a raw history is awkward to query. It also inherits every contract problem in the hub at its sharpest: stored events live as long as the system, so schema change becomes versioning and upcasting, and a bad state is corrected by appending new events rather than editing old ones. Across all four sections the same two questions recur in different vocabulary: what happens when a message arrives twice or out of order, and what a reader sees before derived state catches up.
- Producers and Consumers →
The roles, acknowledgements and consumer groups every later question takes for granted: who sends, who receives, and who confirms what.
- Delivery Guarantees →
The three delivery semantics and why at-least-once is the working default; most other sections are reactions to that choice.
- Idempotent Consumers →
The consumer-side answer to duplicates, and the section where interviewers look for races rather than definitions.
- Transactional Outbox & CDC →
How to change a database and publish an event without one succeeding alone; the first design pattern built directly on the foundations.
- CQRS →
Separate write and read models, and the lag window between them, before adding event sourcing on top.
- Event Sourcing →
Last, because it combines everything above: the log becomes the record, and projections, snapshots and versioning follow from that.
Writing a consumer that is correct only if each message arrives once; after a timeout or restart the second copy charges or ships again.
Checking a processed-ID table and then inserting into it as two separate steps; concurrent duplicates both pass the check unless uniqueness is enforced atomically.
Saving to the database and then publishing to the broker and calling it reliable; a crash between the two leaves them disagreeing.
Calling a choreographed flow loosely coupled while nobody tracks who consumes which event; the coupling moved into the event schema, where it is harder to see.
Retrying a poison message immediately and forever; it stalls everything behind it. Back off, cap the attempts, then route it to a dead-letter queue.
Proposing CQRS or event sourcing for simple CRUD screens without naming the problem it solves; two models and a lag window have real costs.
Treating the read-your-own-writes gap as acceptable because the model is eventually consistent; a user who saves and sees the old value files a bug.
Editing or deleting stored events to absorb a schema change or fix a bug, instead of versioning, upcasting or appending a correcting event.
Most design questions in this hub reduce to a handful of choices. Naming the one you are making, and what decides it, is usually what the interviewer is waiting for. - **Synchronous call versus event.** A call returns an answer and a clear error path, at the price of depending on the other side at runtime. An event removes that dependency and opens a window in which the two sides disagree. Prefer the call when the caller cannot continue without the result. - **Thin versus fat events.** A bare notification keeps the producer the single source of detail but makes consumers call back; an event that carries state lets consumers work alone on a copy that may be slightly old, and it widens the contract. - **Choreography versus orchestration.** Choreography keeps services independent while a flow is short and linear; orchestration gathers branching, timeouts and compensation into one inspectable place, and adds a component that knows about every participant. - **Ordering versus parallelism.** Stronger ordering means fewer independent lanes; more partitions or consumers raise throughput and weaken order across keys. - **Polling versus change data capture.** A polling relay is simple and puts query load and delay on the outbox table; reading the database's change stream is faster and lighter, and needs low-level access and more operational setup.
A few shapes recur across all four sections under different names; spotting them is how you place an unfamiliar question quickly. - **Make the handler safe to repeat.** Deduplication by message identifier, upserts keyed by entity and version checks in projection updaters are one idea: assume any message may arrive again. - **Record first, publish from the record.** The outbox, change data capture and event sourcing itself all make one durable local write the moment of truth and derive publication from it. - **Park what cannot be processed.** Bounded retries with growing delays, then a dead-letter destination with a way to inspect and redrive, keep one bad message from stalling the rest. - **Rebuild from history.** New consumers, repaired projections and new read models replay a retained log from the start, then follow it live. This works only where the history is kept. - **Correlate across hops.** A shared identifier carried on every message is what lets you reconstruct a flow that no single service owns.
explore
- Messaging Foundations59 questions
- Producers and Consumers6 questions
- Message Brokers6 questions
- Pub-Sub vs Queuing6 questions
- Message Queue vs Event Log6 questions
- Event Streaming6 questions
- Ordering and Partitioning6 questions
- Delivery Guarantees5 questions
- Idempotent Consumers6 questions
- Dead-Letter Queues6 questions
- Message Contract Governance6 questions
- Event-Driven Design24 questions
- Notification vs State Transfer6 questions
- Choreography vs Orchestration6 questions
- Transactional Outbox & CDC6 questions
- Event-Driven vs Request-Driven6 questions
- CQRS34 questions
- Command Model6 questions
- Query Model5 questions
- Projections6 questions
- Consistency & Synchronization6 questions
- Event Sourcing Integration6 questions
- Cost of Two Models5 questions
- Event Sourcing33 questions
- Event Store6 questions
- Replay and Rehydration5 questions
- Snapshots5 questions
- Projections and Read Models6 questions
- Versioning and Upcasting5 questions
- Event Sourcing with CQRS6 questions
- API Designskillanchors this topic
- Backend Developerroleanchors this topic
- Full Stack Developerroleanchors this topic
- Java Backend Developerroleanchors this topic
- Kotlin Backend Developerroleanchors this topic
- Software Design & Architectureskillanchors this topic
- Forward Deployed Engineerrole
- Game Developerrole
- Server-Side Game Developerrole
- Software Architectrole
questions
150 · 4 sectionsWhat is a message broker, and why would two services send messages through one instead of calling each other's APIs directly?
basics
~20 sA message broker is a middleman server that receives messages from senders and delivers them to receivers, so the two sides never talk directly. This lets one side keep working even if the other is slow, down, or busy.
In a message queue system, what is a dead-letter queue (DLQ), and why does a consumer route a message there instead of retrying it forever?
basics
~20 sA DLQ is a separate queue where messages that fail processing repeatedly get moved to, instead of blocking the main queue or being retried forever, so the app keeps working while someone looks at the bad message later.
In event-driven messaging systems, what do at-most-once, at-least-once, and exactly-once delivery mean, and which one do most production systems actually rely on by default?
basics
~20 sAt-most-once can lose a message but never repeats it. At-least-once never loses a message but can repeat it. Exactly-once means every message is processed once, no loss and no duplicates. Most real systems use at-least-once plus deduplication.
In an event streaming platform like Kafka, what is a consumer offset, and how does it allow a consumer to replay events it has already processed?
basics
~10 sAn offset is a bookmark marking how far a consumer has read in a partition. Since the log isn't deleted after reading, moving the bookmark backward lets the consumer re-read and reprocess old events.
A message queue delivers the same order-created message to a consumer twice because the consumer crashed after processing it but before acknowledging it. What must the consumer do so processing the message twice doesn't cause a duplicate charge or duplicate shipment?
basics
~20 sIdempotent means doing something twice has the same effect as doing it once. The consumer remembers which messages it already handled (by ID) and skips reprocessing the effect, like charging money again, even if the message arrives more than once.
In a system where several services each own one step of fulfilling an order (reserve inventory, charge payment, arrange shipping), what is the fundamental difference between coordinating those steps with choreography versus orchestration?
basics
~20 sChoreography: each service reacts to events from other services on its own, like dancers who each know their part with no director. Orchestration: one central 'conductor' service tells every other service exactly what to do and when.
In distributed systems, what is the core difference between a request-driven (synchronous request/response) interaction and an event-driven (asynchronous event) interaction between two services?
basics
~20 sIn request-driven, one service asks another and waits for an answer right away, like a phone call. In event-driven, a service announces "this happened" and moves on, and other services react whenever they get around to it, like a text message.
A shipping service publishes a message saying only 'Order 4521 was placed', with no order details attached, and any subscriber that needs more information must call back to the order service's API to fetch it. What is this messaging style called, and how does it differ from a style where the message itself carries the full order payload (items, price, address)?
basics
~10 sIt's called event notification: a thin 'something happened' alert with no data, so listeners must ask for details afterward. The other style, event-carried state transfer, puts all the needed details inside the message itself.
A service needs to save an order to its database and publish an OrderPlaced event to a message broker. Why is calling db.save(order) followed by broker.publish(event) as two separate steps risky, and what could go wrong?
basics
~20 sIf the app writes to the database and then sends a message as two separate steps, one can succeed while the other fails - the app crashing between them means the database and other services disagree about what happened.
A team praises their choreographed order-fulfillment flow as 'loosely coupled' because services only communicate via published events with no direct calls between them. Six months later, changing the shape of the OrderPlaced event breaks three downstream services nobody remembered were listening. What went wrong with the 'loosely coupled' claim?
basics
~20 sChoreography removes direct service-to-service calls, but every subscriber still depends on the exact shape of the events it listens to. That's still coupling - to a shared event contract instead of an API - and it's invisible because nobody 'calls' anywhere, so nobody tracks who depends on what.
In a system that separates commands from events, what is the key difference between a command like `PlaceOrder` and an event like `OrderPlaced`, and why does that naming difference matter for how each one is handled?
basics
~20 sA command asks the system to do something and can be refused, like a request. An event says something already happened and is a fact - you can't refuse a fact, only react to it.
In a CQRS (Command Query Responsibility Segregation) system, commands update a write model and queries read from a separate read model that is synchronized afterward. Why does a user sometimes not see their own change immediately after saving it, and what do engineers call the time window during which this can happen?
basics
~20 sThe read copy of the data updates a little after the write copy, through a background sync step. Until that catches up, queries can show old data. That gap in time is called the lag window (or replication lag).
CQRS separates the model used to handle commands (writes) from the model used to handle queries (reads). Is Event Sourcing a required part of implementing CQRS, or is it a separate, optional choice? Explain how the two typically fit together when a team does use both.
basics
~20 sCQRS just means writes and reads use different models. Event Sourcing means you store every change as an event instead of overwriting data. You can do CQRS without Event Sourcing, but when a team does use Event Sourcing, the event log naturally becomes the write side and the read side is built from those events.
In a CQRS system, what is a 'projection', and why do teams build one instead of just querying the write-side data directly?
basics
~20 sA projection is a read-only copy of your data, shaped for easy querying, that's kept updated by watching events from the part of the system that handles writes. Teams build it because the write data is often shaped for correctness, not for fast, flexible reading.
In a CQRS system, what is a 'read model' (also called a query model) and why is it usually shaped differently from the domain/write model?
basics
~20 sA read model is a copy of your data shaped exactly for what a screen or report needs to show, separate from the data structure used to save changes. It exists so reading is fast and simple, and writing stays focused on correctness.
In an architecture that combines Event Sourcing with CQRS (Command Query Responsibility Segregation), a client submits a command to change data. Walk through what happens on the write side before that change becomes visible to a query.
basics
~10 sThe command side loads the past events for that thing, checks the business rules, creates one new event describing what happened, and appends it. Nothing is overwritten, only added.
What is an event store, and how does its append-only log differ from a traditional database table that you update in place?
basics
~20 sAn event store only lets you add new records, never change or delete old ones. Instead of overwriting a row to reflect current state, it keeps every change as its own permanent entry, in order.
In an event-sourced system, why can't you just edit the shape of an already-stored event when your code's model changes, and what technique is normally used to let old and new event shapes coexist?
basics
~20 sOld events are stored forever and get replayed to rebuild state. If you change what a field means without a plan, old events break the code that reads them. Upcasting is a small translator that turns an old-shaped event into the new shape before your code sees it, so replay still works.
In an event-sourced system, what is a 'projection' and why do we build one instead of querying the raw event log directly for reads?
basics
~20 sA projection is a separate copy of your data, built by reading through the history of every change (events) and turning it into a simple, fast-to-search table - like building an index from a diary instead of re-reading the whole diary every time you need an answer.
In an event-sourced system, an aggregate's current state isn't stored directly - it's derived by replaying its event stream. In plain terms, how does the system reconstruct that state, and what's this process called?
basics
~20 sThe system starts with an empty/default state and applies each stored event for that aggregate, one by one, in the order they happened, updating the state a bit with each event. Doing this to rebuild state is called rehydration; the loop that folds each event into state is the fold/reduce pattern.