skip to content

CQRS

Separating the write model from the read model so each is shaped for its job: commands, projections, and the gap in between. Interviewers probe whether you can say when it is not worth it.

part ofEvent-driven architecture & messagingoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

In a system that separates commands from events, what is the key difference between a command like `PlaceOrder` and an event like `OrderPlaced`, and why does that naming difference matter for how each one is handled?

level: juniorimportance: must knowfreq 75%

answer

  1. verb tense = contract
  2. commands: one handler, can fail
  3. events: many subscribers, immutable fact
  4. imperative vs past-tense naming
  5. request vs record

basics

~20 s

A command asks the system to do something and can be refused, like a request. An event says something already happened and is a fact - you can't refuse a fact, only react to it.

solid answer

~40 s

Commands are imperative, addressed to exactly one recipient: PlaceOrder, CancelSubscription. They express intent about the future and can be rejected - bad input, a broken business rule, insufficient funds - because nothing has happened yet. Events are past-tense facts: OrderPlaced, SubscriptionCancelled. They record something that already succeeded and are immutable; you cannot 'reject' an event, only choose whether to act on it. Structurally this maps to different fan-out and failure semantics: a command has one handler that reports success or failure, while an event can have zero, one, or many subscribers that each react independently and can't veto what already happened. Getting the tense right in naming isn't cosmetic - it tells every future reader whether the message can still fail or is already committed history.

go deeper

for a junior

Should recognize the naming convention, imperative vs past-tense, and give one example of each without confusing them.

for a middle

Should explain that commands can be rejected while events can't, and connect that to why command handlers report success or failure.

for a senior

Should discuss fan-out differences, one handler vs many subscribers, and how that shapes error handling and coupling in a real system.

for a principal

Should discuss how this distinction shapes system-wide contracts - why teams version events differently from commands, and how misnaming a message type causes downstream coupling and replay bugs.

## Two vocabulary types on the write side In a message-driven write side, two vocabulary types coexist, and confusing them is one of the most common design mistakes teams make. - A **command** is imperative in name and intent: `PlaceOrder`, `TransferFunds`, `CancelSubscription`. It is addressed to exactly one recipient — a single handler owns the right to accept or reject it — and it represents something that has not happened yet. Because nothing has happened, a command can legitimately fail: the order might be invalid, the account might lack funds, the subscription might already be cancelled. - An **event**, by contrast, is past-tense: `OrderPlaced`, `FundsTransferred`, `SubscriptionCancelled`. It records something that has already happened and succeeded. Events are immutable historical facts, published to zero, one, or many subscribers who each react independently, and none of them gets a vote on whether the thing happened — it already did. ## Why the distinction exists This distinction exists because a system needs two different conversations happening at different points in time, with different failure semantics. - The **'can this happen'** conversation belongs to commands: it needs authorization, validation against current state, and a clear accept/reject outcome that the sender can act on. - The **'this happened, does anyone care'** conversation belongs to events: it needs broadcast, not permission; consumers subscribe voluntarily and process at their own pace, some immediately, some hours later, some never. Collapsing these into one message type produces systems where nobody can tell, just from looking at a message name, whether it's still safe to reject or already too late. ## The trade-off The trade-off shows up on both sides. | Message type | What it gains | What it gives up | |---|---|---| | **Commands** | buy synchronous certainty — the caller usually gets a definite yes/no back | at the cost of tight coupling: the sender must know the exact command shape and exactly one handler must exist to own it, and the sender typically has to build explicit failure-handling paths for every rejection reason | | **Events** | buy loose coupling and scalability — any number of services can subscribe to `OrderPlaced` without the order-placing code knowing or caring who's listening | at the cost of much harder reasoning: you can no longer look at the code and see who reacts to a given fact, ordering and delivery guarantees become distributed-systems problems, and once published there is no 'undo,' only a further compensating event | ## The classic failure mode: blending the two In production, the classic failure mode is blending the two. 1. A team publishes something named imperatively — say, `ProcessRefund` — onto a shared event-style topic where 'whichever service happens to be subscribed' picks it up, but nobody has designed for the case where two services both claim it, or none does, because event topics don't inherently guarantee single-consumer ownership the way a command dispatch table does. 2. The mirror-image mistake is naming something past-tense — `OrderCancelled` — but routing it as a targeted, single-consumer, must-succeed-synchronously call; consumers built expecting a fire-and-forget notification suddenly have to handle synchronous failure paths for a message whose name told them it was already a done deal. Both mistakes usually surface as confused on-call incidents: a 'command' silently drops because nobody was subscribed, or an 'event' blocks a critical path because its one consumer was down. ## A checkout flow, end to end A concrete example: in an e-commerce checkout flow, the API layer sends a `PlaceOrder` command to the order-management module and waits synchronously for the result, because the shopper needs an immediate 'order confirmed' or 'card declined' response before the page can move on — that's a command relationship, one sender, one handler, a definite answer. Once the order handler validates the request and commits the new order, it raises `OrderPlaced` onto a broker. Inventory service, notification service, and analytics service each subscribe to that event independently; none of them can reject the order at that point, they can only decrement stock, send a confirmation email, or log a metric in reaction to a fact that has already been recorded. If the notification service is down when `OrderPlaced` fires, the order still stands — the customer just gets their confirmation email late, which is a very different failure mode than if `PlaceOrder` itself had failed to reach its handler. ## Different message, different plumbing This also explains why the two message types are usually implemented with different infrastructure primitives even inside the same codebase. - **Commands** are often modeled as direct method calls or point-to-point queues with a single consumer group, because the whole point is that exactly one piece of code owns the decision, and routing a command to more than one handler by mistake, for instance a misconfigured queue with two competing consumers, produces a genuinely dangerous bug: the same `PlaceOrder` might get approved twice by two different handler instances racing each other. - **Events**, on the other hand, are naturally modeled as publish-subscribe topics, because the whole point is that the publisher doesn't know or care how many listeners exist, and adding a fourth subscriber to `OrderPlaced` next quarter should require zero changes to the order-placing code. Recognizing which infrastructure pattern a given message needs is a direct, practical consequence of getting the command-versus-event classification right at design time, not an afterthought bolted on once the messaging layer is already built.

  • Can a command handler publish multiple events?
    Yes - a single command like PlaceOrder might, on success, produce OrderPlaced plus other events if the aggregate's boundary covers more than one concern, or more commonly one event plus follow-up commands dispatched to other aggregates afterward. The handler decides the full set of state changes and events atomically within its own aggregate and transaction boundary.
  • What happens if a command has no handler registered?
    That's typically a hard configuration or routing error at dispatch time, not a business rejection - the caller gets an error like 'unknown command type' rather than a validation failure. This differs from events, where having zero subscribers is a normal, silent, and valid state.
  • Why can't you 'undo' an event the way you can reject a command?
    An event describes something already persisted and, in many designs, already had side effects - other services have reacted to it. Undoing it would mean lying about history; the correct move is to publish a new compensating event, such as OrderCancelled, rather than retract the old one.

A command is like handing a waiter your order - the kitchen can still refuse it ('we're out of salmon'). An event is like the receipt printed after the meal was cooked and served - it's a record of what happened, and nobody can un-print it or argue with it, they can only file it.

saying these in an interview costs you the question

  • Names commands in past tense or events in imperative, e.g. a queue called 'OrderPlace' used as an event
  • Says a command handler can 'fire and forget' with no way to report rejection
  • Treats events as something that can be rejected or fail validation the same way commands do
  • Can't explain why an event can have many subscribers but a command routes to exactly one handler

context

open as a page

In a CQRS (Command Query Responsibility Segregation) system, commands update a write model and queries read from a separate read model that is synchronized afterward. Why does a user sometimes not see their own change immediately after saving it, and what do engineers call the time window during which this can happen?

level: juniorimportance: must knowfreq 80%

basics

~20 s

The read copy of the data updates a little after the write copy, through a background sync step. Until that catches up, queries can show old data. That gap in time is called the lag window (or replication lag).

open as a page

CQRS separates the model used to handle commands (writes) from the model used to handle queries (reads). Is Event Sourcing a required part of implementing CQRS, or is it a separate, optional choice? Explain how the two typically fit together when a team does use both.

level: juniorimportance: must knowfreq 70%

basics

~20 s

CQRS just means writes and reads use different models. Event Sourcing means you store every change as an event instead of overwriting data. You can do CQRS without Event Sourcing, but when a team does use Event Sourcing, the event log naturally becomes the write side and the read side is built from those events.

open as a page

In a CQRS system, what is a 'projection', and why do teams build one instead of just querying the write-side data directly?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A projection is a read-only copy of your data, shaped for easy querying, that's kept updated by watching events from the part of the system that handles writes. Teams build it because the write data is often shaped for correctness, not for fast, flexible reading.

open as a page

In a CQRS system, what is a 'read model' (also called a query model) and why is it usually shaped differently from the domain/write model?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A read model is a copy of your data shaped exactly for what a screen or report needs to show, separate from the data structure used to save changes. It exists so reading is fast and simple, and writing stays focused on correctness.

open as a page

In CQRS (Command Query Responsibility Segregation), the model used to handle writes is separated from the model used to serve reads. What's the main reason a team adopts this split, and what's the most immediate cost it pays for doing so?

level: juniorimportance: must knowfreq 75%

basics

~20 s

Reads and writes often need different shapes: writes need consistency and validation, reads need speed and flexible views. Splitting them lets each side be optimized separately — but now you maintain two models instead of one, and they have to be kept in sync.

open as a page

Walk through, step by step, what a command handler for something like `PlaceOrder` typically does from the moment it receives the command to the moment it finishes, and explain why a single handler is usually scoped to exactly one aggregate.

level: middleimportance: must knowfreq 68%

basics

~20 s

It loads the relevant object from storage, checks the request makes sense given that object's current state, changes the object, and saves it back, all as one all-or-nothing step touching just one thing, so it can't half-fail.

open as a page

A command object like `TransferFunds(fromAccountId, toAccountId, amount)` passes through validation before its handler runs any business logic. What kinds of checks belong in that pre-dispatch validation step, versus what should be left for the handler to check against live state?

level: middleimportance: must knowfreq 72%

basics

~20 s

Before dispatch, check the command's shape is sane - no missing fields, no negative amount. Inside the handler, check things that depend on current data, like whether the account actually has enough money right now.

open as a page

In a CQRS front-end that shows an 'optimistic UI' after a user submits a command - for example, liking a post updates the like count instantly, before the read model confirms it - what has to happen if the command later fails validation, or the eventual read-model update disagrees with what the UI guessed?

level: middleimportance: must knowfreq 70%

basics

~20 s

The UI has to notice the mismatch and fix itself - either roll back to the old value and show an error, or quietly swap in the real value once the read model catches up, so the screen never keeps showing a fact that turned out to be wrong.

open as a page

A user submits a profile update through a CQRS-based app, then is redirected to a 'view profile' page that queries the read model. Name two concrete techniques for making sure that page shows the just-saved change even though the read model may not have caught up yet, and explain how each works.

level: middleimportance: must knowfreq 75%

basics

~20 s

Two options: (1) right after saving, read straight from the write side (or a cache of it) for that user's own data instead of the lagging read model; (2) have the client remember a version number from its last write and make the read model wait until it has caught up to at least that version before answering.

open as a page

In an event-sourced CQRS write model, before a command handler can validate and apply a new command against an aggregate (say, an existing Order), the aggregate's current in-memory state has to be produced somehow — there's no row to just SELECT. How does that reconstitution actually happen?

level: middleimportance: must knowfreq 75%

basics

~20 s

The system pulls every past event for that one specific order from the event log, in the order they happened, and feeds them one by one into a fresh, empty Order object, which updates itself a little with each event. By the end, that object looks just like the order does right now, and only then does the command get checked against it.

open as a page

In a CQRS system where the write side uses Event Sourcing, how does a read model — for example, a 'customer order summary' table used by the UI — get built and kept up to date from the event log? Walk through the mechanism end to end.

level: middleimportance: must knowfreq 85%

basics

~20 s

A separate small piece of code called a projector watches the event log for new events. Whenever it sees an event it cares about (like 'order shipped'), it updates a plain table shaped exactly for what the screen needs to show. The UI only ever reads from that table, never from the raw events.

open as a page

Why must the handler that applies events to a projection be idempotent, and what's a concrete technique to make an 'increment a counter' style update idempotent?

level: middleimportance: must knowfreq 75%

basics

~20 s

Idempotent means applying the same event twice gives the same result as applying it once. It's needed because message delivery isn't perfectly exactly-once — retries or crashes can cause the same event to be processed again, and without idempotency that would double-count or corrupt the read data.

open as a page

Why do CQRS read models commonly use denormalized data structures instead of the normalized tables typical of a write-side relational schema, and what does that cost you?

level: middleimportance: must knowfreq 75%

basics

~20 s

Query models copy data into whatever shape a screen needs and duplicate it in several places, so reads are fast lookups instead of expensive joins. The cost is extra storage and the risk that copies can go slightly out of sync with the real data.

open as a page

In a CQRS architecture, why must the query side never mutate state, even indirectly (e.g. lazy-computing and caching a value inside a query handler)?

level: middleimportance: must knowfreq 65%

basics

~20 s

A query is only supposed to look at data and return an answer — it should never change anything while doing so. If reading data quietly changes it, you lose the ability to trust that 'just reading' is safe to repeat, retry, or run in parallel.

open as a page

Once a team has CQRS in production — separate write and read models kept in sync via projections — what ongoing operational burdens does that dual-model setup create that a single shared model doesn't have, and how do teams typically manage them?

level: middleimportance: must knowfreq 65%

basics

~20 s

Now there are two models to deploy, monitor, and fix bugs in, plus a pipeline that keeps them in sync — and if that pipeline breaks or falls behind, the read side can go stale or wrong without anything obviously crashing.

open as a page

A client's call to submit a `ChargeCard` command times out on the network, so the client's retry logic resends the exact same command. How do you design the command-handling path so the retry doesn't charge the card twice, and where does that check actually have to live?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Give each command a unique ID from the client. Before doing the real work, the handler checks whether it has already processed this exact ID - if yes, it just returns the old result instead of charging again.

open as a page

An order-placement flow in a CQRS/event-driven system needs to reserve inventory, charge payment, and update a read-model 'order status' projection, each owned by a different service with its own write model. No single ACID transaction spans all three. What coordination pattern keeps this consistent, and what happens if the payment step fails after inventory has already been reserved?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Use a saga: a sequence of local transactions where each step tells the next one to proceed, coordinated either by a central orchestrator or by services reacting to each other's events. If payment fails after inventory was reserved, the saga runs a compensating action - releasing the reserved inventory - to undo that earlier step.

open as a page

In a CQRS system where the write side uses Event Sourcing, a user submits a command that changes data, and the client immediately re-queries a read model to display the result — but the change is missing or shows stale data. Why does this happen structurally, and what are the common ways teams handle it?

level: seniorimportance: must knowfreq 70%

basics

~20 s

Saving the change (writing the event) and updating what the screen shows (the read model) are two separate steps done by two separate pieces of code, and the second one takes a little time to catch up after the first. If you check the screen data too fast, you can catch it before it's updated.

open as a page

When a projection's logic has a bug and produced incorrect data, how do you safely rebuild it by replaying the event history, without taking the read side offline or serving inconsistent results mid-rebuild?

level: seniorimportance: must knowfreq 65%

basics

~20 s

You build a brand new copy of the projection from scratch by replaying all the past events into a fresh table, and only switch reads over to it once it's fully caught up — so users keep reading the old (working) version the whole time instead of seeing a half-rebuilt one.

open as a page

A colleague proposes introducing CQRS — separate write and read models — for a new internal admin tool that has low traffic and simple CRUD screens. What's your reasoning for pushing back, and what are the concrete signals that would change your mind?

level: seniorimportance: must knowfreq 70%

basics

~20 s

If reads and writes are low-volume and shaped similarly, splitting into two models just adds complexity (sync pipeline, staleness, more code) with no real benefit — you'd push back and ask what specific problem it's meant to solve.

open as a page

A team is about to apply CQRS with an eventually-consistent read model to a feature that checks a user's account balance before authorizing a withdrawal. As the architect reviewing the design, what makes this a case where you'd push back and require strong (immediate) consistency instead, and what are the concrete ways to get it without abandoning CQRS entirely?

level: principalimportance: must knowfreq 60%

basics

~20 s

Money is a case where showing the wrong, stale number can cause real harm, letting someone overdraw because the read side hadn't caught up yet. For that one check, read the current true value directly instead of trusting the lagging copy, even if the rest of the app still uses the fast, eventually-consistent read model.

open as a page

What's the difference between a 'live' subscription and a 'catch-up' subscription when a projection consumes an event stream, and why would a projection need both?

level: middleimportance: should knowfreq 55%

basics

~20 s

A live subscription gets new events as they happen, like watching a live feed. A catch-up subscription reads through the older, already-happened events first to get up to speed. A projection needs both so it can start from wherever it left off (or from the beginning) and then smoothly switch to real-time.

open as a page

Strict CQRS says a command handler should only accept or reject a command, not return domain data - commands are effectively void. In practice, a client submitting `CreateOrder` usually needs the new order's ID right away. What do teams typically do to reconcile that need with the 'commands don't return data' rule, and what does each option cost?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Either the client makes up the ID itself before sending the command, or the system bends the rule a little and hands back just the new ID, not the full order, as a lightweight acknowledgment.

open as a page

You're on call for a CQRS system where a projector consumes a Kafka topic to update a read-model database. Support reports that some users see stale search results for several minutes after editing a listing. What would you measure to confirm and bound the lag window, and what are two concrete mitigations if the backlog keeps growing under load?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Measure the gap between when an event was written and when the projector actually applied it (processing lag), plus how many unprocessed events are queued up (offset lag). If the backlog keeps growing, either process events faster with more parallel consumers, or make each projector write cheaper, for example by batching, so it can keep pace.

open as a page

You're designing the command/write side of a new CQRS system. Under what conditions is Event Sourcing genuinely worth adopting for that write side, and when would a simpler state-stored write model — a normal table updated in place, with domain events published afterward purely for integration — serve the domain just as well or better?

level: seniorimportance: should knowfreq 60%

basics

~20 s

Use Event Sourcing when you truly need a full history of every change, like an audit trail, or answering 'what did this look like last Tuesday.' If all you need is the current state and maybe a heads-up to other systems when something changes, a normal table that gets updated plus a 'this changed' notification is simpler and usually enough.

open as a page

When choosing a store for a new projection — e.g., relational table vs. search index vs. key-value cache — and deciding whether to maintain multiple projections off the same event stream, what factors drive the decision?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Pick the storage technology that matches how the data will actually be queried — fast lookups by ID want a key-value store, full-text search wants a search index, flexible reporting wants a relational table. It's normal to build several different projections from the same events if different features need different query shapes.

open as a page

A team builds three different read models (a SQL table for a dashboard, an Elasticsearch index for search, a Redis hash for a single-item lookup) all fed from the same stream of domain events. What problem does this design solve, and what new problems does it introduce?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Building several different read models (like a search index, a dashboard table, and a fast key-value lookup) all from the same underlying events lets each one be fast for its own job. The cost is you now have several copies of the truth to keep in sync and debug when they disagree.

open as a page

How does testing a CQRS system differ from testing a service with a single shared model, and what specific strategy would you use to test that a read-model projection ends up correct after a write, given the projection updates asynchronously?

level: seniorimportance: should knowfreq 45%

basics

~20 s

You test the write side and read side separately, then test the connection between them by writing data, waiting for the read side to catch up (or polling), and checking it matches — instead of one instant check like you would with a single database.

open as a page

A command handler for `PlaceOrder` first commits the new order state to its database, and then, as a second separate step, publishes the resulting `OrderPlaced` event to a message broker. What failure mode does this two-step design create, and what are the common ways teams close that gap?

level: principalimportance: should knowfreq 45%

basics

~20 s

If the app crashes right after saving to the database but before sending the event, the save happened but nobody downstream finds out. The write and the event need to succeed or fail together, not as two separate bets.

open as a page

showing 1–30 of 34