skip to content

questions

6

In a CQRS system, what is a 'projection', and why do teams build one instead of just querying the write-side data directly?

level: juniorimportance: must knowfreq 70%

answer

  1. write model vs read model split
  2. denormalized, query-shaped store
  3. eventually consistent
  4. one event source, many projections
  5. checkpoint/offset per projection

basics

~20 s

A projection is a read-only copy of your data, shaped for easy querying, that's kept updated by watching events from the part of the system that handles writes. Teams build it because the write data is often shaped for correctness, not for fast, flexible reading.

solid answer

~40 s

A projection is a denormalized, query-optimized read model built by consuming a stream of domain events and applying them to a store shaped for specific query patterns — flat tables, search indexes, key-value lookups, pre-aggregated counters. In CQRS the write side (the command/aggregate model) is optimized for enforcing invariants and consistency, not for arbitrary reads: it's often normalized, transaction-scoped, and hostile to cross-aggregate joins. Projections decouple the read shape from the write shape, so each query pattern gets its own store (e.g., Postgres for reporting, Elasticsearch for search, Redis for a leaderboard) without touching the write model. The cost is eventual consistency: there's a lag between a command committing and every projection reflecting it, which callers must be able to tolerate.

go deeper

for a junior

Should be able to state that projections are separate read-optimized copies kept in sync via events, and name eventual consistency as the trade-off, without needing to design one.

for a middle

Should describe the mechanism (handler subscribes to events, applies to a store) and be able to justify when a projection is worth the extra infrastructure vs. querying the write side directly.

for a senior

Should discuss multiple projections per event stream, checkpoint/offset tracking, and concrete consistency-handling strategies for read-your-writes scenarios.

for a principal

Should reason about projection strategy at a system level — which stores fit which query shapes, how to bound staleness for SLAs, and how projection design affects team boundaries and operational cost.

## What a projection is **CQRS** (Command Query Responsibility Segregation) splits a system's write path (commands, which mutate state through an aggregate/domain model enforcing invariants) from its read path (queries, which just need to answer questions fast). A **projection** is the concrete artifact on the read side: a specific, materialized read model built by subscribing to a stream of events emitted whenever the write side commits a change, and applying those events to a store shaped for one particular access pattern. It's called a 'projection' because it takes the full richness of the event history and projects it down into whatever narrow shape a given query actually needs — much like a mathematical projection reduces a higher-dimensional object onto a simpler plane. ## How one is actually driven Mechanically, a projection is driven by a **handler**: a piece of code subscribed to an event stream (a message broker topic, an event store's subscription API, or an outbox table being polled). For every event it receives — say, `OrderPlaced` or `OrderCancelled` — it applies a transformation into its store: - upsert a row - increment a counter - add a document to a search index It tracks a position or **checkpoint** (an offset, a sequence number) so that if it restarts, it knows where to resume rather than reprocessing everything or skipping events. The store underneath is whatever best fits the query: - a flat relational table for a dashboard - a document store matching an API response shape one-to-one - a search index for full-text lookup - a Redis counter for a live badge ## Why the two sides are split at all The reason this separation exists is that the write model and the read model have genuinely conflicting design pressures. The write model wants to be normalized, scoped to one aggregate per transaction, and structured to make invariants easy to enforce — it is not optimized to answer 'give me this customer's order history with totals, sorted, paginated, and filterable by status' efficiently. If you try to serve every read need from that same model, you either: - bloat it with query-specific fields it has no business owning, or - pay for expensive joins and aggregations on every request against the transactional store, competing with the writes themselves for capacity. Separating the two lets the read side scale independently — add read replicas, cache aggressively, pick a completely different storage engine — without ever touching the write model's schema or logic, and lets the read shape evolve freely as UI/API needs change. ## The central trade-off The central trade-off is **eventual consistency**. After a command commits, there is a window — often milliseconds, sometimes longer under load or backlog — before the projection reflects that change. Any caller reading from the projection during that window sees stale data. This forces explicit design decisions: does the UI need to see its own just-made change immediately (a 'read-your-writes' requirement), and if so, how is that satisfied? - by returning the result directly from the command response rather than re-querying, or - by having the client wait/poll until the projection reaches a known position Beyond consistency, adopting projections also means adopting more infrastructure: an event bus or subscription mechanism, one or more separate stores, and checkpoint tracking, all of which need monitoring and can fail independently of the write path. ## Failure modes in production In production, a handful of failure modes recur. 1. **Projection lag** spikes under load — the consumer falls behind the event stream, and users start noticing stale reads, sometimes without any alarm firing because 'it's still working, just slow.' 2. **Silent divergence** is a subtler failure: a bug in the handler mishandles one event type (say, an update event is treated like a create), and the projection quietly drifts from what the write side actually reflects, sometimes for a long time before anyone notices the numbers look wrong. 3. **Checkpoint loss** is another: if a consumer crashes and its last-committed position wasn't durably persisted, restart either reprocesses already-applied events (fine only if the handler is idempotent) or skips ahead and misses some. 4. A malformed or unexpectedly-shaped event can throw inside the handler and stall the whole consumer behind it — a **'poison message'** blocking everything queued after it. ## A worked example A concrete example: an e-commerce system's write side enforces order invariants (can't ship what wasn't paid, can't cancel what already shipped) through a normalized Orders write model. - A projection handler subscribed to `OrderPlaced/OrderShipped/OrderCancelled` maintains a denormalized `OrderSummary` table feeding the customer's dashboard. - A second, independent handler on the same events feeds an Elasticsearch index used by support agents to search orders by free text. - A third maintains a Redis counter for 'orders today' shown on an internal ops screen. Each projection is purpose-built, independently scalable, and disposable — if one has a bug, it can be fixed and rebuilt from history without touching the other two or the write model at all.

  • If a projection is eventually consistent, how would you handle a user who just submitted a form and immediately expects to see their own change reflected?
    Common approaches are: have the command handler return the resulting state directly in its response instead of re-querying the projection; route that specific read to a synchronously-updated cache or the write-side store; or track a 'read-your-writes' token (the event's position) and have the client wait/poll until the projection catches up to it. Which one you pick depends on how strict the UX requirement is and how much latency you can tolerate.
  • Can a single write model have zero projections?
    Yes — a bare CQRS split where queries just hit the write-side tables directly is common early on, especially if that shape is 'good enough' for the few queries needed. Projections earn their keep once read patterns diverge enough from the write shape or read volume needs independent scaling; adding them later is a normal evolution, not a design flaw.
  • What's the difference between a projection and a simple database view?
    A DB view is computed synchronously at query time from the current write-side data — always consistent, but limited to what SQL/joins can express and bound by the write DB's engine and load. A projection is materialized ahead of time by an event handler into its own store, which can use a totally different engine (search index, cache, graph DB) and query shape, at the cost of being stale by the processing lag.

Think of a library's card catalog: the shelves (write model) hold the actual physical books in whatever arrangement is easiest to maintain, while the card catalog (projection) is a separate, purpose-built index that's updated whenever books are added or moved, so patrons searching by author or subject don't have to walk every shelf.

saying these in an interview costs you the question

  • describes projections as just a cache with no mention of eventual consistency
  • assumes the read and write model must be the same schema
  • can't explain why you'd want more than one projection
  • thinks projections are updated by the same transaction as the write
  • no mention of an event/change stream driving the update

context

open as a page

Why must the handler that applies events to a projection be idempotent, and what's a concrete technique to make an 'increment a counter' style update idempotent?

level: middleimportance: must knowfreq 75%

basics

~20 s

Idempotent means applying the same event twice gives the same result as applying it once. It's needed because message delivery isn't perfectly exactly-once — retries or crashes can cause the same event to be processed again, and without idempotency that would double-count or corrupt the read data.

open as a page

When a projection's logic has a bug and produced incorrect data, how do you safely rebuild it by replaying the event history, without taking the read side offline or serving inconsistent results mid-rebuild?

level: seniorimportance: must knowfreq 65%

basics

~20 s

You build a brand new copy of the projection from scratch by replaying all the past events into a fresh table, and only switch reads over to it once it's fully caught up — so users keep reading the old (working) version the whole time instead of seeing a half-rebuilt one.

open as a page

What's the difference between a 'live' subscription and a 'catch-up' subscription when a projection consumes an event stream, and why would a projection need both?

level: middleimportance: should knowfreq 55%

basics

~20 s

A live subscription gets new events as they happen, like watching a live feed. A catch-up subscription reads through the older, already-happened events first to get up to speed. A projection needs both so it can start from wherever it left off (or from the beginning) and then smoothly switch to real-time.

open as a page

When choosing a store for a new projection — e.g., relational table vs. search index vs. key-value cache — and deciding whether to maintain multiple projections off the same event stream, what factors drive the decision?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Pick the storage technology that matches how the data will actually be queried — fast lookups by ID want a key-value store, full-text search wants a search index, flexible reporting wants a relational table. It's normal to build several different projections from the same events if different features need different query shapes.

open as a page

At scale, a projection consumes events from a partitioned stream (e.g., a Kafka topic with many partitions) where global ordering isn't guaranteed across partitions, only within each one. How does this constrain how you key and design a projection handler, and what breaks if you get it wrong?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

When a stream is split into partitions for scale, events are only guaranteed to arrive in order within the same partition, not across all of them. So a projection has to make sure all events about the same thing, like the same customer or order, always land in the same partition, or it might process them out of order and end up with wrong, garbled results.

open as a page