In a CQRS system where the write side uses Event Sourcing, how does a read model — for example, a 'customer order summary' table used by the UI — get built and kept up to date from the event log? Walk through the mechanism end to end.
answer
- projector = subscriber to event stream
- checkpoint tracks read position
- one stream -> many projections
- eventual consistency lag
- rebuild by replay on logic change
basics
~20 sA separate small piece of code called a projector watches the event log for new events. Whenever it sees an event it cares about (like 'order shipped'), it updates a plain table shaped exactly for what the screen needs to show. The UI only ever reads from that table, never from the raw events.
solid answer
~60 sA projection is a subscriber process that consumes the event stream — either all events or a filtered subset — in order, and for each event applies a handler that updates a denormalized read store (a SQL table, a document, a search index, a cache). The read store's schema is shaped entirely around the query it needs to answer, not around the aggregate's internal structure, so the same event stream can feed many differently-shaped projections simultaneously. The projector tracks its own read position (a checkpoint) in the stream so it can resume after a restart without reprocessing or skipping events. Because the projector runs asynchronously relative to the command handler that appended the event, the read model is eventually consistent: there's a lag, typically milliseconds to low seconds, between an event being written and the read model reflecting it. If projection logic has a bug or needs to change shape, the standard fix is to discard the read store and rebuild it by replaying the event stream from the beginning.
go deeper
Should describe, at a high level, that a separate process listens to events and updates a table for reading. No need to mention checkpoints or idempotency.
Should explain the projector/checkpoint mechanism and that many projections can read the same stream independently.
Should discuss eventual consistency consequences for API/UX design and operational failure modes like projector lag and poison events.
Should be able to design a rebuild/versioning strategy for projections at scale (blue-green rebuild, dual-write during cutover) and reason about consumer-group scaling of projectors under high event volume.
## What a projection is Once a write model persists state as a stream of events, the read side's whole job is to turn that stream into whatever shape is fastest and simplest for a given query, and it does this through a mechanism usually called a **projection** (or a 'projector' when referring to the running process). A projection is nothing more than a small piece of code that subscribes to events — either the full event stream or a filtered slice of it, such as only events belonging to the `Order` aggregate type — and, for each event it receives, applies a handler function that updates some read-optimized storage. Concretely: - an `OrderPlaced` event handler might insert a new row into an `order_summary` table; - an `ItemAdded` event handler might update a line-items JSON column on that row; - an `OrderShipped` event handler might flip a status column and set a `shipped_at` timestamp. None of this logic runs inside the command path — it's entirely separate code, running in its own process or thread, reading from the event store the same way any other consumer would. ## The load-bearing details The mechanism has a few load-bearing details that are easy to miss. 1. First, projections consume events **in order**, per stream (and often globally too, depending on the guarantee needed), because applying `OrderShipped` before `OrderPlaced` would corrupt the read model. 2. Second, every projector tracks a **checkpoint** — a durable pointer to the last event position it successfully processed — so that if the projector process crashes and restarts, it resumes exactly where it left off instead of reprocessing everything or silently skipping events. 3. Third, and most importantly for why teams adopt this pattern at all, a single event stream can feed an unlimited number of independent projections simultaneously: the same OrderPlaced/ItemAdded/OrderShipped events might simultaneously update a customer-facing order-history table, a warehouse picking-queue table, and a finance revenue rollup, each with a completely different schema, each owned by a different team, none of them aware of or coupled to the others. ## The problem it exists to solve This exists to solve the query problem that a pure event log creates. An append-only log is an excellent write model — full history, no lost updates, natural audit trail — but a poor query model, because answering something as simple as 'list this customer's orders, most recent first' would otherwise require scanning and replaying every event for every order the customer ever placed. Projections pay that replay cost once, incrementally, as events arrive, and store the result in a shape a normal index or SQL query can serve in milliseconds. ## The trade-off The trade-off is consistency and operational surface area. Because the projector runs asynchronously relative to the command handler, there's a real window — typically milliseconds under normal load, but potentially much longer under backpressure or an outage — during which the event exists in the log but hasn't yet been reflected in the read model. Any UI or API that reads from a projection immediately after issuing the command that produced it can observe **stale data**. Handling this well requires deliberate UX or API design. Operationally, every projection is another moving part: - another consumer to monitor for lag; - another checkpoint to persist reliably; - another schema to migrate when query needs change. ## Failure modes in production Failure modes in production tend to cluster around three things: - **projector lag under load** (events pile up faster than the projector can apply them, and the read model falls further behind); - **poison events** (a malformed or unexpected event crashes the projector's handler and, if not isolated, halts the whole projection, silently freezing every read model it feeds); - **drift after a bug fix** (once a projection's handler logic changes, the existing read store reflects the old logic for old data and the new logic only going forward, unless the team rebuilds the store from scratch by replaying the full stream). A well-known real-world pattern here is Marten (a .NET library built on PostgreSQL) or Axon Framework (Java), both of which formalize exactly this projector-plus-checkpoint mechanism as a first-class citizen, letting teams declare a projection and have the framework manage subscription position and rebuild-on-demand.
- What happens to a read model if the projector process crashes mid-stream and is restarted?It resumes from its last durably-stored checkpoint rather than starting over, so it reprocesses at most the events since that checkpoint. As long as the projection's event handlers are idempotent (safe to apply twice), reprocessing a small overlap of already-applied events causes no harm, which is why idempotent handlers are a standard requirement for projection code.
- If two different read models are built from the same event stream, do they update at exactly the same moment?No — each projection runs as its own independent consumer with its own checkpoint and its own processing speed, so one read model can be several events ahead of another at any given instant. Callers should never assume two projections are mutually consistent with each other at a point in time.
- Why would a team choose to rebuild a projection from scratch rather than patch it in place?If the bug affected how historical events were interpreted, patching only fixes new events going forward, leaving old rows permanently wrong; replaying the whole stream from the beginning through the corrected handler produces a read model that's correct for every event ever recorded, old and new alike.
A projection is like a news aggregator site that watches a wire service feed (the event log) and continuously rewrites its own homepage sections — 'Top Stories,' 'Sports,' 'Local' — each shaped completely differently from the raw wire feed and from each other, even though all three are built from the exact same incoming stories.
saying these in an interview costs you the question
- Thinks the read model is updated synchronously as part of the same transaction as the command
- Doesn't mention that a projector needs to track its position/checkpoint
- Assumes there's only ever one read model per aggregate
- Can't explain why stale reads happen right after a write