In an event-sourced cloud service — where the write side persists state changes as an ordered log of events rather than overwriting rows — how are CQRS read-model projections typically built and kept up to date from that event log?
answer
- projector = subscriber + reducer
- log is source of truth, projections are disposable
- multiple projectors, one log
- lag is the tell
- poison-pill events stall a projector
basics
~20 sA separate background process reads new events one by one from the log, in order, and updates a query-friendly copy of the data (a projection) to match. That copy is what queries actually read from.
solid answer
~40 sEach event (e.g., 'OrderPlaced', 'OrderShipped') is appended immutably to an event store/log in the order it occurred. A projector — a subscriber process — reads that log, applies each event to an in-memory or persisted read model incrementally (event-by-event, like a reducer), and writes the updated result into a query-optimized store (document DB, search index, cache). Multiple projectors can consume the same log independently to build different denormalized views for different consumers. Because projection is asynchronous, there's a lag between an event being appended and the corresponding read model reflecting it — this is the eventual consistency inherent to the pattern. Projections are also disposable: since the event log is the source of truth, a projection can be dropped and rebuilt from scratch by replaying the log from the beginning.
go deeper
Should grasp that a background process reads events in order and updates a separate copy that queries actually hit — and that this update isn't instant.
Should describe the projector as a per-event reducer writing to a query-optimized store, and know that projections can be dropped and rebuilt from the log.
Should discuss concrete failure modes — lag, poison-pill events, schema/versioning drift — and how they're monitored and mitigated in production.
Should reason about running many independent projectors at organizational scale: ownership boundaries, replay cost/duration on large logs, and versioning strategy across teams consuming the same event stream.
## What the write side stores In an event-sourced write model, instead of storing **current state** as mutable rows that get overwritten, the system stores every state-changing fact as an immutable event — `OrderPlaced`, `ItemAdded`, `OrderShipped` — appended in order to a per-entity or per-stream log. - The event store is **append-only**: nothing is ever updated or deleted, only added. - Current state, when needed on the write side, is derived by replaying an entity's events in order (or reading the last snapshot plus events since). - CQRS read models — the query-optimized views consumers actually query — are built from that same event log by a separate mechanism called a **projection** (sometimes called a 'read model builder' or 'event handler'). ## How a projection works Mechanically, a projection is a subscriber process that reads events from the log in the order they were appended, typically via a subscription mechanism: - a change feed; - a consumer group on a streaming platform; - a polling cursor against the store. For each event it receives, the projector applies a small piece of logic (structurally like a reducer/fold function: given the current view state and the new event, produce the next view state) and writes the result into a query-optimized store: | Query-optimized store | Shaped for | |---|---| | a document database | a UI list view | | a search index | full-text queries | | a relational table | reporting joins | | a simple cache | a hot lookup path | Because the same event log can be read by multiple independent projectors, a single write (say, `OrderShipped`) can simultaneously update a customer-facing order-status view, an internal fulfillment queue, and an analytics aggregate — each shaped completely differently, each maintained by its own projector, all deriving from the identical stream of facts. ## Two problems solved at once This design exists to solve two problems at once. 1. **First**, it gives the write side a complete, auditable history instead of only a current snapshot — useful for compliance, debugging ('why is this order in this state?'), and temporal queries ('what did this account look like last Tuesday?'). 2. **Second**, and more relevant to the CQRS bridge specifically, it gives you a clean, uniform integration point for building arbitrarily many read views without touching write-side domain logic: any new reporting or UI need is just a new projector reading the same log, not a schema migration on the transactional store. ## The central tradeoff The central tradeoff is **projection lag** and the operational burden of the pipeline itself. Consuming, transforming, and writing events all take time, so there is always some delay — usually milliseconds, but potentially much longer under load, backpressure, or partial outage — between an event being appended and a read model reflecting it. This is the eventual consistency users notice as 'I just did X and don't see it yet.' Production systems mitigate this with lag monitoring/alerting, and, when a specific interaction truly needs read-your-write guarantees, by routing that one read to the write side or a synchronously-updated path instead of the async projection. ## Failure modes Failure modes show up in a few characteristic ways. - A projector can crash or fall behind under load, silently growing lag until someone notices stale data in production. - A **poison-pill event** — one that a projector's logic can't handle, due to a bug or an unexpected event shape — can block that projector's cursor entirely if events must be processed strictly in order, halting all downstream updates for that view until it's fixed or skipped. - **Schema evolution** is another recurring failure mode: if an event's shape changes over time (a field renamed, a new event type introduced), projectors written against the old shape can silently misinterpret new events unless the team maintains explicit event versioning/upcasting. - And because projections are derived, not authoritative, a bug in projector logic doesn't corrupt the source of truth — but it does corrupt what users see until the projection is fixed and rebuilt, which for large logs can itself be a slow, resource-intensive operation. ## Where it shows up A concrete real-world shape of this: a ride-sharing platform's trip service appends events like `TripRequested`, `DriverAssigned`, `TripCompleted` to an event log, often backed in cloud deployments by a managed streaming service. One projector maintains a low-latency 'current trip status' view in a fast key-value store for the rider's live-tracking screen; a completely separate projector aggregates completed trips into a reporting store for finance. Both consume the identical event stream, update independently, and can be rebuilt independently from the log's history if their logic or storage technology changes — which is the whole point of decoupling the read side from the write side.
- What happens to a read model if a projector's logic has a bug and is later fixed?Because the event log is the durable source of truth and the read model is just a derived, disposable artifact, the fix is to correct the projector logic, then replay the entire event log (or the relevant portion) from the beginning to rebuild the read model from scratch. The corrected view then reflects all history correctly, without needing to touch the write side.
- How do teams avoid a slow projector blocking a fast consumer?Each projector typically maintains its own independent cursor position in the log and writes to its own dedicated read store, so one projector falling behind doesn't affect others reading the same log. Consumers query the projector's read store directly, not the log, so a slow projector only makes its own view stale — it doesn't throttle other projections or the write path.
- Why might a team choose different storage technologies for different projections off the same event log?Because each read model is shaped for a specific access pattern — a search index for full-text queries, a key-value store for point lookups by ID, a columnar store for aggregation-heavy reporting — and CQRS's whole benefit is letting each projection pick the storage best suited to its query shape without any of them needing to match the write side's schema or each other's.
Like a bank statement built from a ledger of transactions: the ledger (event log) never changes once written, and your 'current balance' view is just a running total someone recomputes by walking the ledger in order — if that calculation is wrong or out of date, you fix the calculation and rerun it, you never edit the ledger.
saying these in an interview costs you the question
- Thinks projections update the event log rather than being built from it
- Believes a read model, once built, never needs to be rebuilt
- Doesn't mention that projection lag exists or is asynchronous
- Assumes one projector must serve all read use cases
- Can't explain what happens when projector logic has a bug