In a CQRS system, what is a 'read model' (also called a query model) and why is it usually shaped differently from the domain/write model?
answer
- read model = derived copy, not source of truth
- projector/event handler populates it
- eventually consistent with write side
- rebuildable from event log/write DB
- one shape per query need
basics
~20 sA read model is a copy of your data shaped exactly for what a screen or report needs to show, separate from the data structure used to save changes. It exists so reading is fast and simple, and writing stays focused on correctness.
solid answer
~40 sIn CQRS, the query model (read model) is a purpose-built representation of data optimized for a specific read use case — e.g., a denormalized table, document, or DTO shaped exactly like the UI or API response that consumes it. It's built by projecting events or replicating rows from the write side, not by querying the write model's normalized schema directly. Separating it from the command/write model lets each side evolve and scale independently: writes stay normalized and consistency-focused, reads stay denormalized and speed-focused. The read model is disposable — it can always be rebuilt from the source of truth (the write model or the event log).
go deeper
Should be able to say a read model is a separate, simplified copy of data made for reading, and that it's not where changes get made — 'you write to one place, you read from a shaped copy.'
Should explain concretely how the read model gets populated (projector/event handler, CDC, or sync job) and name at least one storage example (materialized view, document, cache).
Should reason about the consistency window between write and read model, know the read model is rebuildable from source of truth, and discuss when to choose which storage technology for a given read model.
Should discuss organizational/system-wide implications: how many read models a system should maintain, reconciliation/observability strategy across all of them, and how read-model design decisions affect the team's ability to evolve the domain model independently.
## What a read model is A **read model** — also called a **query model** or **projection** — is a data structure built and stored specifically to answer one class of question quickly, as opposed to being derived on the fly from a general-purpose, normalized data store. Concretely, it might be: - a SQL table with pre-joined, flattened columns - a JSON document in a document store - a row in `Redis` - a search index in `Elasticsearch` The shape of the read model is dictated by the consumer: if a UI screen shows 'customer name, last order date, total spent,' the read model can store exactly those three fields together in one row, so fetching them is a single indexed lookup rather than a query that joins `customers`, `orders`, and `order_items` and aggregates on the fly. ## How it gets populated The read model is populated by a process separate from the request that reads it. 1. **In an event-sourced system**, a 'projector' or 'denormalizer' subscribes to the stream of domain events (`OrderPlaced`, `OrderShipped`, `CustomerAddressChanged`) and, for each event, updates the read model's stored rows to reflect the new state — appending, updating counters, or recomputing derived fields. 2. **In a non-event-sourced CQRS system**, the same idea can be achieved with change-data-capture from the write database, database triggers, or an application-level 'after commit, also update the read table' step. Either way, the read model is a consequence of the write side, not an independent source of truth: if it is deleted, it can in principle be rebuilt by replaying history or re-syncing from the write store. ## Why the two sides are pulled apart This separation exists because commands and queries have opposing shape pressures. | The side that changes state | The side that answers questions | |---|---| | The command/write side wants normalization, strong consistency, and a schema that enforces business invariants (foreign keys, uniqueness constraints, transactional boundaries around an aggregate) so that a state transition is safe. | The query side wants the opposite: no joins at request time, no invariant-checking overhead, and a shape that mirrors what the caller will do with the data — a screen, a report, an API response. | Trying to serve both needs from one schema forces compromises on both sides: - either the write schema gets denormalized (risking data integrity) - or every read pays the cost of joining and computing at request time (hurting latency and scalability) CQRS's query-model pattern resolves the tension by letting each side have the schema it wants and paying a small synchronization cost to keep them related. ## The central trade-off The central trade-off is **consistency versus performance and complexity**. Because the read model is updated asynchronously after the write model changes, there is a window — often milliseconds, sometimes longer under load — during which the read model is stale relative to the write model. Client code and product requirements have to tolerate 'read your own writes' not being instantaneous, or the system has to add compensating UX (optimistic UI updates, polling, or routing a user's own just-written data back through the write path temporarily). There is also real infrastructural cost: - every read model is another copy of data that needs its own storage, its own migration story, and its own monitoring - every new read model is another piece of denormalization logic that can drift out of sync with the source of truth if the projector has a bug or misses an event ## Failure modes Failure modes cluster around synchronization. - A projector can crash mid-batch and leave a read model half-updated. - An event can be delivered twice and double-count an aggregate field like a running total. - An event can be missed entirely (e.g., a message bus outage) and leave the read model permanently behind until someone notices and triggers a rebuild. - The read model's schema can silently diverge from what the projector code assumes after a refactor, producing wrong-but-plausible-looking data that isn't caught by tests. Because read models look like 'just another table,' teams sometimes treat them as authoritative and start writing to them directly from application code, which breaks the entire model: now there are two paths that can change the same data, and the read model is no longer safely rebuildable. ## Seeing it in one system A concrete example: an online store's 'My Orders' page is backed by an `OrderHistorySummary` read model — one row per order with the customer name, item count, total, and status already flattened in, instead of querying the normalized orders/order_items/customers tables at request time. When a payment event fires, a projector updates that row's status field within a few hundred milliseconds. If the projector goes down for an hour, the fix is to bring it back up and let it catch up from its last processed offset, or, in the worst case, wipe the read model and replay the full order history to rebuild it — the write-side event log or database remains the ground truth throughout.
- If a read model can always be rebuilt from the write side, why not just delete it and query the write model directly whenever you need something new?Because rebuilding is often expensive and slow at scale, and the write model's normalized schema isn't shaped for arbitrary ad hoc queries — you'd be paying the join/aggregation cost on every request instead of once at write time. The read model exists precisely to avoid making every read pay for the write side's schema design; deleting it just moves the cost back to request latency.
- What has to be true about the projector logic for a read model to be safely rebuildable?It needs to be a pure, deterministic function of the source history — given the same sequence of events or the same write-side state, it must always produce the same read-model output. It also needs to be idempotent, so replaying overlapping or duplicate events doesn't corrupt the result.
- How would you detect that a read model has silently drifted from the write model in production?Run a periodic reconciliation job that recomputes a checksum or spot-checks a sample of records against the source of truth and alerts on mismatches, rather than relying on someone noticing wrong data by accident.
A read model is like a restaurant's printed menu board: the kitchen's inventory and recipes are the real source of truth, but customers don't dig through pantry shelves to figure out what's available — the board is a pre-formatted summary kept updated by staff, and if it's ever wrong or lost, someone reprints it from what the kitchen actually has.
saying these in an interview costs you the question
- Says the read model is 'the database' with no mention of it being derived from something else
- Assumes the read model is always perfectly in sync with the write model (no mention of eventual consistency)
- Can't explain how the read model gets populated/updated
- Treats the read model as safe to write to directly from arbitrary application code
- No sense that read models can be rebuilt/regenerated