Why do CQRS read models commonly use denormalized data structures instead of the normalized tables typical of a write-side relational schema, and what does that cost you?
answer
- duplicate data, pre-shaped per query
- cost = storage + write amplification + staleness
- one read model per real access pattern, not per screen
- fan-out-on-write example (feeds)
- propagation should be async, not inline in the write transaction
basics
~20 sQuery models copy data into whatever shape a screen needs and duplicate it in several places, so reads are fast lookups instead of expensive joins. The cost is extra storage and the risk that copies can go slightly out of sync with the real data.
solid answer
~40 sDenormalization means storing the same underlying facts multiple times, pre-shaped and pre-joined for a specific query, instead of normalizing to eliminate duplication. On the read side of CQRS this is deliberate: a query like 'show a customer's order history with product names and totals' would otherwise require joining several normalized tables at request time, which gets expensive at scale. By maintaining a flattened, purpose-built projection updated whenever the underlying data changes, reads become O(1) lookups. The cost is storage duplication, the need to keep every denormalized copy in sync (more moving parts, more code to maintain), and eventual consistency — the projection reflects the write model with some lag. It's the classic space/consistency-for-speed trade, and it's why teams build one read model per real query pattern rather than one generic one.
go deeper
Should be able to say that denormalized data is duplicated on purpose to make reads faster, in plain terms.
Should explain the specific costs — storage duplication, write amplification, eventual consistency — and give an example of a query that's expensive normalized but cheap denormalized.
Should reason about when denormalization is worth it versus not (traffic patterns, query complexity), and discuss keeping propagation asynchronous to avoid coupling write latency to N projection updates.
Should discuss read-model sprawl as an organizational risk, propose governance (only build a projection per justified access pattern), and connect the trade-off to concrete system examples like fan-out-on-write feeds.
## What denormalization means on the read side **Denormalization** on the read side of CQRS means deliberately storing the same underlying facts more than once, each time pre-shaped for a particular query, rather than storing each fact exactly once in a normalized schema and reconstructing views at request time via joins. Concretely: instead of a normalized set of tables — `orders`, `order_items`, `products`, `customers` — linked by foreign keys, a read model might store a single `OrderSummary` row per order that already contains the customer's name, the list of product names and quantities, and a precomputed total, all copied in at write time. If a customer has 500 orders, the customer's name is now duplicated across 500 rows. That duplication is deliberate: it means answering 'list this customer's orders with product names' is one indexed read of one table, with no join and no runtime aggregation, however large the dataset gets. ## Why teams accept the duplication The reason to do this is **query efficiency and predictable latency at scale**. A normalized schema is optimized for write-side integrity — updating a customer's name in one place updates it everywhere it's referenced — but that same design means every read that needs a 'flattened' view must pay a join cost at request time. - Joins across large tables get slower as data grows. - Complex read patterns (aggregations, cross-entity lookups, full-text search) often need indexes or storage engines the write-side database wasn't built for. CQRS's answer is: don't try to serve every kind of read from one general-purpose schema; instead, build a purpose-specific projection per real query pattern, computed once when data changes rather than recomputed on every read. This is sometimes phrased as **'one read model per view'** — a search page, a dashboard, and a detail page can each have their own denormalized shape, even though they're all derived from the same underlying write-side facts. ## What it costs The cost side of this trade is threefold. 1. **First, storage:** every denormalized copy is additional bytes, and a system with many read models for the same underlying entity can end up storing the 'same' data five or six times over in different shapes. 2. **Second, write amplification:** a single write-side event (e.g., a customer renames themselves) may need to propagate updates into every read model that copied the customer's name — one write becomes N updates, and each of those N updates is a place a bug or an outage can cause drift. 3. **Third, eventual consistency:** because updating the denormalized copies happens asynchronously after the write, there's a window where different read models — or the read model and the write model — briefly disagree, which every consumer of that data has to be designed to tolerate. ## Failure modes Failure modes follow directly from write amplification and duplication. If the process responsible for propagating a change to a given projection fails partway — crashes after updating three of five read models — the system now has silently inconsistent read models until something notices and repairs it, and there's often no built-in signal that this happened, because each read model still 'looks' like valid data. As the number of read models grows, so does the operational burden — every new denormalized view is: - another schema to migrate when a business rule changes - another set of projector code to maintain - another thing to keep in mind when the underlying write-side data model changes shape A common anti-pattern is **'read model sprawl,'** where the team keeps adding projections faster than they can keep them tested and consistent, until nobody fully trusts any of them and staff resort to querying the write database directly 'just to be sure,' defeating the pattern's purpose. ## Where the trade-off shows up at scale A well-known real-world example of this trade-off is a social-media timeline built with **fan-out-on-write**: rather than computing 'show me my feed' by querying all the accounts a user follows at read time (expensive as follow-counts grow), the system denormalizes by writing a copy of each new post into every follower's precomputed timeline read model at post time. This makes reading a timeline a cheap single lookup, at the cost of a write that can fan out into millions of copies for a popular account, and a real engineering problem around handling celebrities' posts differently (hybrid push/pull) because pure fan-out-on-write doesn't scale to that write amplification. The same tension — cheap, denormalized reads bought with more expensive, more failure-prone writes — shows up in any CQRS system that builds dedicated read projections, just usually at a smaller scale than a global social feed.
- If a system needs five different read models for the same underlying entity, how do you decide which ones are actually worth building versus over-engineering?Build a dedicated read model only for query patterns with real, distinct performance or shape requirements — a search page and an admin table have genuinely different needs, but two screens that show nearly the same fields with minor formatting differences can usually share one projection. The rule of thumb is: build one per access pattern that would otherwise require a join or aggregation the normalized schema can't serve cheaply, not one per screen.
- How do you keep write amplification from becoming a bottleneck when one write-side change needs to update several denormalized copies?Make the fan-out asynchronous and independently retryable per read model — publish one event and let each projector consume it on its own schedule — rather than doing all N updates synchronously inside the original write transaction. That keeps the write path fast and isolates a slow or failing projector from blocking the others.
- What's a lighter-weight alternative to full denormalization when the query pattern doesn't justify a dedicated read model?A database materialized view or a cache-aside layer over the normalized data can give most of the read-speed benefit for infrequent or low-traffic queries, without committing to maintaining a fully separate projector and store.
It's like keeping a photocopy of a document at every desk that needs to reference it, instead of making everyone walk to the one filing cabinet. Everyone reads faster, but now every edit to the original means walking around updating every photocopy, and if you miss a desk, that person is working from an outdated copy without knowing it.
saying these in an interview costs you the question
- Thinks denormalization means 'no downside, just faster'
- Can't name what has to keep the copies in sync
- Doesn't mention eventual consistency as a cost
- Suggests building a separate read model for every single screen without weighing the cost
- No awareness of write amplification when a single fact changes