When choosing a store for a new projection — e.g., relational table vs. search index vs. key-value cache — and deciding whether to maintain multiple projections off the same event stream, what factors drive the decision?
answer
- match store to query pattern, not write shape
- relational/document/kv/search/graph archetypes
- N projections = N handlers, N infra pieces
- independent lag per projection
- projection sprawl / ownership drift
basics
~20 sPick the storage technology that matches how the data will actually be queried — fast lookups by ID want a key-value store, full-text search wants a search index, flexible reporting wants a relational table. It's normal to build several different projections from the same events if different features need different query shapes.
solid answer
~50 sStore choice should follow the query pattern, not the write model's shape: point lookups by key favor a key-value store or cache, full-text/faceted search favors a search index, ad-hoc reporting and joins favor a relational store, graph traversal favors a graph DB. Because a projection is just a derived materialization, there's no reason to force one shape to serve every read need — it's normal and often preferable to run several independent projections off the same event stream, each optimized for one access pattern (e.g., a search index for admin search, a flat summary table for a customer dashboard, a Redis counter for a live badge). The cost is N times the infrastructure, N handlers to keep correct and idempotent, and N places that can drift or lag independently, so the decision balances query-fit gains against the operational multiplication of maintaining more moving read-side pieces.
go deeper
Should recognize that different kinds of queries (search vs. lookup vs. reporting) might want different kinds of storage.
Should be able to match a couple of common store types to the query patterns they fit and explain why one store rarely serves everything well.
Should articulate the operational cost side of multiple projections (N handlers, independent lag, ownership) alongside the query-fit benefits, and make a reasoned call for a given scenario.
Should set cross-team guidance on when a new projection is justified, own patterns for detecting/pruning sprawl, and think about system-wide consistency guarantees when features compose data from multiple independently-lagging projections.
## What the choice actually is Choosing a store for a projection means picking the persistence technology used to hold that projection's materialized data. The core principle in CQRS is that this choice should be driven **entirely by how the data will be queried**, not by whatever technology the write side already uses — CQRS explicitly permits, and often expects, the read store to diverge from the write store, since decoupling them is the whole point of the pattern. ## The store archetypes A few common store archetypes recur and each fits a different access pattern. | Store | Access pattern it serves | |---|---| | **A relational store** | gives flexible ad hoc querying, joins, and transactional guarantees, making it a reasonable default when query patterns aren't fully settled yet or reporting-style access is needed | | **A document store** | fits data that's naturally fetched and updated as one self-contained nested unit — for example, an order together with its embedded line items, matching an API response shape one-to-one without needing joins | | **A key-value store or cache** | is built for fast, O(1)-ish point lookups by a known key, and is poor at range or secondary-attribute queries | | **A search engine** | is built for full-text search, faceting, and relevance ranking, at the cost of weaker consistency and transactional guarantees | | **A columnar or OLAP-style store** | fits analytics workloads with large scans and aggregations | | **A graph database** | fits relationship-traversal queries | Running 'multiple projections' means having separate handlers — possibly each with its own consumer group and checkpoint, even when reading the same underlying event stream — that each independently write into one of these stores. ## Why one shape rarely serves every need This exists because a single query shape rarely serves every read need well. Forcing all reads back through one general-purpose, normalized store reproduces the exact problem CQRS was meant to solve in the first place — slow joins, a shape that fits no particular feature especially well. Because adding a new projection is comparatively cheap — it's just another independent consumer of the same event stream, requiring no coordination with the write model's schema — the natural design that emerges is **'one projection per genuinely distinct query need,'** letting each feature or team own its read shape without having to negotiate schema changes with everyone else who happens to read related data. ## Benefits, and the costs that multiply The benefits are real: each projection can be scaled, cached, indexed, and evolved completely independently of the others, and can be changed or dropped without touching anything else. The costs are equally real and multiply with each additional projection: - more infrastructure to run, monitor, back up, and secure; - N separate handlers, each of which needs to be correctly idempotent and versioned, multiplying the surface area for bugs and drift; - a harder system-wide consistency story, because different projections can lag the event stream by different, uncorrelated amounts, meaning two features can present a user with two different 'current' pictures of the same underlying event at the very same moment. Building a genuinely new projection is also bounded by the rebuild/replay cost discussed elsewhere — if the event history is large, standing up a brand-new projection isn't instantaneous. ## Failure modes A handful of failure modes show up repeatedly. - **'Projection sprawl'** is common — a projection outlives the feature that needed it, nobody remembers who owns it or why it still exists, and it keeps quietly consuming infrastructure and consumer capacity. - **Two sibling projections can show contradictory numbers** in the same UI — a dashboard total from one projection disagreeing with a list count from another — purely because they lag independently, which damages user trust even though nothing is actually 'wrong.' - **Teams sometimes pick a store for familiarity rather than fit** — cramming a search feature into a relational LIKE-query projection instead of standing up a real search index — and the mismatch is often only discovered once real usage volume exposes the performance gap. - **As projection count grows**, a shared event broker or consumer capacity budget can become under-provisioned, causing all projections to start lagging together rather than independently. ## A worked example A concrete, realistic example: a ride-sharing platform might run - a relational projection for 'trip history' that supports joins and filters for a rider's account page, - a search-engine projection for support agents to search across trip and message text, - a fast key-value projection holding a driver's current live location and status for sub-second reads. All three are fed by the same underlying `TripRequested/DriverAssigned/TripCompleted` event stream. This general pattern — a shared event backbone fanning out into several purpose-built read stores — is broadly consistent with how large event-driven backends at companies like Uber and Lyft have been described in public engineering write-ups.
- Why might two different projections built from the same event stream show a user two slightly different 'current' numbers at the same moment?Each projection has its own independent consumer/checkpoint and can lag the event stream by a different amount depending on its processing cost, store latency, or recent incidents, so at any instant one projection may have processed events another hasn't yet — there's no built-in guarantee that sibling projections stay in lockstep with each other.
- What's a concrete reason to prefer a document store over a relational table for a specific projection?When the read shape matches a single, self-contained nested structure that's always fetched and updated as one unit — like an order with its embedded line items exactly as the API returns it — a document store avoids joins entirely by storing that shape directly, at the cost of being worse at ad hoc cross-entity queries a relational store handles naturally.
- How would you decide whether a new read requirement justifies a brand-new projection versus reusing an existing one?Check whether the existing projection's shape and store can serve the new query with acceptable performance and without contorting its schema for an unrelated use case; if the access pattern genuinely differs, like needing full-text search when the existing one is a flat summary table, or reuse would create tight coupling between unrelated features, a new dedicated projection is usually worth the added operational cost.
It's like a newsroom producing the same reported facts into different formats for different audiences — a full written article (relational), a searchable archive (search index), and a quick breaking-news ticker (cache) — each optimized for how that audience consumes it, all sourced from the same underlying reporting.
saying these in an interview costs you the question
- defaults to relational for every projection regardless of query pattern
- doesn't recognize that multiple independent projections can lag by different amounts
- treats adding a new projection as free with no operational cost
- can't name at least two different store archetypes and a query pattern each fits