Which categories of managed cloud services are commonly used to implement the event store and the projection pipeline in a CQRS-plus-event-sourcing architecture, and what role does each typically play?
answer
- four pieces: log, delivery, compute, read store
- CDC/change feed as lightweight alternative to a dedicated event store
- serverless functions as projectors
- read store chosen per query shape
- retention limits cap how far back you can replay
basics
~20 sCloud providers offer managed 'always-on' logs (streaming services) to hold the events, managed databases with built-in 'change feeds' to trigger updates automatically, serverless functions to run the update logic, and managed read stores (like document databases or search services) to hold the fast, query-ready copies.
solid answer
~40 sFour service categories typically appear. (1) Managed append-only event stores/streaming platforms hold the ordered event stream durably and let multiple consumers subscribe independently. (2) Change-data-capture / change-feed features on managed databases let teams derive an event stream from writes to a conventional store without a dedicated event store. (3) Serverless compute (functions triggered by new events) commonly implements the projector logic itself, scaling automatically with event volume. (4) Read-optimized managed stores — document databases, search services, caches, or data warehouses — host the resulting projections, chosen per consumer's query shape. Together these let a team assemble the pipeline from managed building blocks instead of operating an event store and projection workers themselves.
go deeper
Should recognize that cloud providers offer ready-made services for holding an event stream and for storing fast read copies, without needing to name specific products.
Should name the four pipeline roles (log/feed, delivery, projector compute, read store) and give at least one example service category for each.
Should compare CDC-on-existing-database versus dedicated event store as a real architectural decision, and discuss cost/retention tradeoffs of the managed approach.
Should weigh vendor lock-in, cross-service cost model at scale, and retention/rebuildability guarantees as first-class architectural risk when choosing the managed pipeline for a cloud platform.
## The four moving pieces CQRS with event sourcing needs four moving pieces working together: 1. somewhere durable to append events in order; 2. a way to deliver those events to interested consumers; 3. compute to run the projection logic; 4. somewhere query-optimized to land the resulting read models. Cloud platforms offer managed services for each piece, and most production implementations assemble the pipeline from these rather than hand-rolling an event store from scratch. ## The durable, ordered log The first piece is the durable, ordered event log itself. This is typically implemented with a managed streaming/pub-sub platform — a hosted Kafka-compatible service, or a cloud-native equivalent — that guarantees events are: - appended in order (usually per-partition or per-entity-stream); - retained durably; - readable by multiple independent consumers at their own pace. The defining property that makes these suitable as an 'event store' is that a consumer's read doesn't remove or alter the data — many projectors can each maintain their own read position ('offset' or 'cursor') into the same stream without interfering with each other. An alternative, increasingly common path is to skip a dedicated event log and instead use a **change-data-capture (CDC)** or 'change feed' feature built into a managed operational database: writes to a normal table or document store are also emitted, in order, as a change stream that downstream projectors can subscribe to. This lets a team get event-driven projections without redesigning the write side around an explicit event-sourcing model — a lighter-weight, more common entry point for many cloud teams than 'full' event sourcing. ## The projector The second piece is the projector — the code that consumes each event and updates a read model. In cloud-native deployments this is very commonly implemented as **serverless, event-triggered compute**: - a function that's automatically invoked per incoming event (or per small batch); - scales its concurrency with event volume; - is billed per invocation rather than for an always-on worker. This fits the projection workload well because event arrival is often bursty, and a function-based projector avoids provisioning a persistent worker fleet sized for peak load. Teams with stricter ordering, stateful-aggregation, or very-high-throughput needs sometimes instead run a dedicated, always-on stream-processing job rather than a per-event function, trading some operational overhead for tighter control over batching, state, and ordering guarantees. ## The read-model store The third piece is the read-model store, chosen per consumer's access pattern rather than uniformly. | Managed read store | Access pattern | |---|---| | A document database | suits a UI that needs to fetch a whole denormalized 'order summary' object by ID in one round trip | | A managed search service | suits full-text or faceted queries a document store handles poorly | | A key-value/cache service | suits ultra-low-latency point lookups on a hot path | | A managed data warehouse | suits heavy aggregation and reporting queries across many events | Because CQRS decouples the read side from the write side's schema entirely, nothing forces all projections onto the same storage technology — and in practice they usually aren't, precisely because each is optimized for a different query shape. ## Why this matters for a cloud team Why this matters for a cloud team specifically: assembling the pipeline from managed services shifts a large amount of operational burden — durability, replication, scaling the log itself, patching brokers — onto the cloud provider, leaving the team responsible mainly for projector logic and read-store schema design. The tradeoff is **cost model and lock-in**: managed streaming/CDC/serverless services bill on throughput, invocation count, and storage retention, which can get expensive at very high event volumes, and moving off a specific provider's event-store or CDC implementation later is nontrivial work. ## Failure modes Failure modes specific to this managed-service composition include: - CDC feeds silently dropping or reordering events during a source database failover if the feature's ordering/delivery guarantees aren't well understood; - serverless projector functions timing out or throttling under a sudden event burst, causing lag spikes; - retention limits on the underlying log meaning that if a projection falls behind for too long, the events it needs may have already expired and can no longer be replayed from that log, requiring a longer-retention archive tier for full rebuildability. ## Where it shows up A concrete real-world shape: a retail order system uses a managed database's native change feed to emit order-table writes as an ordered event stream, a set of serverless functions subscribed to that feed to update a 'current order status' document store for the customer app and a separate search-indexed store for internal support-agent lookups, with the change feed's extended retention tier kept available so either projection can be fully rebuilt from scratch if its logic changes.
- What's the practical difference between building on a dedicated event store versus a database's change-data-capture feed?A dedicated event store models events as the explicit source of truth from day one, giving strong guarantees about ordering, retention, and being designed for many independent replayable consumers. A CDC/change feed instead derives events from ordinary table writes on an existing operational database, which is a lighter lift for teams not ready to redesign around full event sourcing, but usually comes with tighter retention windows and delivery guarantees that are the database vendor's, not a purpose-built event log's.
- Why might a team choose a stateful stream-processing job over per-event serverless functions for a projector?When the projection logic needs to maintain running aggregates across many events with strict ordering and exactly-once-style guarantees, a dedicated stream-processing job can hold that state efficiently in memory across the stream, whereas invoking a stateless function per event forces re-fetching or re-persisting intermediate state on every call, which is slower and more expensive at high volume.
- What operational risk does event retention introduce for rebuildability?If a projector falls far enough behind, or a team decides to rebuild a projection long after events were emitted, and the underlying log or change feed has already expired those old events, there's nothing left to replay from — the projection can't be fully reconstructed. Teams mitigate this by archiving events to longer-retention storage specifically to preserve full replay capability.
Like a newswire service: the wire (managed log/CDC feed) carries stories in order to many subscribing newsrooms; each newsroom (a serverless projector) has its own editor turning wire stories into a locally-formatted product — a print layout, a website feed, a radio script — stored wherever suits that format best.
saying these in an interview costs you the question
- Assumes an event store must always be a purpose-built, dedicated product
- Doesn't know CDC/change feeds are a lighter-weight alternative on existing databases
- Thinks all projections should share one storage technology for consistency
- Unaware that streaming/CDC services have retention limits affecting replay
- Can't name any concrete category of cloud service used at any stage of the pipeline