When a brand-new projection needs to be built for a system that already has years of events in its log, how does a 'catch-up subscription' take it from zero to fully up to date, and then keep it live afterward?
answer
- two phases, one code path
- position zero vs checkpoint
- no snapshot-then-listen race
- throttle to protect source store
- lag that never closes is a bad sign
basics
~20 sIt reads through the entire history of past events first, from the very beginning, building up the data step by step, and once it reaches the present, it switches to just listening for new events as they happen - like binge-watching a show to get caught up, then watching new episodes live.
solid answer
~40 sA catch-up subscription has two phases sharing one code path: a historical replay phase, where the subscriber reads events from position zero (or a checkpoint) in strict order at high throughput, applying each to the projection's store; and a live phase, where once it reaches the tail of the log it switches to receiving new events as they're appended, typically via the same read API so no separate 'live' code path is needed. The transition point matters: because new events can be appended while the historical read is still catching up, the subscription must keep consuming until it truly reaches the current tail, not just 'the position it captured at start time,' or it will miss events written during the catch-up window.
go deeper
Should convey the basic two-step idea: read old stuff first, then switch to listening for new stuff, in order.
Should explain the mechanics precisely - checkpoints, ordered replay, the transition to live - and why a naive snapshot-then-listen approach loses events.
Should discuss throughput and throttling trade-offs, snapshotting to shorten replay, and the operational impact of large replays on a shared event store.
Should reason about this at the level of designing a rollout for a brand-new projection against a live, high-volume, years-old production log - scheduling, backpressure, parallelization strategy, and monitoring for 'lag that never closes.'
## What the pattern is for A **catch-up subscription** is the standard pattern event stores (`EventStoreDB`, Kafka consumers, Axon's `TrackingEventProcessor`, and similar) use to let a consumer start reading from any point in a stream's history and transparently continue into live traffic, without needing two separate implementations for 'replay old stuff' and 'listen to new stuff.' ## The two phases, one read loop Mechanically, it works in two phases over one continuous read loop. 1. In the **historical phase**, the subscriber requests events starting at a given position - position zero for a brand-new projection, or a saved checkpoint for one resuming after downtime - and the store streams them back in strict commit order, typically in batches for throughput, while the subscriber applies each event to its projection store and periodically persists its checkpoint. Because this phase can take anywhere from seconds to hours depending on log size, systems throttle or parallelize it (per-partition or per-aggregate-stream concurrency) to avoid starving the source store's I/O for other live readers. 2. Once the subscriber's read position reaches the position that was 'now' at some earlier instant, it does not simply stop; it must keep consuming, because writers kept appending events the whole time it was catching up. 3. Only once it observes it is caught up to the true current tail - no more backlog - does it transition into the **live phase**, where it typically receives push notifications or long-polls for new events appended after that point, applying them with the same handler code used during replay. ## Why the naive alternative loses events The reason this pattern exists is that a naive alternative - snapshot the current tail position first, replay only up to there, then separately start a live listener from that snapshot - has a **race condition**: events appended between taking the snapshot and starting the live listener would be silently skipped. Catch-up subscriptions solve this by never treating 'caught up' as a point-in-time decision made once; they keep consuming continuously through the transition, so the boundary between historical and live is fuzzy by design and no event is lost in the handoff. This is also what makes projections rebuildable at any time: dropping a projection's store and starting a fresh catch-up subscription from position zero is functionally identical to how the projection was built the first time. ## The trade-off The trade-off is largely about time and resource contention versus operational simplicity. - **The upside is huge**: one subscription implementation handles both bootstrap and steady-state operation, and a projection can be added, at any time, to a system that has been running for years, by simply starting a new catch-up subscription - no special-casing, no coordination with the write side. - **The downside** is that catch-up for a large log is not fast or free: replaying millions of events sequentially, applying a handler per event (which may itself do I/O, like an upsert into a database), can take hours, during which the new projection is unavailable to serve reads, and the read pressure on the source event store competes with every other live subscriber pulling from the same store. Systems mitigate this with **snapshotting** (periodic full-state snapshots so replay resumes from a recent snapshot instead of position zero) at the aggregate level, and with **parallel replay** across independent partitions or streams where ordering only needs to be preserved within a stream, not globally. ## Failure modes Failure modes show up in a few recognizable ways. 1. If a subscriber persists its checkpoint **too eagerly** - before the corresponding write to the projection store is durably committed - a crash and restart can silently skip events, since the subscription resumes past events it never actually applied; conversely, checkpointing too late (after the store write, with a delay) risks reprocessing events on restart, which is why idempotent handlers matter here just as they do for ordinary live delivery. 2. Another common issue is **'catch-up that never catches up'**: if the live write rate outpaces replay throughput, the subscriber's position asymptotically approaches but never reaches the tail, which shows up in monitoring as ever-growing lag rather than a one-time bootstrap delay. 3. A third is **resource starvation**: an unthrottled catch-up subscription doing a full historical replay on a shared, multi-tenant event store can degrade latency for every other live consumer reading from that store concurrently, which is why production systems rate-limit or schedule large replays during low-traffic windows. ## A concrete instance A concrete real-world instance: EventStoreDB's client SDKs expose native catch-up subscription APIs that do exactly this - pass a starting position (or 'start') and the client library handles the historical-to-live transition transparently, which is precisely how teams bootstrap a brand-new read model against a production event store that already has years of history without writing custom replay logic.
- Why can't you just snapshot the current log position, replay up to it, then separately attach a live listener from that snapshot?Because events keep being appended while the historical replay is in progress, so any snapshot-then-listen approach has a race window where events written during the replay are neither included in the replay nor caught by the live listener, and get silently dropped. A true catch-up subscription avoids this by never treating 'caught up' as a single point-in-time decision.
- What's the operational risk of starting a brand-new catch-up subscription against a production event store with ten years of history?The full historical replay can take a long time and generates sustained heavy read load on the source store, which can degrade latency for other live consumers sharing that store. Teams typically throttle replay throughput, run it during low-traffic windows, or use periodic aggregate snapshots so replay doesn't have to start from the very first event.
- How does snapshotting interact with catch-up subscriptions?A snapshot is a periodically stored full-state checkpoint for an aggregate or stream, so replay can resume from the most recent snapshot plus only the events after it, instead of from position zero. This shortens catch-up time dramatically for long-lived streams, at the cost of extra storage and snapshot-management complexity.
Like joining a live group chat that has years of history: you first scroll up and read every past message in order to get context, and only once you've reached the very latest message do you switch to watching new messages arrive in real time - and if people keep posting while you're still scrolling, you keep reading until there's truly nothing left unread.
saying these in an interview costs you the question
- describes catch-up as simply 'reading everything once at startup' with no mention of transitioning to live traffic
- assumes replay and live delivery use different, unrelated code paths
- doesn't recognize the race condition in snapshot-then-listen
- thinks catch-up subscriptions are only relevant the very first time a projection is created
- no awareness that replay load can affect other consumers of the same store