skip to content

What's the difference between a 'live' subscription and a 'catch-up' subscription when a projection consumes an event stream, and why would a projection need both?

level: middleimportance: should knowfreq 55%

answer

  1. catch-up = replay from checkpoint to now
  2. live = push new events as they happen
  3. seamless handoff at the tail avoids gaps
  4. buffer live events during catch-up to avoid missing the gap
  5. new/rebuilding projections start in catch-up

basics

~20 s

A live subscription gets new events as they happen, like watching a live feed. A catch-up subscription reads through the older, already-happened events first to get up to speed. A projection needs both so it can start from wherever it left off (or from the beginning) and then smoothly switch to real-time.

solid answer

~40 s

A catch-up subscription reads events sequentially from a given position — often the beginning, or a saved checkpoint — up through whatever has already been written, at whatever throughput the store/consumer can sustain; it's used when starting a new projection, recovering after downtime, or rebuilding. A live subscription instead delivers new events to the consumer as they're appended, typically with lower latency, often via push rather than bulk pull. In practice a projection usually needs to transition from catch-up to live seamlessly: start in catch-up mode from the last checkpoint, and once it reaches the current head of the stream, switch to live mode so it doesn't miss events written in between — with careful handling of that handoff so no event is skipped or double-applied at the boundary.

go deeper

for a junior

Should be able to describe in plain terms that catch-up means reading old events first and live means getting new ones as they happen.

for a middle

Should know when each mode is used (new/rebuilding projection vs. steady-state) and that a real implementation needs to transition between them.

for a senior

Should identify the gap/race risk at the catch-up-to-live handoff and describe a concrete strategy (buffer-then-drain, position-based dedup) to avoid missing or double-processing events.

for a principal

Should evaluate this trade-off across event-store choices (explicit dual-mode APIs vs. Kafka's implicit continuous-offset model) and factor operational concerns like isolating catch-up read load from live production traffic.

## The two modes of subscription An **event stream**, in this context, is an ordered sequence of events (per topic, per aggregate, or per some other grouping) that a projection reads and processes. A **subscription** is the mechanism a projection handler uses to receive events from that stream. Two distinct modes of subscription show up across event-driven systems: - **catch-up**, meant for reading a backlog of already-existing events - **live**, meant for receiving newly-appended events as they happen ## How each one behaves Mechanically, the split looks like this: | Mode | Behaviour | |---|---| | **Catch-up** | starts at a specific position — the very beginning of the stream, or a checkpoint the projection previously saved — and reads sequentially forward through everything that already exists, typically in batches, optimized for throughput rather than latency, until it reaches whatever was the 'current head' of the stream when it started | | **Live** | by contrast, is optimized for low latency: as soon as a new event is appended, it's pushed (or nearly-immediately polled) to the subscriber, so the consumer learns about it within milliseconds rather than waiting for a bulk-read cycle | Concretely, systems like EventStoreDB expose a live subscription API that pushes new writes to connected subscribers, and change-data-capture systems like MongoDB Change Streams push document changes as they commit. ## Why one projection needs both Both modes exist because a projection has two very different jobs at different points in its life. 1. A brand-new projection, or one being rebuilt after a bug fix, has to process potentially millions of historical events before it's useful at all — doing that through a low-latency, one-at-a-time push path would be needlessly slow and wasteful of connection/notification overhead, so a bulk, throughput-optimized catch-up read is the right tool. 2. Once a projection has processed everything that already existed, though, continuing to run bulk reads on a polling interval to look for new data is wasteful and adds unnecessary latency compared to being pushed new events as they arrive — so mature event-store client libraries typically expose a single subscription abstraction that internally starts in catch-up mode and transitions to live mode automatically once it reaches the tail of the stream. ## The trade-off The trade-off is exactly that split of concerns: - **Catch-up mode favors throughput** at the cost of freshness, which is perfectly fine because during a rebuild or initial load, staleness relative to 'right now' is already expected and accepted. - **Live mode favors freshness** at the cost of added coordination complexity, particularly around the moment of transition between the two. Some simpler systems skip this elegance entirely and just poll on a fixed short interval for everything, trading a bit of latency for a lot less implementation complexity — this is a legitimate and common simplification, especially with Kafka, where a single continuous consumer poll loop naturally reads old records and then new ones without the application needing to think of them as separate modes at all. ## Where it breaks 1. **The handoff itself** is the classic failure mode. If a projection only starts its live listener after fully finishing catch-up, any events appended to the stream during the time catch-up was still running — after it captured its notion of 'current head' but before the live listener actually started — can be missed entirely. The safe pattern is to start buffering live events before or during catch-up, finish catch-up, then drain the buffer while deduplicating by event position so nothing appended during the overlap is skipped or double-applied. 2. **A second, related failure** is catch-up resuming from a stale or incorrect checkpoint after a restart — if the last-processed position wasn't durably persisted before a crash, catch-up either reprocesses a chunk of already-applied events (safe only if the handler is idempotent) or, worse, skips ahead past events that were never actually applied. 3. **A third failure** is running a large catch-up read against the same infrastructure and connection budget as live production traffic, which can starve production reads or writes of I/O capacity while the rebuild churns through history. ## How real systems expose it - **EventStoreDB's catch-up subscription API** is a direct, named example of this pattern: it reads from a given position through to the live end of the stream and then keeps delivering new events without the caller having to manage the transition explicitly. - **MongoDB Change Streams** offer an analogous idea via resume tokens: a consumer can catch up on everything since a given token and then continue receiving live changes. - **Kafka**, notably, doesn't expose these as two separate APIs at all — a consumer group simply starts reading from a stored offset (or 'earliest') and naturally proceeds toward the tail as one continuous poll loop, which is exactly why Kafka-based projections tend to be simpler to reason about on this specific point than systems with an explicit dual-mode subscription API.

  • What's the risk if a projection switches from catch-up to live only after fully finishing the catch-up read?
    Any events appended to the stream during the time the catch-up read was running but after it captured its snapshot of 'current head' can be missed entirely, because the live subscription only starts listening after catch-up is declared done. The fix is to start buffering live events before or during catch-up, then reconcile/dedup by position once catch-up reaches the same point.
  • Why might a Kafka-based projection not need an explicit 'catch-up subscription' API the way an event-store-based one does?
    A Kafka consumer group just reads from a stored offset (or 'earliest') forward as one continuous poll loop — there's no separate 'bulk historical read' mode versus 'live push' mode at the API level, so the transition from replaying old records to consuming new ones is implicit rather than something the application code has to orchestrate.
  • If catch-up is reading from a checkpoint that turns out to be stale by a day due to a crash before the checkpoint was persisted, what's the consequence?
    The projection reprocesses roughly a day's worth of already-applied events, which is safe only if the handler is idempotent; without idempotency this reprocessing would double-apply changes such as double counts or duplicate rows across that whole window.

It's like joining a video call late: you first skim the meeting recording from where you left off to get context (catch-up), then once you're caught up you switch to watching the live feed so you don't miss what's said next — and a good system makes sure it starts recording the live feed before you finish skimming, so nothing said in between falls through the cracks.

saying these in an interview costs you the question

  • treats catch-up and live as unrelated features rather than two phases of the same subscription lifecycle
  • doesn't mention the handoff/gap risk between the two modes
  • assumes a fresh projection can just subscribe live and be correct
  • can't say why bulk historical reads and low-latency pushes have different performance profiles

context