When a projection subscribes to read events from an event store to build a read model, what ordering guarantees can it actually rely on, and where do they break down?
answer
- per-stream: strict, gap-free order
- cross-stream: no defined order unless a global feed exists
- global position feed = system-wide commit order
- catch-up subscription = checkpoint + resume
- checkpoint must be transactional with the read-model write
basics
~20 sWithin one entity's own history, events always arrive in the exact order they happened. But if you're watching many entities at once, events from different ones can arrive interleaved in almost any order relative to each other, so you can't assume a global timeline unless the store gives you one explicitly.
solid answer
~50 sAn event store guarantees strict, monotonic ordering within a single stream — reading order-91a3 from position 0 always returns its events in exactly the order they were committed, and a subscription to that stream delivers them in that same order, gap-free. It does not, by default, guarantee any particular relative ordering between two different streams' events; order-91a3's third event and order-77b1's first event have no defined 'which came first' unless the store also exposes a global commit-order feed. Projections that need a consistent cross-aggregate view subscribe to that global feed and process it in commit order, using a persisted checkpoint/position so a restart resumes exactly where it left off rather than reprocessing or skipping events — this checkpointed resubscription is usually called a catch-up subscription. Where this breaks down: consumer lag under load, checkpoint loss on crash, and events published to external systems losing the store's total order if they're re-partitioned differently downstream.
go deeper
Should know that one entity's own events always come back in the order they happened.
Should know that different streams don't have a guaranteed relative order unless there's an explicit global feed, and roughly what a subscription is.
Should explain catch-up subscriptions with checkpointing in detail and identify the checkpoint/read-model transactionality failure mode.
Should reason about consumer lag, retention limits, and ordering loss across system boundaries (e.g., relay into Kafka) at an architecture level.
## The guarantee you get: per-stream order An event store's foundational ordering guarantee is per-stream: within a single named stream, such as `order-91a3`, every event has a strictly increasing position (0, 1, 2, ...) assigned at commit time, and reading — whether a one-shot read of the whole stream or a live subscription to new events on it — always delivers events in exactly that position order, with no gaps and no reordering. This guarantee follows directly from how appends work: the store serializes writes to a given stream (this is the same mechanism optimistic concurrency control relies on — one writer commits at a time per stream, advancing the version by one), so 'the order events were committed in' and 'the order they'll be read back in' are the same thing by construction. ## The guarantee you do not get: cross-stream order What the store does not guarantee, by default, is any defined relative ordering between two different streams. If `order-91a3` gets its third event at 10:00:01.200 and `order-77b1` gets its first event at 10:00:01.100 — a few milliseconds earlier in wall-clock time — there's no built-in mechanism telling a consumer 'process order-77b1's event before order-91a3's third event,' because they're on different streams and the store's core guarantee is scoped to a stream, not to the whole system. This is intentional, not a limitation to work around: streams are independent because the aggregates they represent are independent, and forcing a total order across all of them would mean serializing all writes system-wide, killing the whole point of per-aggregate partitioning (unbounded write concurrency across unrelated entities). ## The second ordering: a global commit-order feed For consumers that do need a cross-aggregate view — most read-model projections do, since a read model like 'orders dashboard' needs events from every order stream — event stores expose a second ordering: a global, monotonically increasing commit-order feed across all streams. EventStoreDB calls this the `$all` stream: every event ever committed to the store, in the literal order it was committed, regardless of which per-aggregate stream it belongs to, addressable by a global position. A projection subscribes to `$all` (or the store's equivalent), processes events as they arrive in that global commit order, and — critically — persists a **checkpoint** after processing each event or batch, recording the global position it has successfully processed up to. If the projection process crashes or is redeployed, on restart it reads its last saved checkpoint and resubscribes from that exact position, rather than starting over from the beginning or picking up wherever 'now' happens to be. This checkpointed resume pattern is what's usually called a **catch-up subscription**: - 'catch up' on everything since the checkpoint (a bounded historical read), - then transition seamlessly into 'live' mode consuming new events as they commit, - with the consumer's code path treating both phases identically. ## Where the guarantees break down Where these guarantees break down in practice falls into a few recurring categories. 1. **First, checkpoint loss**: if the projection commits its checkpoint asynchronously or in a separate transaction from updating the read model itself, a crash between the two can leave the read model updated but the checkpoint stale (causing a harmless re-processing of the same events, tolerable if the projection logic is idempotent) or the checkpoint advanced but the read-model update lost (causing silent data loss in the projection, which is much worse and usually requires the checkpoint-then-update to be transactional or the projection logic to be genuinely idempotent so re-processing recovers it). 2. **Second, consumer lag**: under heavy write load, a slow projection can fall arbitrarily far behind the global feed's head; if downstream retention isn't unbounded, a lagging consumer can lose events it never got a chance to process. 3. **Third, cross-system re-partitioning**: if events are relayed out of the event store into something like Kafka, and the Kafka topic is partitioned by a key other than the aggregate ID (or the store's global order isn't preserved in how messages are produced), the store's careful per-stream and global-commit ordering can be scrambled on the way out, so a downstream Kafka consumer sees an order that matches neither guarantee the original store provided. ## Where it shows up A concrete real-world shape: a MonthlyStatement projection in a banking system subscribes to EventStoreDB's `$all` stream starting from its last saved checkpoint, folds every Deposited/Withdrawn/FeeCharged event across all account streams into per-account running totals in a Postgres read table, and commits the new checkpoint in the same database transaction as the read-table update — so a crash mid-batch either re-applies a handful of already-applied (idempotent) events on restart, or loses nothing, but never silently skips events past a stale checkpoint.
- If a projection's checkpoint update and its read-model update aren't in the same transaction, what specific failure can occur?A crash between the two can leave the checkpoint advanced but the read-model write lost, which silently skips that event forever on restart since the projection believes it already processed up to that position. The safe pattern is either a shared transaction for both, or making the projection logic idempotent so re-processing from a slightly-behind checkpoint is harmless.
- Why doesn't an event store just give every event in the whole system a single global order by default, instead of only per-stream order?Enforcing one total order across all streams would require serializing every write system-wide, defeating the entire purpose of partitioning by aggregate — unrelated aggregates could no longer be written concurrently. Per-stream ordering plus an optional separate global-commit feed gives you both properties without forcing that trade-off on every write.
- What happens to an event store's ordering guarantees if events are relayed into a downstream system like Kafka on a different partition key than the aggregate ID?The relative order the original store guaranteed — per-stream strict order, or system-wide commit order via a global feed — is not automatically preserved by the relay; if the downstream topic is partitioned differently, events can be reordered relative to both of those guarantees, so consumers of the relayed topic can no longer assume either property unless the relay is specifically built to preserve it.
Like reading letters from one pen pal in the order they were postmarked (always reliable) versus trying to interleave letters from ten different pen pals into 'the one true order everything happened' — you need a shared, global postmark log across all ten mailboxes to do that at all, not just careful reading of any one mailbox.
saying these in an interview costs you the question
- assumes any two events anywhere in the store have a well-defined relative order without a global feed
- doesn't know what a catch-up subscription / checkpoint is
- commits checkpoints outside any transaction/idempotency guard tying them to the actual read-model update
- thinks per-stream ordering implies system-wide total ordering