A customer-profile service publishes event-carried state transfer messages ('CustomerUpdated') that include the customer's full current profile in each payload, delivered over a message broker that does not guarantee ordering across partitions. A downstream analytics service applies each event's payload directly onto its local copy as soon as it arrives. What can go wrong, and what technique fixes it?
answer
- broker ordering only within partition
- no version check = last-received wins, not last-happened wins
- monotonic version/sequence number guard
- notification style is order-agnostic
- CDC uses LSN as version
basics
~20 sIf an older update arrives after a newer one, the analytics copy can be overwritten with stale data. Adding a version number or timestamp to each event and only applying it if it's newer fixes this.
solid answer
~50 sBecause the broker doesn't guarantee ordering, a CustomerUpdated event carrying an earlier state can be delivered after one carrying a later state, and a consumer that blindly applies whatever arrives last will end up holding stale data that looks fine but is wrong - a classic out-of-order overwrite. The standard fix is to attach a monotonic version number or a producer-side timestamp to every event and have the consumer compare it against the version it already has, discarding any event whose version is not newer. Some systems go further and version each field or use vector clocks for concurrent writes, but a simple per-entity sequence number is usually enough when there's a single writer. This is a variant of the general 'last-write-wins with a real clock' problem, and it's specific to event-carried state transfer because event notification never has this issue - the consumer always re-fetches current state directly.
go deeper
Should recognize that 'later event overwrites earlier data' is a bug and have a rough intuition that some kind of ordering or timestamp check is needed.
Should name the mechanism precisely (broker doesn't guarantee cross-partition ordering) and propose a version-number guard as the fix.
Should distinguish single-writer version counters from multi-writer conflict scenarios and know when a simple counter isn't sufficient.
Should connect this to broader patterns like CDC/LSN-based versioning or CRDTs, and design detection (reconciliation, snapshot events) as a defense-in-depth measure, not just prevention.
## Ordering guarantees in most message brokers Ordering guarantees are one of the most consequential and most frequently overlooked details in event-carried state transfer designs, and the failure mode described in this scenario - an older payload silently overwriting a newer one - is common enough to have a name in distributed systems circles: the **out-of-order** (or 'stale write') **overwrite problem**. To understand why it happens, start with the mechanics of most message brokers. - **Kafka**, for example, guarantees ordering only within a single partition; if a producer publishes two updates for the same customer to different partitions (which can happen if the partition key isn't consistently the customer ID, or if a producer instance restarts and a rebalance occurs), a consumer can receive them out of publish order. - Even within a single partition, **retries** after a transient publish failure can reorder messages relative to a strict wall-clock sequence. - **SQS standard queues** make no ordering promise at all. So 'the events will arrive in the order they happened' is an assumption that holds only under specific, easy-to-violate conditions, and event-carried state transfer designs are especially exposed because each event is treated as a full replacement for the consumer's local copy. ## The mechanism of the bug The mechanism of the bug is simple once you see it: the consumer's apply logic is effectively `localCopy = event.payload`, with no check on which event is 'more recent' in domain terms. If event V2 (say, an address correction made at 10:00:05) is processed first, and event V1 (the original profile as of 10:00:00) arrives a moment later due to network jitter or a retry, the consumer now believes the pre-correction address is current, and nothing in the system flags this as wrong - the write succeeded, the payload was well-formed, and the local copy simply holds outdated truth. This is why it's called a **silent failure mode**: no exception is thrown, no retry is triggered, and the bug is typically discovered only when a customer complains or a reconciliation job runs. ## The fix: compare recency, not arrival The fix follows directly from naming what's missing: a way to compare 'how recent' two events are that doesn't depend on delivery order. 1. **Attach a version at the producer.** The standard technique is to attach a monotonically increasing version number (or sequence number) to the entity at the producer, incrementing it on every state change, and embedding that version in the event payload. 2. **Guard the apply.** The consumer then applies a simple guard: before overwriting its local copy, check `if (incomingEvent.version > localCopy.version) applyUpdate() else discard()`. This makes the consumer's apply logic commutative and idempotent with respect to delivery order - it converges to the correct state regardless of what order the events actually arrive in, as long as it eventually sees the highest version. 3. **Sequence number over wall-clock timestamp.** A wall-clock timestamp can substitute for a version number if clocks are reliably synchronized (e.g., via NTP with tight bounds), but a producer-side sequence number is more robust because it sidesteps clock-skew risk entirely. 4. **Multiple concurrent writers.** For systems with multiple concurrent writers to the same entity (not a single producer), a simple counter isn't enough, and teams reach for vector clocks or CRDTs to detect and merge concurrent writes rather than just picking a winner by recency. ## Why event notification escapes this This trade-off exists specifically because event-carried state transfer treats each event as a snapshot rather than a query result. Event notification never has this problem in the same form: because the consumer calls back to the producer to fetch current state at the moment it needs it, it always gets whatever the producer's database currently holds, and ordering of the notification messages themselves is largely irrelevant to correctness (multiple 'something changed' notifications for the same entity just trigger redundant, idempotent re-fetches, all converging on the same current answer). This is one of the underappreciated ways event notification is architecturally simpler even though it costs a runtime dependency. ## Where it shows up A concrete instance of this is how many CDC (change-data-capture) pipelines, such as Debezium streaming database row changes into Kafka, rely on the source database's log sequence number (LSN) as exactly this kind of version marker, so downstream consumers building materialized views can detect and discard out-of-order or duplicate change events. Teams that skip this and naively apply CDC payloads as they arrive routinely discover data-quality bugs in their analytics warehouse weeks later, tracing back to a rebalance or retry that reordered a handful of update events for a small number of records.
- Why doesn't event notification suffer from this same out-of-order overwrite problem?Because the consumer never applies a stale payload directly - it just re-fetches current state from the producer whenever it acts, so redundant or reordered notifications simply trigger redundant, idempotent re-reads that all converge on the same current truth. Ordering of the 'something changed' pings doesn't matter to correctness.
- What if two different services can both write updates to the same customer profile concurrently - does a simple version number still work?Not reliably, because a single incrementing counter assumes a single writer picking the next version; with concurrent writers you can get real conflicts where two updates both claim to be 'next'. Systems in that situation typically reach for vector clocks, CRDTs, or route all writes through one owning service to keep a single source of versioning truth.
- How would you detect that this bug has already happened in a running system?A periodic reconciliation job that compares each consumer's local copy against the producer's authoritative current state (or its version number) will surface any records that have silently diverged. Emitting occasional full-snapshot events that consumers can use to self-correct is a common complementary technique.
It's like getting two postcards from a trip with no dates on them - if the one mailed first arrives second, you might think the trip ended somewhere it didn't, unless each postcard is numbered so you know which to trust.
saying these in an interview costs you the question
- Assumes message brokers always deliver events in the order they were published
- Doesn't propose any versioning/sequencing fix, just says 'use a reliable queue'
- Confuses this with the general at-least-once-delivery/duplicate problem instead of ordering
- Claims event notification has the same out-of-order risk
- Suggests wall-clock timestamps solve it without acknowledging clock-skew risk