In an eventually consistent data store, two replicas each accept a write to the same key while a network partition separates them. When the partition heals and the store uses last-write-wins (LWW) to reconcile, what happens to the value, and what has to happen for that to work reliably?
answer
- timestamp wins
- clock skew -> silent loss
- no merge, no notify
- per-column vs per-row granularity
- HLC reduces skew window
basics
~20 sLast-write-wins keeps whichever write has the newer timestamp and throws away the other one. For this to work, every replica needs a timestamp on each write, and those timestamps must be trustworthy enough to compare across machines.
solid answer
~50 sLWW resolves conflicting writes by attaching a timestamp (wall-clock or a hybrid logical clock) to every write and keeping the one with the higher timestamp when replicas reconcile after a partition heals. It's attractive because it needs no application logic and no extra metadata beyond a timestamp and node/replica id as tie-breaker. The catch is that 'reliably' hides a hard requirement: the timestamps have to be comparable across nodes, which means clocks need to be reasonably synchronized (NTP, or ideally hybrid logical clocks). If clocks drift, a write that happened later in real time can lose to one that happened earlier just because its node's clock was ahead. LWW also silently discards the losing write entirely — no merge, no notification, just data loss — which is fine for some fields (e.g., a 'last seen' timestamp) and dangerous for others (e.g., a shopping cart).
go deeper
Should be able to state that LWW keeps the value with the later timestamp and that the other write is lost; doesn't need to know about HLCs or per-column granularity yet.
Should know that comparability requires synchronized clocks and be able to name at least one scenario where a causally-later write loses due to skew.
Should be able to reason about per-field vs per-row granularity, propose HLCs or vector clocks as alternatives depending on the data shape, and know which real systems default to LWW.
Should be able to set organization-wide guidance on when LWW is an acceptable default vs when it requires an explicit risk sign-off, and design monitoring to detect silent loss at scale.
## How LWW picks a winner Last-write-wins (**LWW**) is the simplest conflict-resolution strategy an eventually consistent store can use: when two replicas hold different values for the same key because they each accepted a write while the other was unreachable, the store compares a timestamp attached to each write and keeps the value with the larger timestamp, discarding the other entirely. The timestamp is usually written client-side or coordinator-side at write time — either: - a raw **wall-clock reading**, - or increasingly a **hybrid logical clock (HLC)** that combines a physical clock with a logical counter so that ties and small clock skew can be broken deterministically. Some systems also carry a replica or node id as a final tie-breaker when timestamps are exactly equal, so the outcome is deterministic no matter which replica performs the comparison. Once the partition heals and replicas exchange their versions (via anti-entropy, read-repair, or hinted handoff), each replica runs the same comparison and converges on the same winner — that determinism is what makes LWW usable in a leaderless, multi-master system with no coordinator to arbitrate. ## Why it exists LWW exists because most other conflict-resolution techniques require either: - a **central authority** to serialize writes (which sacrifices availability during partitions, the thing eventual consistency is trying to preserve), - or **application code** that understands how to merge two divergent values (extra engineering work the application team may not want to do for every field). LWW sidesteps both: it needs no consensus round and no domain-specific merge function, just a comparable timestamp already present on the write. That makes it the default conflict policy in systems like **Cassandra** (per-cell timestamps) and **Riak** (as one of several selectable policies), and it's attractive whenever the 'right' answer really is 'whichever write happened most recently,' such as a user's last-known IP address or a cache entry. ## The trade-off The trade-off is that LWW buys simplicity by being willing to lose data. The losing write isn't merged, queued for review, or even logged by default — it simply vanishes, and neither client that made the concurrent writes is told anything happened. This is acceptable for values where only the latest state matters, but it is the wrong tool whenever a write represents an operation rather than a value replacement — two concurrent 'add item to cart' writes are not readings of the same underlying fact, and picking one over the other silently drops a customer's item. ## Failure modes 1. **Clock-skew-induced causality violation** — the dominant failure mode. Because LWW orders writes by timestamp rather than by actual happened-before relationships, a write that is causally later (a client saw the first write's result and then issued a follow-up) can still lose if its node's physical clock lags behind the other node's. A concrete pattern that shows this in production: a user updates their profile, the update briefly appears, then an older cached write from a lagging replica 'wins' during reconciliation and the user's change disappears — without any error being raised, which makes it a support-ticket generator rather than something caught in testing. 2. **Timestamp granularity** — a second failure mode. Many implementations timestamp whole values or whole rows rather than individual fields, so an update to field A on one replica can wholly overwrite a concurrent update to field B on another replica, even though the two changes don't logically conflict at all. ## Cassandra in practice A well-known real-world instance is Cassandra's default conflict resolution: every column (cell) write carries a client-supplied or server-assigned microsecond timestamp, and on read (or during compaction/repair) the cell with the highest timestamp wins per-column. This per-column granularity mitigates the whole-row-clobber problem somewhat, but Cassandra explicitly warns operators to keep NTP-synchronized clocks tight, because two writes issued within the same clock-skew window can resolve in the 'wrong' order relative to what actually happened. Teams that need correctness guarantees stronger than 'usually right, occasionally silently wrong' either: - pair LWW with hybrid logical clocks to shrink the skew window, - or move the field to a strategy that preserves both writes (vector clocks plus application merge) instead of discarding one.
- What is a hybrid logical clock and why do systems prefer it over raw wall-clock timestamps for LWW?A hybrid logical clock (HLC) combines a physical clock reading with a logical counter, so it advances with real time but also strictly increases on every event and can incorporate causality information from messages it receives. This lets a store break ties deterministically and ensures that if event A causally happens before event B, A's HLC timestamp is guaranteed to be smaller, which raw NTP-synced wall clocks cannot promise under skew. It doesn't eliminate LWW's fundamental data-loss behavior, but it shrinks the window in which clock skew can invert the 'true' order of unrelated concurrent writes.
- If LWW is risky, why do widely used systems like Cassandra still ship it as the default?Because for the majority of workloads — caches, session state, telemetry, denormalized read models — only the freshest value matters and losing a stale concurrent write is a non-event, not a bug. It also requires zero extra storage or client complexity compared to vector clocks and siblings, which matters at Cassandra's scale where every byte of metadata per cell adds up. Teams that need stronger guarantees are expected to opt into lightweight transactions or handle merge explicitly rather than pay that cost on every write.
- How would you detect that LWW is silently dropping writes in production?Track write and read metrics per key alongside comparing values you expect against actual stored state, or instrument the reconciliation path to log/count every time a write is discarded due to a timestamp comparison. Some teams sample discarded writes and replay them against an audit log to see how often causally-later writes lose. Without this instrumentation, LWW data loss is invisible because there's no error path, only a quietly different final value.
Like two people editing a shared whiteboard from different rooms with clocks that aren't perfectly synced — whoever's watch says the later time gets to keep their writing, even if the other person actually wrote second in real life.
saying these in an interview costs you the question
- Says LWW never loses data as long as timestamps exist
- Assumes wall-clock timestamps are always causally ordered
- Doesn't know LWW requires clock synchronization to be meaningful
- Applies LWW to counters/collections instead of last-value fields
- Thinks the losing write is logged or recoverable by default