skip to content

questions

6

In an eventually consistent data store, two replicas each accept a write to the same key while a network partition separates them. When the partition heals and the store uses last-write-wins (LWW) to reconcile, what happens to the value, and what has to happen for that to work reliably?

level: juniorimportance: must knowfreq 75%

answer

  1. timestamp wins
  2. clock skew -> silent loss
  3. no merge, no notify
  4. per-column vs per-row granularity
  5. HLC reduces skew window

basics

~20 s

Last-write-wins keeps whichever write has the newer timestamp and throws away the other one. For this to work, every replica needs a timestamp on each write, and those timestamps must be trustworthy enough to compare across machines.

solid answer

~50 s

LWW resolves conflicting writes by attaching a timestamp (wall-clock or a hybrid logical clock) to every write and keeping the one with the higher timestamp when replicas reconcile after a partition heals. It's attractive because it needs no application logic and no extra metadata beyond a timestamp and node/replica id as tie-breaker. The catch is that 'reliably' hides a hard requirement: the timestamps have to be comparable across nodes, which means clocks need to be reasonably synchronized (NTP, or ideally hybrid logical clocks). If clocks drift, a write that happened later in real time can lose to one that happened earlier just because its node's clock was ahead. LWW also silently discards the losing write entirely — no merge, no notification, just data loss — which is fine for some fields (e.g., a 'last seen' timestamp) and dangerous for others (e.g., a shopping cart).

go deeper

for a junior

Should be able to state that LWW keeps the value with the later timestamp and that the other write is lost; doesn't need to know about HLCs or per-column granularity yet.

for a middle

Should know that comparability requires synchronized clocks and be able to name at least one scenario where a causally-later write loses due to skew.

for a senior

Should be able to reason about per-field vs per-row granularity, propose HLCs or vector clocks as alternatives depending on the data shape, and know which real systems default to LWW.

for a principal

Should be able to set organization-wide guidance on when LWW is an acceptable default vs when it requires an explicit risk sign-off, and design monitoring to detect silent loss at scale.

## How LWW picks a winner Last-write-wins (**LWW**) is the simplest conflict-resolution strategy an eventually consistent store can use: when two replicas hold different values for the same key because they each accepted a write while the other was unreachable, the store compares a timestamp attached to each write and keeps the value with the larger timestamp, discarding the other entirely. The timestamp is usually written client-side or coordinator-side at write time — either: - a raw **wall-clock reading**, - or increasingly a **hybrid logical clock (HLC)** that combines a physical clock with a logical counter so that ties and small clock skew can be broken deterministically. Some systems also carry a replica or node id as a final tie-breaker when timestamps are exactly equal, so the outcome is deterministic no matter which replica performs the comparison. Once the partition heals and replicas exchange their versions (via anti-entropy, read-repair, or hinted handoff), each replica runs the same comparison and converges on the same winner — that determinism is what makes LWW usable in a leaderless, multi-master system with no coordinator to arbitrate. ## Why it exists LWW exists because most other conflict-resolution techniques require either: - a **central authority** to serialize writes (which sacrifices availability during partitions, the thing eventual consistency is trying to preserve), - or **application code** that understands how to merge two divergent values (extra engineering work the application team may not want to do for every field). LWW sidesteps both: it needs no consensus round and no domain-specific merge function, just a comparable timestamp already present on the write. That makes it the default conflict policy in systems like **Cassandra** (per-cell timestamps) and **Riak** (as one of several selectable policies), and it's attractive whenever the 'right' answer really is 'whichever write happened most recently,' such as a user's last-known IP address or a cache entry. ## The trade-off The trade-off is that LWW buys simplicity by being willing to lose data. The losing write isn't merged, queued for review, or even logged by default — it simply vanishes, and neither client that made the concurrent writes is told anything happened. This is acceptable for values where only the latest state matters, but it is the wrong tool whenever a write represents an operation rather than a value replacement — two concurrent 'add item to cart' writes are not readings of the same underlying fact, and picking one over the other silently drops a customer's item. ## Failure modes 1. **Clock-skew-induced causality violation** — the dominant failure mode. Because LWW orders writes by timestamp rather than by actual happened-before relationships, a write that is causally later (a client saw the first write's result and then issued a follow-up) can still lose if its node's physical clock lags behind the other node's. A concrete pattern that shows this in production: a user updates their profile, the update briefly appears, then an older cached write from a lagging replica 'wins' during reconciliation and the user's change disappears — without any error being raised, which makes it a support-ticket generator rather than something caught in testing. 2. **Timestamp granularity** — a second failure mode. Many implementations timestamp whole values or whole rows rather than individual fields, so an update to field A on one replica can wholly overwrite a concurrent update to field B on another replica, even though the two changes don't logically conflict at all. ## Cassandra in practice A well-known real-world instance is Cassandra's default conflict resolution: every column (cell) write carries a client-supplied or server-assigned microsecond timestamp, and on read (or during compaction/repair) the cell with the highest timestamp wins per-column. This per-column granularity mitigates the whole-row-clobber problem somewhat, but Cassandra explicitly warns operators to keep NTP-synchronized clocks tight, because two writes issued within the same clock-skew window can resolve in the 'wrong' order relative to what actually happened. Teams that need correctness guarantees stronger than 'usually right, occasionally silently wrong' either: - pair LWW with hybrid logical clocks to shrink the skew window, - or move the field to a strategy that preserves both writes (vector clocks plus application merge) instead of discarding one.

  • What is a hybrid logical clock and why do systems prefer it over raw wall-clock timestamps for LWW?
    A hybrid logical clock (HLC) combines a physical clock reading with a logical counter, so it advances with real time but also strictly increases on every event and can incorporate causality information from messages it receives. This lets a store break ties deterministically and ensures that if event A causally happens before event B, A's HLC timestamp is guaranteed to be smaller, which raw NTP-synced wall clocks cannot promise under skew. It doesn't eliminate LWW's fundamental data-loss behavior, but it shrinks the window in which clock skew can invert the 'true' order of unrelated concurrent writes.
  • If LWW is risky, why do widely used systems like Cassandra still ship it as the default?
    Because for the majority of workloads — caches, session state, telemetry, denormalized read models — only the freshest value matters and losing a stale concurrent write is a non-event, not a bug. It also requires zero extra storage or client complexity compared to vector clocks and siblings, which matters at Cassandra's scale where every byte of metadata per cell adds up. Teams that need stronger guarantees are expected to opt into lightweight transactions or handle merge explicitly rather than pay that cost on every write.
  • How would you detect that LWW is silently dropping writes in production?
    Track write and read metrics per key alongside comparing values you expect against actual stored state, or instrument the reconciliation path to log/count every time a write is discarded due to a timestamp comparison. Some teams sample discarded writes and replay them against an audit log to see how often causally-later writes lose. Without this instrumentation, LWW data loss is invisible because there's no error path, only a quietly different final value.

Like two people editing a shared whiteboard from different rooms with clocks that aren't perfectly synced — whoever's watch says the later time gets to keep their writing, even if the other person actually wrote second in real life.

saying these in an interview costs you the question

  • Says LWW never loses data as long as timestamps exist
  • Assumes wall-clock timestamps are always causally ordered
  • Doesn't know LWW requires clock synchronization to be meaningful
  • Applies LWW to counters/collections instead of last-value fields
  • Thinks the losing write is logged or recoverable by default

context

open as a page

A distributed key-value store attaches a vector clock to every value it stores, one counter per replica. When two versions of the same key are compared during a read, how does the store use the vector clocks to decide whether one version happened-before the other, or whether the two are concurrent and represent a real conflict?

level: middleimportance: must knowfreq 70%

basics

~20 s

A vector clock is a small list of counters, one per replica, attached to a value. By comparing the lists between two versions, the system can tell if one grew directly out of the other, or if they happened independently and truly conflict.

open as a page

When automatic conflict resolution (like LWW or vector-clock causality checks) cannot safely pick a winner because two writes are truly concurrent and semantically different, systems like Amazon's Dynamo return both versions to the application as 'siblings.' Concretely, how does an application perform a semantic merge on such siblings, using a shopping-cart example, and what does the application have to guarantee for the merge to be safe?

level: seniorimportance: must knowfreq 65%

basics

~20 s

When a database can't safely decide which of two conflicting writes is 'right,' it hands both versions to the application, and the application's own code — which understands what the data means — merges them into one sensible result, like combining two shopping carts instead of picking one and losing items from the other.

open as a page

Last-write-wins is popular because it's simple and requires no application involvement. Walk through at least three concrete ways LWW can silently lose committed data or violate causality in production, and describe one situation where you would deliberately choose LWW anyway despite these risks.

level: seniorimportance: must knowfreq 60%

basics

~20 s

Last-write-wins can quietly throw away real changes when clocks disagree, when a whole record gets overwritten instead of just the changed part, or when a 'delete' loses to an older 'update.' It's still fine to use when only the newest value ever matters, like a cache or a heartbeat.

open as a page

A team building a Dynamo-style store advertises 'vector clocks' for conflict detection, but a reviewer says what they've actually implemented is a version vector, and that the distinction matters for correctness. What is the difference between a vector clock and a version vector in this context, and what unbounded-growth problem do both share in a system with many nodes or clients?

level: middleimportance: should knowfreq 40%

basics

~20 s

A 'version vector' is basically the same idea as a vector clock, but it counts per storage replica instead of per client, which keeps it small and stable no matter how many different users write to the data.

open as a page

In a leaderless replicated store using vector-clock-based conflict detection, deletes are often implemented as tombstone writes rather than physically removing data. Explain why plain deletes are dangerous under concurrent conflict resolution, and what problems tombstones introduce over time that operators must manage.

level: principalimportance: should knowfreq 35%

basics

~20 s

If you just erase a deleted record, a slightly-late update from before the delete can make it reappear, because there's nothing left to compare against. So systems keep a 'this was deleted' marker (a tombstone) around for a while instead of erasing right away — but keeping markers around forever wastes space, so they eventually get cleaned up too.

open as a page