Last-write-wins is popular because it's simple and requires no application involvement. Walk through at least three concrete ways LWW can silently lose committed data or violate causality in production, and describe one situation where you would deliberately choose LWW anyway despite these risks.
answer
- clock skew flips order
- whole-row clobber
- delete vs update race -> zombie writes
- tombstones for safe deletes
- fine for freshest-value-only fields
basics
~20 sLast-write-wins can quietly throw away real changes when clocks disagree, when a whole record gets overwritten instead of just the changed part, or when a 'delete' loses to an older 'update.' It's still fine to use when only the newest value ever matters, like a cache or a heartbeat.
solid answer
~50 sLWW's core risk is that it converts a rich causal question ('which write logically supersedes the other?') into a single scalar comparison ('which timestamp is bigger?'), which breaks in at least three concrete ways: clock skew between nodes can make a causally-later write lose to a causally-earlier one; coarse granularity (whole-row or whole-document LWW) can silently discard an unrelated field change bundled with the losing write; and delete-vs-update races can resurrect deleted data or permanently drop an update, depending on which one's timestamp happens to be later. None of these produce errors — they produce a quietly different final state, which is what makes LWW dangerous for anything the business cares about losing. It's still the right choice for high-volume, low-stakes fields where only the freshest value matters and paying for vector clocks or a merge function isn't worth it — session heartbeats, presence/last-seen fields, ephemeral caches, or metrics where an occasional lost sample doesn't change the aggregate.
go deeper
Should be able to name at least one way LWW can lose data (e.g., clock skew) when prompted.
Should be able to describe two of the three failure modes concretely and know deletes need tombstones, not plain LWW.
Should discuss all three failure modes with mechanism-level detail and articulate a clear rule for when LWW is/isn't acceptable.
Should be able to set a policy (which classes of fields default to LWW vs require merge/CRDTs) and design tombstone retention / clock-sync SLAs that bound the risk organization-wide.
## What the single comparison discards Last-write-wins collapses conflict resolution to a single comparison — pick the write with the larger timestamp — which is attractive because it requires no extra metadata beyond a clock reading and no application involvement, but that same simplicity is exactly what causes its failure modes: it discards the causal information that would let the system tell the difference between 'this write is genuinely newer' and 'this write merely has a larger number attached to it.' ## Three failure modes 1. **Clock-skew-induced causality violation** — the first and most cited failure mode. If node A's physical clock runs ahead of node B's, a write accepted by A can carry a timestamp later than a write accepted by B even when B's write causally happened after A's — for instance, a client reads A's write and then deliberately issues a follow-up write to B a second later in real time, yet A's skewed clock still stamps its (earlier) write with a larger number. LWW then keeps A's stale write and discards B's intentional follow-up, and nothing in the system flags this as wrong: it's a valid-looking write with a valid-looking timestamp, just the wrong one. This is why systems that rely on LWW invest heavily in tight clock synchronization (NTP, physical time devices like Spanner's TrueTime, or hybrid logical clocks that bound the skew window), because the correctness of LWW degrades directly with how much clocks can disagree. 2. **Granularity** — the second failure mode. Many LWW implementations compare and swap an entire value — a whole row, a whole document — rather than individual fields, because tracking timestamps at finer granularity costs more metadata. When two clients concurrently modify different fields of the same record starting from the same base version, whole-value LWW keeps one client's entire document and discards the other's entirely, even though the two changes didn't logically conflict at all — a user changing their bio and, seconds later on a different device, changing their avatar can end up with only one of the two changes surviving, purely because of write ordering unrelated to the actual conflict. Cassandra mitigates this specifically by timestamping at the column (cell) level rather than the row level, which narrows but doesn't eliminate the problem — a delete-then-recreate of a whole row, or a batch of column writes issued together, can still exhibit whole-value-like clobbering. 3. **The delete/update race** — the third failure mode, and often the most operationally painful. Because a delete is just another write competing on the same timestamp axis, a concurrent update and delete of the same key resolve however their timestamps happen to compare, with no special-casing for the fact that 'deleted' is a different kind of state than 'has a value.' If the delete's timestamp is later, the record correctly disappears; if a stale update happens to carry a later timestamp than the delete (which can happen with clock skew, or with delayed message delivery through a queue), the deleted record can silently reappear — a 'zombie' write — well after the user thought it was gone. This is one of the reasons deletes in LWW-based systems are usually implemented as tombstones with their own retention and compaction rules rather than as physical removal, precisely so a late-arriving update can still be compared against and correctly lose to a tombstone within its retention window. ## Why the damage stays invisible None of these three failure modes produce an error the application can catch — they all manifest as a quietly different final state, discovered later (if ever) as a confused support ticket, a re-appearing deleted item, or a change that 'didn't stick.' ## When LWW is still the right choice That said, LWW remains the right engineering choice for a large class of data precisely because these failure modes are tolerable there: - **session heartbeats and presence/'last seen' fields** only ever care about the freshest value, so losing a stale concurrent write is definitionally correct behavior, not a bug; - **high-volume metrics or cache entries** can tolerate an occasional lost sample without changing any aggregate a human relies on; - **the operational simplicity** — no per-key metadata growth, no sibling management, no merge functions to write and test — is a real, ongoing cost saving at scale. The engineering judgment call is asking, for each field, 'if this write silently disappears with no error and no notification, does anyone's outcome change?' — if the answer is no, LWW is the pragmatic default; if the answer is yes (money, inventory, user-authored content, anything representing an operation rather than a value snapshot), it's the wrong tool and the cost of vector clocks plus a merge function, or a CRDT, is money well spent.
- How do tombstones fix the delete/update race, and what new problem do they introduce?A tombstone is a marker write recording 'this key was deleted at time T' rather than physically removing the data, so a late-arriving update with an earlier timestamp than the tombstone correctly loses to it during comparison instead of silently resurrecting the record. The new problem is that tombstones themselves must be retained for some window before being physically purged (compacted away), and if that window is too short, a very late update can still resurrect data after the tombstone is gone; if it's too long, tombstones accumulate and bloat storage and read performance, which is exactly the operational trade-off systems like Cassandra tune via a grace-period setting.
- Would using a hybrid logical clock instead of a raw wall clock eliminate the clock-skew failure mode?It significantly narrows the window in which skew can invert causal order, because an HLC advances based on message timestamps it observes and is guaranteed to strictly increase across any causally related pair of events, not just events on the same node. It doesn't eliminate the risk entirely for truly concurrent, non-communicating writers, since HLCs still have a physical-time component that can drift between nodes that never exchange a message establishing the relationship.
- If a team decides LWW is unacceptable for a field, what's the cheapest next step up before jumping to full vector clocks and application merge?Field-level (rather than whole-row) timestamps, if the store supports them, remove the granularity failure mode cheaply without adding a new mechanism. For the delete/update race specifically, adopting tombstones with a sane retention window addresses that one failure mode directly. Only if the data genuinely needs to preserve two independent concurrent writes (not just 'apply the more recent field-level change correctly') is the jump to vector clocks plus a merge function actually necessary.
It's like a shared inbox where only the message with the latest postmark is kept and all others are shredded unread — fine if the inbox only ever needs 'the newest memo,' disastrous if two people mailed genuinely different, both-important updates on the same day.
saying these in an interview costs you the question
- Claims LWW is always wrong / never appropriate
- Doesn't know deletes need special handling (tombstones) under LWW
- Thinks tighter clock sync fully eliminates LWW's risk rather than narrowing it
- Can't name a scenario where LWW is the correct pragmatic choice
- Assumes LWW failures would show up as visible errors