An order document on the store keeps losing recently written fields under load, and nothing errors: how do you confirm overwritten changes?
answer
- no error exists to look for
- present and complete, but older
- rule out deadline and stale reads
- log the value read beside the value written
- reproduce with concurrent writers
basics
~20 sYou will not find it in the store's errors, because none are raised. Confirm it from outside: compare the entry against an independent record, log the value each caller read beside the value it wrote, and reproduce the loss with concurrent writers.
solid answer
~40 sStart by accepting that the store has nothing to tell you: an unprotected read-modify-write ends in `last-writer-wins` and both writes are acknowledged, so there is no conflict, no rejection and no log line anywhere. That rules the store's error metrics out as evidence and means the investigation happens on the caller's side. Three moves, in order. Rule out the cheaper explanations first — an entry that reached its deadline, a read served by a lagging replica, a write path that was never called. Then make the loss observable: log, on every write, the value the caller read and the value it wrote, and count writes whose starting value had already been superseded. Finally, reproduce it deliberately with concurrent writers and compare the final value with the expected one. Only then choose a remedy.
go deeper
Recall the signature: the document is still there and still consistent, but one writer's recent fields have gone back to older values, and nothing anywhere reported an error.
Explain why the store cannot help: both writes succeeded, so no conflict exists to count. The evidence has to be built on the caller's side by recording what was read alongside what was written.
Demonstrate the ordered investigation: rule out expiry, a stale read and a skipped write first, then instrument the read-write dependency, then reproduce with concurrent writers before changing the write path.
Ask why this was only noticed downstream. A tier that cannot report a lost change needs a deliberate detection strategy for the invariants kept on it, or those invariants belong somewhere that can refuse a write.
## Why the store cannot tell you This defect is defined by its silence. Two callers read the same document, each writes back a complete document, and the later write replaces the earlier one. Both operations did exactly what they were asked, so both were acknowledged; no version was compared, nothing was refused, and the store's error counters stay at zero. Searching the store's logs is the first thing most engineers do and the one thing guaranteed not to help. Say that out loud in an interview. A candidate who asks for the conflict metric has not internalised that there is nothing to conflict with. ## Rule out the cheaper explanations first Several unrelated failures also look like "fields we wrote are not there", and each is cheaper to eliminate than a concurrency investigation: 1. **The entry reached its deadline.** If the document carries a lifetime, a missing document is expiry, not a lost update. Check whether the whole entry vanished or only some fields regressed — regression to an older value is the signature here, absence is not. 2. **A stale read.** Where the tier is replicated and reads may be served by a replica, a read taken shortly after a write can legitimately return the older value without anything having been lost. Confirm whether reads and writes are going to the same node. 3. **The write never happened.** A branch that silently skipped the write, a serialisation that dropped an unknown field, a caller that wrote a subset of the document. Confirm from the caller's own logs that the write was issued with the field present. 4. **Eviction under memory pressure.** As with expiry, this removes the entry rather than reverting a field. The distinguishing signature of an overwritten change is that the document is **present, complete and internally consistent — but at an older state for one writer's fields**. ## Making the loss visible Since nothing in the store records the dependency between a read and a write, the caller has to record it: - On every read-modify-write, log an identifier for the caller, the value (or a hash of it) that was **read**, and the value that was **written**. - A lost update then has a fingerprint: two writes whose recorded prior values are identical, with different results, close together in time. The second one overwrote the first. - Counting those pairs gives you a rate, which is what turns a suspicion into a number you can act on and later verify against. - This is pure instrumentation — it changes nothing about the outcome, which is what you want while you are still establishing the cause. There is a stronger variant: switch one write path to present a version token, and count the refusals. That both measures the rate and prevents the loss, so use it when you are confident of the diagnosis and are ready to fix it; use the logging variant when you are still proving it. ## Reproducing it deliberately A test with one writer will always pass, which is why this reaches production in the first place. Reproduce it with the shape it really has: - Several concurrent writers, each doing read, modify, write against one entry, with a deliberate delay between the read and the write to widen the window. - Assert on the **final state against the expected state**, not on whether each write succeeded — every write will succeed, which is the point. - Run it against the same store configuration as production, since whether the server can interpret the value decides which remedies you can even test. ## From measurement to remedy Once the rate is established, the fix is the ordinary selection between the three remedies available on a store with no transaction to open: move the change onto the server where the server can interpret the value, declare what you read so a stale write is refused instead of landing, or hold an expiring claim around the cycle. Which one fits depends on the value's shape, the contention rate and how long the protected work runs. One thing that is not a fix: adding a retry. A retry needs something to have failed, and nothing failed — an unprotected write always succeeds. Retrying an operation that already reported success changes nothing at all. ## What an interviewer is listening for - The store is ruled out as a source of evidence, explicitly and early. - Cheaper explanations are eliminated before concurrency is blamed, and the signature that distinguishes them is named. - The instrumentation described records the **dependency** between the read and the write, which is the thing nobody else is recording. - The remedy is chosen after the measurement, not proposed as the first move.
- What is the cheapest change that turns this silent loss into a signal?Log, on each write, the prior value the caller read alongside the value it is writing, and count writes whose prior value had already been replaced. That measures the rate without changing behaviour. If you are ready to fix it at the same time, present a version token instead and count the refusals — the refusal rate is the loss rate, and the loss stops.
- Why does wrapping the write in a retry not help here?Because nothing failed. The losing write succeeded and was acknowledged; there is no error to catch and no failed call to repeat. Retrying only makes sense once the write can be refused — which is what presenting a version token, or declaring the entry you read, actually buys you.
- The team says the document is only written by one service, so this cannot be a lost update. How do you check?One service is not one writer. Count concurrent instances, background jobs, retried requests and any administrative tool that touches the same entry, then check whether two write cycles from that service ever overlap in the logs. The race needs two overlapping cycles, not two codebases.
saying these in an interview costs you the question
- Searches the store's error log for a conflict never raised
- Concludes the store dropped a write it acknowledged
- Blames replication lag where reads and writes hit one node
- Assumes an entry lifetime removed the fields without checking
- Tests the write path with one writer and calls it correct
- Adds a retry around a write that already succeeded