skip to content

A nightly job overwrote a day of raw events in a store with many-nines durability - why did that durability figure not help?

level: middleimportance: must knowfreq 58%

answer

  1. hardware, not you
  2. the write was successful
  3. replication copies intent faithfully
  4. a replica is not an earlier moment
  5. retained and out of the writer's reach

basics

~20 s

Durability defends the bytes against hardware, not against you. The store did exactly what it was asked and then replicated the overwrite faithfully to every copy, so a high durability figure is not a backup and never rewinds a bad write.

solid answer

~50 s

Durability answers one question: will the bytes the store acknowledged still be reconstructable later. It defends against disks, machines and whole failure domains, and it defends them by copying whatever the store currently holds. A bad write is not a failure from the store's point of view - it is a successful request, so every copy converges on the new bytes, the near ones immediately and the far one after its replication lag. Nothing about a many-nines figure implies an older state exists anywhere. What protects you here is a genuinely different thing: a copy that is retained for a period and that the principal doing the writing cannot change, so the earlier state is still available after the mistake. That is a backup, and the honest answer in an interview is to name the distinction rather than to reach for more replication.

go deeper

for a junior

Remember the direction: durability protects bytes from hardware failures, not from a job that writes the wrong thing. A successful bad write is not a store failure.

for a middle

Explain why replication makes it worse rather than better: every copy converges on the new bytes, and a deletion is copied just as faithfully. Then name what a backup adds.

for a senior

Show the recovery reasoning: can the source replay the window, what earlier state is retained, and which identity is allowed to change it. Mention that an untested restore is an assumption.

for a principal

Set the standard: which datasets are systems of record, what retention each gets, and who may alter the retained copy. On a large raw zone, retaining everything forever is a budget decision, not a safety one.

## What durability actually defends A durability figure is a statement about **loss**: the probability that an object the store acknowledged can no longer be reconstructed from any copy it holds. The threats it is built against are physical and statistical - a disk that fails, a machine that dies, media that decays silently, and at wider scopes a whole zone or region going away. The defence is redundancy plus repair: keep copies, notice a missing one, rebuild it. Notice what is missing from that list. Nothing in it is about **intent**. The store has no opinion about whether a write is what you meant. ## Replication is faithful, and that is exactly the problem When the nightly job wrote over the day of raw events, the store received a well-formed, authorised request and carried it out. From there the machinery works against you with perfect reliability: - The near copies take the new bytes before the write is acknowledged. - A copy in another region takes them too, after its replication lag, so the far copy converges on the damage rather than preserving the earlier state. - The repair process keeps every copy consistent with the current object, which is now the wrong one. The same is true of deletion: replication copies a delete faithfully. A replica is a defence against a failure domain. A backup is a defence against a mistake. They are different products of different mechanisms, and calling replication a backup is the single most expensive misunderstanding in this material. ## The threat table | Threat | What defends against it | | --- | --- | | A failed disk or machine | Copies inside one zone, plus repair | | The loss of a whole zone | Copies spread across zones in the region | | The loss of a region | A copy in a second region | | A bad write or an accidental delete by your own job | A retained copy the writer cannot modify | | Credentials of the writing workload being misused | A retained copy under different control from the writer | The top three rows are what a durability figure summarises. The bottom two are not, and no number of nines moves them. ## What makes a stored copy a backup Four properties, and a copy that lacks any of them is a replica wearing the word: 1. **It represents an earlier moment.** It is a point you can go back to, not a mirror of the current state. A mirror of the current state contains the mistake. 2. **It is retained.** It survives for a defined period regardless of what happens to the original, which is what gives you time to notice. 3. **The writer cannot rewrite it.** If the same identity that corrupted the objects can also change or remove the copy, one bad job or one stolen credential takes both. 4. **It has been restored from at least once.** An untested restore is an assumption, not a defence. ## The landing-zone case For a raw-event landing zone the first question is actually upstream: **can the events be replayed?** If the ingest source retains the day, the cheapest recovery is to re-run ingest for that window, and the store is not the system of record for that period at all. That answer is often better than any storage feature, and an interviewer will respect it more than a reflex about more copies. If the events cannot be replayed, then the landing zone is the system of record, and it needs the properties above - a retained earlier state, out of reach of the ingest identity. The cost argument is also honest: raw zones are large and mostly cold, so a retained history of everything is expensive, and the usual shape is a shorter retention on the raw zone plus a longer one on the curated output that downstream consumers actually depend on. ## How to say it in an interview Lead with the direction of the mechanism: durability defends bytes against hardware, replication copies intent faithfully, and neither one is a backup. Then name the property you actually need - an earlier state, retained, that the writing principal cannot touch - and finish by asking whether the source can simply be replayed, which is the answer that costs nothing.

  • Would a copy in a second region have saved the overwritten day?
    No, only delayed it. The far copy is applied after the write is acknowledged, so it converges on the same bad bytes once replication catches up. The lag is a few minutes of accident, not a design you can rely on, and nobody notices a corruption inside that window on purpose.
  • What is the cheapest recovery for a raw landing zone after a bad write?
    Replaying from the source, when the upstream still retains the window. That makes the ingest source the system of record for that period and costs you a re-run rather than a stored history of everything. It only works if the retention upstream is long enough to cover the time it takes to notice.
  • Why does it matter which identity can change the retained copy?
    Because the threat model includes the writer. If the same credential that corrupted or deleted the objects can also modify the retained copy, one bad job or one compromised workload removes both the damage and the recovery. Separating who can write from who can alter the retained copy is what makes the second one a real defence.

A fireproof safe protects the document inside it from the building burning down. It does nothing about someone with the key crossing out a paragraph.

saying these in an interview costs you the question

  • Calls cross-region replication a backup
  • Thinks many nines of durability means data cannot be lost by any cause
  • Plans to recover from the far copy, which replicates the overwrite too
  • Assumes the store rejects a write that damages existing data
  • Believes more copies reduce the risk of a bad write