skip to content

Every object is replicated to a second region, yet a mistaken delete removed the documents there too — why?

level: seniorimportance: must knowfreq 58%

answer

  1. a copy obeys the same instruction
  2. failure domains, not mistakes
  3. the destination converges on the source
  4. lag is not a recovery window
  5. previous state is what survives a wrong call

basics

~20 s

Replication copies operations, so a deletion is reproduced faithfully at the destination. It defends against losing a failure domain, never against a wrong call. Versioning, a retention lock or a genuinely separate backup are what survive a mistake.

solid answer

~40 s

Replication is a pipeline that applies what happened at the source to the destination — it keeps the second copy *current*, not *older*. A deletion is an operation like any other, so it propagates, and the destination ends up in the same state as the source. That makes replication a defence against losing a failure domain, a region or a zone, and against latency or residency requirements; it is not a defence against an operator deleting the wrong prefix, because the operator's intent is exactly what gets copied. What survives a mistake is something that keeps the *previous* state: versioning at either end, a retention lock that forbids removal, or a backup held under different credentials. Whether a delete marker replicates by default differs between providers, so it must be verified rather than assumed.

go deeper

for a junior

Recall that a replica is a current copy, not an older one: whatever you do at the source is applied at the destination, deletions included. A second copy is not an undo.

for a middle

Explain the split: replication addresses losing a failure domain, locality and residency, while versioning, retention locks and backups address a wrong instruction. Be able to say why replication lag is not a recovery window.

for a senior

Diagnose the specific path — the delete propagated, a marker propagated, versions were purged, or the destination was never versioned — and say which setting decides each. Volunteer the drill that would have caught it.

for a principal

Set the standard: which copies exist for place failure and which for instruction failure, where the separately-credentialed backup lives, and a recurring drill whose result is a measured recovery time rather than a diagram.

## Replication copies instructions, not moments The mental model that causes this incident is that a replica is a *spare*. It is not. A replication configuration is an ongoing pipeline: an operation lands at the source, and the platform applies an equivalent operation at the destination, usually asynchronously and usually within a short but unguaranteed delay. The destination's job is to **converge on the source**. That single property explains the outcome: - A write propagates, so the destination gets the new object. - An overwrite propagates, so the destination gets the corrupt render too. - A delete propagates, so the destination loses the key as well. Nothing in the design distinguishes an intended delete from a mistaken one, because that distinction exists only in the operator's head. ## What replication is genuinely for It is a real and valuable control — for a different class of risk: - **Failure-domain loss.** A zone or a region becomes unavailable or loses its copies; the other copy is intact because it was never in that domain. - **Locality.** Readers far from the source get a copy nearer to them. - **Residency.** A copy is deliberately placed in a jurisdiction that a rule requires. So the correct sentence is that replication protects against **the place** failing, and versioning, locks and backups protect against **the instruction** being wrong. Mixing the two up is the most common storage-safety error in the field. | Threat | Replication | Versioning | Retention lock | Separate backup | |---|---|---|---|---| | A region or zone is lost | yes | no | no | yes, slowly | | Someone deletes the wrong prefix | no | yes | yes | yes | | A job overwrites everything with garbage | no | yes | yes | yes | | Credentials are stolen and versions purged | no | no | yes | yes, if separately held | | A corrupt render is propagated | no | yes | partly | yes | ## Why the destination did not keep an older copy There are several ways the second copy ends up empty, and they are worth separating because the fixes differ: 1. **The delete replicated as a delete.** The destination applied it and the key is gone there too. This is the ordinary case. 2. **A delete marker replicated.** On a versioned pair, the marker itself is an operation; if markers are replicated, the destination's key is hidden exactly as the source's is. Providers differ on whether marker replication is on by default, and that setting is easy to assume in either direction — verify it rather than believing it. 3. **Versions were purged at the source.** A version-targeted delete removes bytes. Depending on configuration, the destination may retain its own versions or may not — so the destination sometimes is recoverable, and that is a configuration outcome, not a guarantee. 4. **The destination never had versioning on.** A replication pair does not imply matched settings on both ends. If the destination store is unversioned, it has only ever held current objects. Note what falls out of the list: replication *lag* is not a recovery window. It is not guaranteed to be long, nobody notices a mistake inside it reliably, and racing a pipeline is not a recovery plan. ## What actually survives a wrong delete - **Versioning at either end**, which turns the delete into a marker and keeps the content beneath it. - **A retention lock** on the retained versions, which refuses the removal outright rather than making it reversible. - **A backup held under different credentials**, in a different account or a different trust boundary, with its own retention. This is the one that also survives credentials being stolen, because the point is not distance but *separation of control*. None of these is what a very high advertised durability figure is about, either. Durability describes the bytes surviving hardware and media failures; it makes no promise about an operation you asked for. ## Proving it rather than believing it The reason this scenario keeps happening is that the protection was never tested. The drill is small and it is the thing a senior candidate should volunteer: 1. Write a throwaway object, wait for it to appear at the destination, then delete it at the source. 2. Check the destination: is the key gone, hidden by a marker, or intact? 3. Recover it by the route you believe you have — promoting a version, removing a marker, or restoring from the backup — and time how long that takes for one key. 4. Multiply by the number of keys a realistic incident touches, and see whether the answer is still acceptable. An organisation that has run this once knows which of its copies are *current* and which are *previous*, which is the entire distinction this question is about.

  • The destination store has versioning of its own. Does that save you?
    Sometimes, and only by configuration. If the delete arrives as a marker and the destination retains previous versions, the content is recoverable there. If markers are not replicated but version purges are, or the destination was never versioned, nothing older is held. It is a setting to verify, not a property to assume.
  • If replication does not protect against a mistake, what does it actually buy?
    Survival of a failure domain — a zone or a region becoming unavailable or losing its copies — plus locality for distant readers and placement for a residency rule. Those are real and none of them are addressed by versioning. The two controls are complements, not substitutes.
  • Why does a backup in another account beat a replica in another region here?
    Because the threat is control, not distance. A replica shares the operation stream and often the credentials that issued the bad call; a backup under separate credentials with its own retention does not apply that call at all. Separation of control is what survives both a mistake and a stolen key.

A machine that mirrors your filing cabinet into another building copies the shredding as faithfully as the filing. What saves you is the drawer of superseded copies, not the second building.

saying these in an interview costs you the question

  • Calls cross-region replication the backup strategy
  • Thinks the destination keeps the pre-delete state
  • Assumes delete markers never replicate, so the second copy is safe
  • Expects replication lag to provide a usable recovery window
  • Cites a high durability figure as protection against a wrong delete
  • Believes a replica gives point-in-time restore