Why does a standby replica of a managed store give a near-zero recovery point yet fail to protect against a mistaken bulk delete?
answer
- faithful copying is the feature
- a replica has no notion of intent
- replication survives loss, retention survives mistakes
- rewind to a moment inside the window
- restore yields a fresh store to cut over to
basics
~20 sA replica copies committed writes, so it reproduces the mistaken delete as faithfully as any other write. Only a point-in-time restore rewinds the data to a moment before the mistake, at the cost of a much longer recovery time.
solid answer
~40 sA standby replica exists to survive the loss of the failure domain the primary sits in, and it does that by applying the primary's committed writes continuously — which is why its recovery point is close to zero and why a mistaken delete reaches it in the same moment. Logical corruption is not a failure the replica can see; to it, the delete is a valid write. **Point-in-time restore** is the mechanism for that case: it rebuilds the data as of a chosen moment inside a retention window, so you can aim just before the bad write. The price is recovery time — a restore produces a fresh store you must cut traffic over to, and its duration grows with the amount of data. Real designs keep both, because they answer different questions.
go deeper
Remember the one-line split: a replica protects against losing a copy, a restore protects against a copy being changed wrongly. Say that a mistaken delete reaches the replica immediately.
Explain the mechanics: committed writes applied continuously against a base capture plus a change record replayed to a chosen moment, and what each of those buys in recovery point and recovery time.
Show the incident judgment — which mechanism matches which failure, why promotion after a bad write makes things worse, and how you reconcile writes made to the damaged store after the corruption.
Argue the portfolio: what retention window is worth buying given realistic detection time, which services get both mechanisms, and what the combination costs to run and to rehearse.
## Two mechanisms, two different enemies The pair of recovery numbers is bought with two mechanisms that are constantly confused, because both produce "another copy of the data". A **standby replica** is a second running copy of the store that continuously applies the primary's committed writes. Under **synchronous replication** the primary does not acknowledge a write until the standby has it, so the replica is never behind and the recovery point is effectively zero; under **asynchronous replication** the standby trails by a **replication lag** and the recovery point is roughly that lag. Recovery is by **promotion**: the standby becomes the primary. The enemy it defeats is the loss of the failure domain the primary sits in. A **point-in-time restore** rebuilds the data as it stood at a chosen moment. It is normally assembled from a base capture plus the change record that follows it, which is why the moment can be chosen with fine granularity rather than only at capture boundaries. The enemy it defeats is **logical corruption**: a mistaken bulk delete, a migration that rewrote the wrong rows, a defect that quietly corrupted values for hours. ## Why replication cannot save you from a mistake Replication is deliberately faithful. It has no notion of intent, so a delete issued by an authorised session is simply a committed write, and a correctly functioning replica applies it as promptly as it applies anything else. That faithfulness is the feature: a replica that diverged from its primary would be useless for failover. The consequence is the rule worth memorising: > Replication protects against losing a copy. Retention protects against changing a copy wrongly. The corollary is that a mirrored store is **not** a backup, however many copies exist, and "we replicate to three places" is not an answer to "what happens when someone runs the wrong statement". ## Comparing the two along the numbers they buy | | Standby replica | Point-in-time restore | |---|---|---| | Recovery point | zero under synchronous replication, roughly the lag otherwise | the granularity of the change record, inside the retention window | | Recovery time | short — promote and repoint clients | long — provision, copy, replay, cut over, verify | | Protects against | loss of the failure domain holding the primary | a bad write, a bad migration, a corrupting defect | | Does not protect against | anything committed by an authorised caller | anything older than the retention window | | Cost shape | a second store running continuously | storage of captures and change records, plus the restore itself | ## What the retention window bounds A restore can only reach back as far as the **retention window** the service keeps. Two consequences follow, and both show up in interviews: - A corruption discovered after the window has rolled past is unrecoverable from that store, whatever the durability of the underlying storage. Slow-burning logical corruption is therefore far more dangerous than a loud failure, because detection time competes with the window. - Extending the window is a real purchase, and providers differ in how far the window can be pushed and in what it costs, so the number is a design decision rather than a default to accept quietly. ## What a restore actually hands you A restore does not rewind the existing store in place. It normally produces a **fresh store** holding the chosen state, sitting at a new address, while the damaged original is still running. That is deliberate — it preserves the evidence and lets you compare — but it means the restore is not the end of the recovery. Somebody has to decide the store is correct, cut traffic over to it, and deal with everything that was written to the damaged original in the meantime. In the payments case, that reconciliation is often the slowest part: the ledger has been taking writes since the bad delete, and rewinding to a moment before it discards those too. This is also why restore duration grows with the amount of data and with the distance between the base capture and the chosen moment: more to copy, and more change record to replay. ## Using both, on purpose A regulated payments ledger typically carries both, for different lines in the recovery table: 1. A closely following standby, so the failure of the primary's failure domain costs seconds of work and a promotion. 2. Restore points with a retention window long enough to cover realistic detection time for a bad write. 3. A written decision about which mechanism answers which incident, because reaching for the wrong one under pressure is how a recoverable mistake becomes permanent — promoting a standby after a bad delete simply installs the corrupted state as the new primary. The reporting copy derived from that ledger usually carries neither at full strength, because it can be rebuilt from the ledger.
- A bad delete has been replicated. Why is promoting the standby the wrong reflex?Because the standby holds the deleted state too, so promotion installs the corruption as the new primary and costs you the damaged original you might have wanted for comparison. The mechanism that matches this incident is a restore aimed at a moment before the delete, followed by reconciling anything legitimately written since.
- What decides how far back a point-in-time restore can reach?The retention window of the captures and the change record kept alongside them. Anything older than the window cannot be reached from that store at all, which is why detection time matters: a corruption found after the window has rolled past is unrecoverable there, regardless of how durable the underlying storage is.
- Does a restore give you the same recovery point everywhere inside the window?Broadly yes where a continuous change record is kept alongside the base captures, since the target moment can be chosen finely. Where only periodic captures exist, the reachable moments are the capture boundaries and the effective recovery point is the interval between them. Which of those you have is worth checking before quoting a number.
A mirror on your desk reproduces whatever you write, including the word you wrote by mistake. A photocopy taken every hour is what lets you go back to before you wrote it.
saying these in an interview costs you the question
- Thinks a standby replica is a backup
- Says replication filters out mistaken or malicious writes
- Believes a restore rewinds the damaged store in place
- Assumes restore duration is independent of data volume
- Offers promotion as the answer to logical corruption
- Ignores the retention window when promising how far back you can go