skip to content

What is a deliberately delayed replica, one configured to stay (say) one hour behind the primary, and which failure does it protect against that a normal standby and a nightly backup do not?

level: seniorimportance: nice to knowfreq 26%

answer

  1. receives now, applies an hour later
  2. anti-mistake, not anti-crash
  3. first action: stop apply
  4. extract rows vs promote at a point
  5. not a backup substitute, not an HA target

basics

~20 s

A delayed replica receives the change stream immediately but deliberately withholds applying it for a fixed interval, so it holds a live copy of the database as it was an hour ago. That gives you a fast rewind after a destructive human or application error, which normal standbys replicate instantly.

solid answer

~50 s

A delayed replica streams changes normally but applies them only after a configured delay, for example one hour. It is not for read scaling and not for failover freshness; it is an anti-mistake device. A normal standby is useless against an accidental `DELETE` without a `WHERE` clause or a bad migration, because it faithfully replicates the destruction within milliseconds. A backup does protect you, but restoring a large database plus replaying logs to just before the mistake can take hours. A delayed replica is already running and already holds the pre-mistake state, so recovery is minutes: stop its apply immediately (the clock is literally ticking), then either read the lost rows out of it and copy them back, or advance it to the moment just before the bad statement and promote it. Costs: a full extra copy, no read value for fresh data, and an operational trap if you forget to stop apply. It complements, and never replaces, backups.

go deeper

for a junior

Know it exists and what it is for: a copy deliberately kept behind so a mistaken DELETE can be undone.

for a middle

Explain the mechanism (data received now, applied later), and contrast it with a standby and with restoring a backup.

for a senior

Own the runbook: stop apply immediately, choose extraction versus promotion, size the window, and account for the extra copy and its non-role in failover.

for a principal

Position it in the recovery strategy: which error classes it covers versus point-in-time restore and offsite backups, what it costs, and what recovery-time objective it actually buys.

## What it is A delayed replica is an ordinary physical replica with one behaviour changed: it receives the primary's change stream in real time but deliberately holds each change for a configured interval before applying it. Configured with a one-hour delay, its visible data is the primary as of one hour ago, and it stays exactly one hour behind as time passes. PostgreSQL expresses this with a recovery delay setting on the standby; MySQL with a delay parameter on the replication channel. Crucially the data is already transferred; only the apply is deferred. So the replica is not missing anything, it is holding the future in its log. ## The failure it addresses Replication protects against machine and site failure. It does not protect against a valid instruction that you did not mean: a `DELETE` or `UPDATE` that missed its `WHERE` clause, a deploy whose migration drops or rewrites a column, an admin script run against production instead of staging, or an application bug that mass-updates rows. These are correct-looking transactions, so every standby applies them faithfully and instantly. Your high-availability setup does not help; it replicates the damage. Backups do help, but on the timescale of restore plus log replay. For a multi-terabyte database that can be hours of downtime, and point-in-time recovery is an all-or-nothing rewind of the whole database, discarding the good writes that happened after the mistake. ## Why a delayed replica is different Because it is a live server holding the pre-mistake state, it gives you options within minutes: - **Extract**: query the delayed replica for the rows that were destroyed and copy just those back into the primary. This is surgical, keeps every unrelated write made since, and is usually the right move for a single bad statement. - **Promote**: stop its apply just before the offending transaction, let it finish recovery to that point, and promote it as the new primary. This is the fast full rewind when the damage is broad, at the cost of losing legitimate writes made after that point. ## The operational realities The delay is a countdown, not a safety net. The moment you notice the mistake, the first action is to stop apply on the delayed replica; otherwise, at the end of the delay, it will apply the destructive change too and your copy is gone. That makes the runbook part of the design: who is paged, what the single command is, and how it is tested. Sizing the delay is a trade-off. Too short and nobody notices the incident in time; too long and the copy is far behind, which makes extraction messier (more subsequent legitimate changes to reconcile) and makes promotion lose more data. Delays between one and twenty-four hours are typical, and some teams run two, a short one and a long one. Cost: a full extra copy of the database, plus storage for the pending change stream. It carries no read-scaling value for current data, and it must not be counted as a failover target, since promoting it means losing the delay window of writes. It also does not replace backups. It protects only within its delay window and only against logical errors; it does not protect against corruption that predates the window, against losing the whole environment, or against deletion of live infrastructure. Backups must remain, offsite and tested. ## How it interacts with the change stream type With physical (block-level) replication the delayed copy is the whole cluster, so extraction means querying it and shipping rows back. With logical replication you can be more selective about which tables the copy holds, at the cost of not having a promotable full replica. Either way, verify that the mechanism actually delays apply rather than merely delaying acknowledgement. ## How to present it Say what it is in one sentence, name the specific failure class (human and application error, which normal replicas faithfully copy), contrast the recovery time against restore-from-backup, and show operational awareness: stop apply first, choose extract versus promote, size the window, and keep backups anyway.

  • Why does an ordinary hot standby not protect against a bad DELETE?
    Because the DELETE is a legitimate committed transaction. Replication exists to make standbys identical to the primary, so it applies the deletion within milliseconds. Standbys defend against node, disk, and site failure, not against instructions the operator did not intend.
  • When would you extract rows from the delayed replica instead of promoting it?
    When the damage is narrow, for example one table or one statement, and there has been substantial legitimate write traffic since. Extraction restores just the lost rows and keeps everything else, whereas promotion rewinds the entire database to the delayed point and discards every write made after it.

It is an undo buffer with an expiry timer: the mistake is queued to be copied too, and you have until the timer runs out to hit stop.

saying these in an interview costs you the question

  • Calling it a backup replacement
  • Counting it as a failover target for high availability
  • Forgetting that the destructive change is already in its log and will be applied when the delay elapses unless apply is stopped
  • Thinking the delay works by not transferring data, and therefore that it lowers network cost
  • Setting a delay of a few minutes, which nobody can react within

context