skip to content

A runbook claims a database recovery point objective of 5 minutes and a recovery time objective of 30 minutes. How would you determine whether those numbers are actually true, and what continuous telemetry would tell you when they stop being true?

level: seniorimportance: should knowfreq 40%

answer

  1. recovery point = observable now; recovery time = must be exercised
  2. byte lag, archive backlog, backup age, degraded-sync flag
  3. time every drill phase, not just the total
  4. stop the clock at a successful client write
  5. trend restore time vs data growth; alert on stale proof

basics

~20 s

Measure, do not assume. Run timed restore and failover drills to get real recovery times, and monitor replication byte lag, archive backlog, and backup age continuously as live proxies for the recovery point. Alert when either exceeds the promised number, and trend restore times against data growth.

solid answer

~60 s

Two different verification problems. **Recovery point** is continuously observable, because its inputs are live: replication byte lag on each standby, age of the newest successfully archived log segment, archiver backlog depth, and age of the newest verified backup. I alert when any of these exceeds the promised 5 minutes, and I report the p99 over time rather than the average - a promise is a ceiling. **Recovery time** cannot be observed; it must be **exercised**. So: scheduled game days that fail over a real cluster and timed automated restore drills into a scratch environment, with the clock started at the injected failure and stopped when a client successfully writes and reads. Record every phase separately - detect, decide, fence, restore, replay, repoint, warm - because the phase breakdown tells you what to fix, and a single total tells you nothing. Then watch for **drift**: restore time grows with data volume, so I trend measured restore duration against database size and forecast the date the 30-minute promise breaks. I also treat a stale drill ("last proven 9 months ago") as an alertable condition in its own right.

code

text · 9 lines
text
failure injected            00:00
detected (alert fired)      00:04:10   <- 4m detection
decision to fail over       00:19:30   <- 15m human decision  ** dominant **
old primary fenced          00:20:05
standby promoted            00:20:25
first successful app write  00:24:50   <- 4.5m client reconnect
p95 latency back to normal  00:31:40   <- cache warm-up

declared RTO 30m | measured 31m40s to normal latency -> breach

go deeper

for a junior

Say that the only way to know is to try it - restore a backup and time it - and that replication lag and backup age are the everyday signals to watch.

for a middle

Separate the two: the recovery point is measurable live from lag, archive backlog and backup age; the recovery time has to be exercised in a drill, timed end to end.

for a senior

Add phase-level instrumentation, precise clock start and stop definitions, tail statistics rather than averages, degraded-mode alerting, and trending restore duration against data growth to forecast the breach.

for a principal

Treat it as a governance question: per-tier thresholds, staleness of proof as an alertable control, who is accountable when measurement crosses the declared line, and whether the engineering or the promise changes.

## Two claims, two very different verification methods A recovery-point promise is about a property the system has **right now** - how far behind the surviving copies are - so it can be measured continuously from live telemetry. A recovery-time promise is about behaviour during an event that has not happened, so it can only be established by **inducing the event**. Conflating them is the most common gap in a disaster-recovery programme: teams monitor lag diligently and have never once timed a restore. ## Measuring the recovery point continuously Instrument every path by which data leaves the primary, and alarm on the worst one that a given failure scenario would rely on: - **Replication lag in bytes**, per standby: the difference between the log position generated on the primary and the position flushed or replayed on the standby. Bytes, not seconds, is the honest measure of outstanding data; time-based lag inflates misleadingly on an idle primary because it is derived from the last replayed transaction's timestamp. - **Archive freshness**: the age of the newest successfully archived log segment, and the count of segments waiting to be archived. A failing archive command is the classic silent degradation - backup jobs keep reporting success while the recoverable point drifts backwards by hours, and the primary's disk fills as retained log accumulates. - **Backup age and verification age**: how old the newest backup is, and how long since one was proven restorable. - **Degraded-mode flags**: whether a nominally synchronous standby is currently disconnected and the primary has fallen back to asynchronous commit. That state is a live breach of a zero-loss promise and must page, not merely log. Report these as percentiles over a window. The promise is a ceiling, so p99 or max is the relevant statistic and the average is close to meaningless. ## Measuring the recovery time by exercising it There is no telemetry for this. There are only drills, in increasing fidelity: 1. **Automated restore drill.** On a schedule, restore a real backup into a scratch environment, replay to a target moment, start the engine, and run data assertions. Fully automated, so it can run weekly or nightly and produce a time series of restore duration. This validates the restore-and-replay portion. 2. **Failover drill.** Promote a standby in a staging or, better, production-equivalent topology, including fencing and client repointing, and measure how long until an application client successfully writes. 3. **Game day.** Inject a realistic failure without warning the responders, require them to use the runbook, and time from injection to verified recovery. This is the only exercise that measures detection and human decision time - typically the largest components - and the only one that surfaces the missing credential, the expired certificate, and the step that only one person knows. Instrument the drill with **phase timestamps**: failure injected, failure detected, decision made, fencing complete, data recovery started and finished, log replay finished, first successful client write, latency back to normal. A single total is a scoreboard; the phase breakdown is a work list. If detection is 18 minutes of a 26-minute total, buying faster storage is wasted money. Define the stop condition precisely and in the users' terms: not "the engine opened" but "a representative application transaction committed and read back successfully at acceptable latency". Teams that stop the clock at engine startup systematically under-report. ## Watching for drift Both numbers decay silently: - **Data growth** lengthens fetch, restore, and replay roughly linearly, so a 30-minute restore at 400 GB becomes hours at 3 TB with no change in configuration. Trend measured restore duration against database size and forecast when the promise breaks, then act before it does - more frequent base backups to shorten replay, parallel restore, faster storage, or moving that tier onto a standby-based recovery path. - **Topology and tooling changes** - a new region, a rotated key, an upgraded backup tool, a changed storage class - can break the recovery path while every routine metric stays green. Any such change should invalidate the last proof and trigger a fresh drill. - **Staleness itself is a risk.** Track "time since this cluster's last proven restore" as a first-class metric with a per-tier threshold, and treat breaching it as an incident, because it is the metric that tells you whether the rest of your evidence is still credible. ## What good looks like A dashboard per cluster showing: current and p99 replication byte lag, archive backlog and freshness, newest backup age, newest verified-restore age, last measured restore duration with its phase breakdown, and the declared objectives drawn as lines on the same charts. When the measurement crosses the declared line, either the engineering or the promise has to change - and saying so plainly is exactly the judgement an interviewer is testing.

  • What single metric best serves as a live proxy for the current recovery point exposure?
    Replication lag measured in bytes of write-ahead log not yet flushed on the surviving copy, taken as the maximum across the copies that would survive the scenario you care about. It directly represents the committed data that exists only on the primary. Pair it with archive backlog for the archiving path, and alert on the maximum of the two rather than the average of either.
  • Why do teams systematically under-report their measured recovery time?
    They start the clock when the incident was acknowledged rather than when the failure occurred, and they stop it when the database engine opens rather than when a real application transaction commits and reads back at acceptable latency. Both boundaries omit large components - detection and human decision at the front, client repointing and cache warm-up at the back. Drills must define both boundaries in user-visible terms.

saying these in an interview costs you the question

  • Presenting the runbook's declared objectives as if they were measured facts.
  • Monitoring replication lag in seconds only, which is misleading on an idle primary and hides the real byte exposure.
  • Never alerting on archive backlog, so a broken archiver degrades the recovery point invisibly while backups look healthy.
  • Reporting average lag or average restore time instead of the tail, when the promise is a ceiling.
  • Assuming a recovery time measured two years ago still holds after the database tripled in size.

context