Break down everything that consumes wall-clock time between a database primary dying and the application serving writes again. Which of those components usually dominate, and which can be engineered away?
answer
- clock starts at failure, not at the page
- detect / decide / fence / recover / repoint / warm / verify
- human approval is often the biggest slice
- standby kills restore+replay; proxy kills DNS delay
- logical damage: failover useless, restore required
basics
~20 sDetect, decide, fence the old primary, promote or restore, replay logs, repoint clients, warm caches, verify. Detection and human decision usually dominate, then restore and replay if there is no standby. Automation removes detection and decision time; a warm standby removes restore and replay.
solid answer
~1 minThe recovery clock starts at the failure, not at the alert. Components: 1. **Detect** - health checks converge. Seconds if automated, tens of minutes if a human notices. 2. **Decide** - automatic policy or a human approval step. Human approval is often the single largest slice. 3. **Fence** - stop the old primary accepting writes, so promotion cannot cause split brain. Skipping it is faster and sometimes catastrophic. 4. **Recover the data** - promoting a caught-up standby is seconds; restoring from backup is fetch + decrypt + decompress + write + replay, scaling with data volume and log distance from the base backup. 5. **Repoint clients** - DNS TTL, connection-pool reconnect and retry behaviour, proxy or service-discovery convergence. This one silently eats minutes when TTLs are long or pools cache resolved addresses. 6. **Warm up** - cold buffer cache and cold plan caches make the service technically up but functionally slow. 7. **Verify** - checks before declaring recovery. Engineer away: detection and decision through automated failover with a quorum-based leader election; restore and replay through a pre-built warm standby; repointing through a proxy or virtual IP instead of DNS. What remains irreducible is fencing, a little consensus latency, and warm-up.
code
text · 8 lineshost failure, warm standby present
detect 20s | decide (auto) 5s | fence 10s | promote 15s | clients reconnect 30s | warm-up ~2m
total to healthy: ~3 min
bad migration, restore required
detect 25m | decide 10m | identify target moment 15m
fetch 1.8TB 40m | restore+decrypt 35m | replay 6h of WAL 50m | verify 15m | repoint 5m
total: ~3h55mgo deeper
List the main steps - notice the failure, switch to another database, point the application at it - and know that promoting an existing standby is far faster than restoring a backup.
Give the full component breakdown with rough magnitudes, and explain why a warm standby removes restore and replay while automation removes detection and decision delay.
Emphasise the parts teams forget: the clock starting at failure, human approval as the dominant slice, fencing, client-side reconnection behaviour, cache warm-up, and recovery time degrading with data growth.
Discuss the trade you are actually making - automated failover speed versus the blast radius of a wrong promotion - and how you state recovery-time commitments per failure scenario across a portfolio.
## The clock starts at the failure A recovery-time objective is measured from the moment the service stops working, not from the moment a human noticed. This distinction alone reclassifies many systems: a 90-second failover triggered by a page that a sleeping engineer acknowledges after 18 minutes is an 18-minute recovery. ## The components, in order **1. Detection.** How long until something concludes the primary is gone. Determined by health-check interval, failure threshold, and network timeouts. Automated detection converges in seconds to tens of seconds. The tension is that aggressive thresholds cause spurious failovers on transient blips - and a spurious failover is itself an outage - so detection time is a deliberate trade against false positives, not something to minimise blindly. **2. Decision.** Automatic policy or human approval. Many organisations require a human in the loop for a database failover, especially cross-region, because the downside of a wrong promotion (split brain, lost writes) is severe. That choice can be right, but it must be costed honestly: paging, context-gathering, and approval routinely add 10-30 minutes, which typically makes it the largest single component in any system that has a standby. **3. Fencing / STONITH.** Before promoting, you must guarantee the old primary can no longer accept writes - revoke its network access, take its storage lease, or power it off. Without fencing, a primary that was merely partitioned comes back and both nodes accept writes, producing divergence that is far more expensive than the outage. Fencing costs seconds to a minute and is not optional. **4. Data recovery.** Two very different worlds: - *Promote a standby.* A caught-up standby needs only to finish applying what it has and open for writes: seconds. This is why standbys, not backups, are the tool for meeting a tight recovery-time objective. - *Restore from backup.* Time = fetch from storage (bandwidth-bound, worse from archival tiers) + decrypt + decompress + write to local storage (bound by disk write throughput) + replay all log since the base backup (single-threaded on many engines, and often the surprise dominant term when the base backup is old) + any post-recovery work. For a multi-terabyte database this is hours. Reducing it means more frequent base backups (shorter replay), parallel restore, faster storage, and larger network pipes. **5. Repointing clients.** The database being up is not the service being up. Time here comes from DNS record TTL and negative caching, application connection pools that cache resolved addresses or hold dead connections until a long timeout, proxies or service-discovery layers converging, and retry and circuit-breaker back-off windows that keep clients from reconnecting promptly. This component is routinely under-modelled: teams measure the database promotion and forget the five minutes clients spend not noticing. Using a proxy, virtual IP, or service mesh endpoint that flips instantly - rather than DNS - largely eliminates it, as does configuring pools with short connection validation and prompt retry. **6. Warm-up.** The new primary starts with a cold buffer pool and cold caches; query latency can be several times normal until working-set pages are read in. If "recovered" means "serving at acceptable latency", warm-up belongs inside the recovery time. Keeping the standby's cache warm by having it serve read traffic helps materially. **7. Verification and declaration.** Sanity checks, confirmation that writes succeed, comms to stakeholders. Small but real, and worth scripting. ## What dominates, and what you can remove In practice: with a standby, **detection plus human decision plus client repointing** dominate; the actual promotion is trivial. Without a standby, **restore plus replay** dominate and scale with data volume, which is why recovery time quietly degrades as the database grows - a system that met its objective at 200 GB may miss it badly at 3 TB with no change other than growth. Engineerable: - Automated leader election with a quorum-based coordinator removes detection and decision latency, at the price of trusting automation and needing a robust fencing story. - A warm standby, kept close, removes restore and replay entirely for infrastructure failures. - A proxy or virtual IP removes DNS and client-cache delay. - More frequent base backups shorten replay for the restore path. - Rehearsal removes the improvisation time that dominates any un-drilled procedure. Irreducible: fencing, consensus round trips, some warm-up, and the residual risk-management time a human decision buys you. ## The failure class that breaks the model All of the above addresses infrastructure failure. **Logical** damage - a dropped table, a bad migration - is replicated to every standby instantly, so failover does nothing and the recovery path is a restore. That is why a system can have a 60-second recovery time for host failure and a six-hour recovery time for a bad deploy, and why mature answers state recovery time per scenario rather than as one number.
- Why is skipping fencing before promoting a standby dangerous, even though it saves time?If the old primary was only partitioned rather than dead, it may still be accepting writes from clients that can still reach it. Promoting a standby then yields two primaries whose histories diverge, and reconciling divergent writes afterwards is manual, error-prone, and often impossible without data loss. Fencing - revoking network access, seizing the storage lease, or powering the node off - is what makes promotion safe, so the seconds it costs are not optional.
- Why can a system meet its recovery-time objective for a failed host but badly miss it for a bad migration?Host failure is repaired by promoting a standby, which takes seconds because the data already exists elsewhere. A bad migration is logical damage that replicates to every standby immediately, so no standby holds a clean copy; recovery means restoring a backup and replaying the log to just before the damage, plus the time to identify the exact moment. That path scales with data size and is typically orders of magnitude slower, so recovery-time objectives must be stated per failure scenario.
- How would you shorten the client-repointing portion of a failover?Put a proxy, virtual IP, or service-discovery endpoint in front of the database so the address clients hold never changes and the flip is instantaneous, rather than relying on DNS with its TTL and negative caching. Then tune the client side: short connection-validation and socket timeouts, pools that discard connections on error rather than blocking, and retry with bounded back-off so clients reconnect promptly instead of waiting out a long circuit-breaker window.
saying these in an interview costs you the question
- Starting the recovery clock when the alert was acknowledged rather than when the failure occurred.
- Quoting only the promotion time and ignoring detection, human decision, client reconnection, and cache warm-up.
- Promoting a standby without fencing the old primary, risking split brain.
- Assuming replicas protect against a dropped table or bad migration, so failover is the answer to every scenario.
- Treating recovery time as a fixed property, when restore-based recovery grows with data volume and quietly breaches the objective as the database grows.