skip to content

Your disaster-recovery plan claims a 30-minute recovery time objective for failing a service over to a secondary region. How would you verify that claim in practice, and what typically turns out to be wrong the first time you actually try it?

level: seniorimportance: must knowfreq 55%

answer

  1. a number nobody measured is a guess
  2. start the clock at detection, not at typing
  3. the standby was never sized for real traffic
  4. DNS, pools and caches lag the switch
  5. nobody ever tested failing back

basics

~20 s

An untested recovery time objective is an estimate, not a number. Verify it by executing a real failover with a clock running from detection to service restored, in a low-traffic window with a tested way back — then treat the measured time, not the plan's, as the truth.

solid answer

~50 s

You verify a recovery time objective by doing the failover for real and timing it end to end — from the moment the failure is detectable to the moment the service is genuinely serving users, not from when someone starts typing commands. Run it in a low-traffic window, with a rehearsed way back and a named person who can call it off. What breaks the first time is rarely the database promotion itself: the standby was scaled to a fraction of production and buckles when it takes full traffic, DNS TTLs and client connection pools keep sending traffic to the dead endpoint, secrets or configuration were never replicated, an IAM role or a regional quota is missing, caches are cold so the recovery looks like a second outage, and the deploy pipeline you would need to push a fix only exists in the primary. Also: almost nobody has tested failing back, and the measured recovery time is usually a large multiple of the stated one.

go deeper

for a junior

Know what RTO and RPO mean — maximum tolerable downtime and maximum tolerable data loss — and be able to say that neither is real until someone has actually performed a recovery and timed it.

for a middle

Explain how to time a failover honestly from detection to restored service, and name several concrete things that break: undersized standby, DNS and connection-pool inertia, unreplicated config, cold caches, missing regional quota or roles.

for a senior

Show the risk judgment: when to run a full production failover versus a partial traffic shift, what safety controls and exit criteria you set beforehand, and how you handle a measured time that badly misses the published objective.

for a principal

Own the honesty policy across the organization: which services must demonstrate a measured recovery objective on a schedule, who is allowed to publish an untested number, and how the cost of a warm standby is weighed against the business consequence of a real regional loss.

## The claim and why it is usually fiction **RTO (recovery time objective)** is the maximum tolerable time to restore service after a disaster. **RPO (recovery point objective)** is the maximum tolerable amount of data loss, expressed as time — an RPO of five minutes means you accept losing up to five minutes of writes. Both are commitments. Neither is a measurement until someone has executed the recovery with a clock running. A plan that says "30 minutes" and has never been exercised is an estimate made by the person who wrote the runbook, in a quiet room, assuming everything else works. This is the single most interview-worthy fact about disaster recovery: **the number is not real until it has been measured, and the first measurement is almost always much larger than the claim.** ## How you actually verify it **Define the clock honestly.** RTO runs from when the disaster becomes detectable to when users are genuinely served again. Teams flatter themselves by starting the clock when an engineer begins running commands, which hides detection, paging, mobilization, and the decision to fail over — often the majority of the elapsed time. Record the whole timeline with timestamps. **Pick the fidelity you can afford.** In descending order of truthfulness: - Full production failover, low-traffic window, announced, with a rehearsed way back. - Partial failover: shift a slice of production traffic to the secondary, or fail over one dependency at a time. - Full failover in an environment built the same way as production by the same automation — worth much less if the environments diverge. - Everything up to the irreversible step: page for real, assemble, verify access, check replication lag, dry-run the promotion. **Set the exit criteria before you start**, so the exercise cannot be quietly declared a success: what latency and error rate count as "restored", how long the secondary must sustain full traffic, and what data loss you will accept and verify. **Measure RPO too, not just RTO.** After the failover, actually check what was lost: compare row counts or sequence positions, look at the replication lag at the moment of promotion, and confirm the number matches what the plan promised. **Then fail back**, and time that as well. Failback is the half of the plan that has typically never been written, let alone rehearsed, and a team that can fail over but not fail back has bought a one-way ticket. ## What breaks the first time Almost never the part everyone worried about. The recurring findings: - **The standby is undersized.** It was provisioned at 10–20% of primary to save money. The runbook says "fail over"; it does not say "scale first, and here is how long that takes". Under real traffic the secondary saturates and you have now converted a regional failure into a total one. This is why redundancy is expressed as headroom — N+1 means one unit can be lost while the rest still serve peak — and a standby that cannot serve peak is not redundancy. - **Traffic does not move when the record does.** DNS TTLs, resolver and JVM-level caching, long-lived connection pools and keep-alive sessions all keep pointing at the dead endpoint. Ten minutes of a thirty-minute budget can disappear here. - **Cold everything.** Caches are empty, connection pools are unwarmed, autoscalers start from a low floor. The first minutes after "recovery" look like a fresh outage, and if the exit criterion was "traffic is flowing" you will declare victory too early. - **Config and secrets did not replicate.** Data was replicated because someone owned that; a feature-flag store, a certificate, a signing key or an environment variable was not. - **Access and quota.** The role that can promote the replica was removed in a permissions cleanup. The secondary region's account quota does not permit the instance count you need. Neither shows up on paper. - **Control-plane dependencies live in the failed region.** The deploy pipeline, the artifact registry, the bastion, even the runbook or the identity provider you log in with. If your ability to fix things depends on the thing that is broken, your RTO is unbounded. - **The decision is slow.** Even with a perfect mechanism, the team spends twenty minutes debating whether this is bad enough to justify failing over, because the criteria were never written down. - **Nobody has tested failback,** and returning to the primary risks losing writes taken in the secondary. ## The decision this exercise forces A production DR drill spends real risk to buy a real number. The counter-argument — "failing over might cause an outage" — is true, and it is exactly the point: if the plan is capable of causing an outage on a quiet Tuesday with the whole team watching, it will certainly cause one during a genuine regional failure at 3am with one tired engineer. You choose whether to discover that under controlled conditions or uncontrolled ones. When the answer is genuinely "we cannot fail over production", say so and take the partial: shift a percentage of traffic, exercise a single dependency, or run everything up to the irreversible step. A measured partial beats an unmeasured whole. ## Closing the loop The output is a measured RTO and RPO, plus owned findings. If the measurement is 70 minutes against a 30-minute claim, you have two legitimate choices and must pick one explicitly: **fund the work** that closes the gap (pre-scaled standby, shorter TTLs, replicated config, a written decision threshold), or **change the published objective** to the truth. Continuing to publish 30 while knowing it is 70 is the failure mode — it misleads every downstream team that planned around it. Then re-run the same drill after the fixes, because the fixes are untested too.

  • Your measured recovery time comes in at 70 minutes against a published 30-minute objective. What do you do?
    Pick one of two honest options and say so publicly. Either fund the work that closes the gap — pre-scale or warm the standby, shorten TTLs, replicate config and secrets, write down the decision threshold so debate does not eat twenty minutes — with owners and dates, or restate the objective as the measured number so dependent teams stop planning around a fiction. What you must not do is keep publishing 30. Then re-run the drill after the fixes, because a fix is also untested.
  • How does a drill measure the recovery point objective rather than just the recovery time?
    Write known, identifiable traffic right up to the moment of the fault, then after promotion check what actually survived: compare sequence positions or row counts against what the primary acknowledged, and record the replication lag at promotion time. That gives you a measured data-loss window to compare with the promised RPO. It also forces the question nobody enjoys — what the business does with the lost writes, and whether there is a reconciliation path.
  • Why is it a problem when the deploy pipeline or the identity provider lives only in the primary region?
    It makes your recovery time unbounded rather than merely long. If shipping a fix, logging into a console, or pulling an artifact requires the region that just died, the team cannot act at all — no runbook step helps. The check is to walk the recovery path asking, at every step, which region that step depends on; anything that answers "the one that failed" is a finding, and the drill is the only thing that reliably asks the question.
  • When is it legitimate to verify a failover in a pre-production environment instead of production?
    When the environment is built by the same automation from the same definitions and differs only in scale, a pre-production failover honestly tests the mechanics: the promotion, the runbook steps, the tooling. It cannot test capacity, real client behaviour, DNS and connection-pool inertia, or production quotas and permissions. Use it to remove the mechanical defects cheaply, then still run something in production — even a partial traffic shift — for the parts scale and real clients own.

saying these in an interview costs you the question

  • The runbook documents the RTO, so we meet it
  • Start the clock when the engineer runs the first command
  • The standby exists, so it can take the load
  • Failover is tested; failback will just work
  • We can't drill it in production, so we can't measure it

context