Define Recovery Point Objective (RPO) and Recovery Time Objective (RTO) for a database, and give a concrete example of a system where the two numbers are very different.
answer
- RPO = data loss backwards; RTO = downtime forwards
- RPO from replication/archive cadence
- RTO from detection + failover + restore
- ledger: RPO ~0, RTO 30 min
- objective vs achieved - measure it
basics
~20 sRPO is how much recent data you can afford to lose, measured in time before the failure. RTO is how long the system may stay unavailable before it is serving again. RPO is about data loss; RTO is about downtime.
solid answer
~60 s**RPO - Recovery Point Objective**: the maximum acceptable amount of data loss, expressed as time. An RPO of 5 minutes means that after a disaster you accept losing up to the last 5 minutes of committed writes. It is set by how frequently data is durably copied off the primary - synchronous replication, log shipping cadence, backup frequency. **RTO - Recovery Time Objective**: the maximum acceptable time from failure to the service being usable again. It is set by detection time, failover automation, restore and replay speed, and the human steps in between. They are independent axes. A financial ledger might demand RPO near zero - no committed payment may vanish - while tolerating 30 minutes of downtime; the business can pause trading but cannot lose a transaction. Conversely, a public content site might tolerate an hour of data loss (regeneratable content) but need an RTO of 60 seconds because downtime is immediately visible to every user. Both are **business commitments**, not engineering preferences, and both should be measured, not merely declared.
go deeper
Give the two definitions crisply and one contrasting example. Being clear that RPO is data loss and RTO is downtime is most of the credit at this level.
Connect each number to its mechanism - RPO to synchronous versus asynchronous replication and archive cadence, RTO to detection, failover automation, and restore speed - and note detection time sits inside RTO.
Distinguish objective from achieved, insist on drill-measured numbers, and point out per-data-store RPO differences and the cross-store consistency problem after partial recovery.
Frame both as negotiated business commitments with a cost curve, discuss how tiering the estate keeps the expensive guarantees where the loss actually hurts, and how the commitments are evidenced.
## The two numbers Every disaster-recovery conversation is anchored on two quantities, both measured in time but describing different things. **Recovery Point Objective (RPO)** answers: *how far back in time may we be thrown?* If a failure happens at 14:00 and the most recent durable, recoverable state is 13:55, you have lost 5 minutes of writes. RPO is the maximum such loss the business accepts. RPO = 0 means no committed transaction may be lost. **Recovery Time Objective (RTO)** answers: *how long may we be down?* If the failure is at 14:00 and the service accepts writes again at 14:25, the achieved recovery time is 25 minutes - which includes detecting the failure, deciding to act, executing failover or restore, replaying logs, warming caches, and repointing the application. RTO is the maximum acceptable value of that total. A useful mental picture: on a timeline of the incident, RPO extends **backwards** from the failure (how much history is lost) and RTO extends **forwards** (how long the outage lasts). ## What sets each RPO is set by how often data is made durable somewhere that survives the failure: - Synchronous replication with commit acknowledgement from a second node gives an RPO of zero for the failures that node survives. - Asynchronous replication gives an RPO equal to the replication lag at the moment of failure - typically sub-second, but unbounded during a lag spike. - Continuous log archiving gives an RPO of roughly the archive interval plus transfer time. - Nightly backups alone give an RPO of up to 24 hours. RTO is set by the recovery path: - Promoting a warm standby that is already caught up: seconds to a couple of minutes, mostly detection plus routing. - Restoring a physical backup and replaying log: minutes to many hours, scaling with data size. - Restoring a logical dump: usually the slowest, because indexes and constraints must be rebuilt. ## Why they are independent The classic mistake is to treat them as one dial. Two examples make the independence obvious. **Payment ledger.** Losing a settled payment is a compliance and reconciliation disaster; being unable to accept new payments for 20 minutes is an inconvenience with a queue behind it. So: RPO near zero (synchronous commit to a second site, accepting the write-latency cost), RTO moderate (a careful, possibly human-approved failover is acceptable). **High-traffic content or feed site.** The content is re-derivable from upstream systems and losing an hour of derived data is annoying but survivable; being offline for an hour is on the front page. So: RPO an hour (cheap asynchronous replication, periodic snapshots), RTO under a minute (fully automated failover, multiple live replicas). **Analytics warehouse loaded nightly.** RPO of a day is fine because the source of truth is elsewhere and the load can be rerun; RTO may be a day too. Both loose, and that is a legitimate, deliberate, cheap position. ## Objective versus achieved RPO and RTO are *objectives* - targets negotiated with the business. The matching measured facts are sometimes called the achieved or actual recovery point and time. The gap between them is the real risk. A runbook claiming a 15-minute RTO while the last timed restore drill took four hours is not a 15-minute system; it is a four-hour system with a 15-minute aspiration. Interviewers care a lot about whether a candidate distinguishes the promise from the measurement. ## Related terms worth knowing - **MTTD / MTTR** (mean time to detect / recover) are the operational averages; RTO is a committed ceiling, not an average, and it usually starts at the moment of failure, not at the moment someone noticed. - **RPO applies per data store.** A service with a zero-RPO database and a 24-hour-RPO object store still loses a day of files; consistency across stores after a partial recovery is its own hard problem. - **Service-level agreements** normally express availability, and RTO plus incident frequency is what determines whether that availability is achievable. The short answer to give: RPO is acceptable data loss measured backwards from the failure; RTO is acceptable downtime measured forwards; they are set by replication and archiving on one side and failover and restore speed on the other; they are business commitments and must be measured, not assumed.
- Can a system have an RPO of zero and an RTO of zero at the same time?RPO of zero is achievable with synchronous commit to another site, at the cost of write latency and availability coupling. RTO of exactly zero is not achievable in practice - detection and failover always take some time - but active-active or multi-primary designs with client-side retry can push observed downtime to seconds or less. The honest answer is that both can be pushed very low, at steeply rising cost in latency, complexity, and money.
- Where does detection time fit into these numbers?It sits inside RTO, and it is frequently the largest single component. A failover that executes in 30 seconds but is triggered by a human who noticed an alert 20 minutes later yields a 20-minute recovery, not a 30-second one. Silent logical failures such as a bad migration are worse, because detection can take hours - which is also why RPO for those scenarios is governed by backup retention rather than replication.
Saving a document you are writing. RPO is how much typing you lose when the laptop dies - set by how often it autosaves. RTO is how long until you are typing again - set by how fast you can get a working machine.
saying these in an interview costs you the question
- Swapping the two definitions, or describing RPO as 'how long recovery takes'.
- Treating RPO and RTO as a single 'disaster recovery level' rather than independent axes with different costs.
- Quoting the runbook's target as if it were a measured capability, with no restore drill behind it.
- Claiming replicas alone give RPO zero for all failure types, ignoring that replication faithfully propagates a dropped table.
- Starting the RTO clock when someone noticed the outage instead of when the failure occurred.