skip to content

What does a recovery point objective promise about a managed service, and what does a recovery time objective promise instead?

level: juniorimportance: must knowfreq 70%

answer

  1. two numbers, one shared instant
  2. one looks back, one looks forward
  3. how much work against how long down
  4. capture frequency against procedure duration
  5. business states it, engineering prices it

basics

~20 s

A recovery point objective caps how much recent work the business accepts losing, measured backwards from the failure. A recovery time objective caps how long the service may stay unavailable, measured forwards from the same moment.

solid answer

~40 s

Both numbers are measured from the instant a service is declared lost, and they point in opposite directions. The **recovery point objective** looks backwards: how much recent work may be lost, usually expressed as a duration of writes. The **recovery time objective** looks forwards: how long the service may be unavailable before it is serving correct traffic again. They are bought with different mechanisms, so one can be excellent while the other is dreadful — a replica that trails by a second gives a tiny recovery point, but if promoting it and repointing every client takes a day, the recovery time is awful. Both are business statements that engineering prices, set per service rather than once for the whole estate.

go deeper

for a junior

Memorise the two directions: the recovery point looks backwards at lost work, the recovery time looks forwards at lost availability. Being able to say which is which, without mixing them, is the whole first-screen answer.

for a middle

Explain what sets each number in practice — capture frequency and replication lag for one, the full procedure duration for the other — and give an example where one is excellent and the other is bad.

for a senior

Show that you have measured a real procedure against its stated numbers, and that you count provisioning, replay, cutover and verification inside the recovery time rather than just the copy.

for a principal

Frame the pair as a purchase: state the cost curve as either number approaches zero, and who signs for the increment. Explain why different services in one system carry different numbers.

## The two numbers, and the instant they share Recovery objectives always come in a pair, and both are measured from the same instant: the moment a service is declared lost. A **recovery point objective** points backwards from that instant to the most recent state you can actually get back — it is the width of the window of work you accept losing. A **recovery time objective** points forwards from that same instant to the moment the service is serving correct traffic again — it is the length of outage you accept. Because they share a starting point and point in opposite directions, they get collapsed into one vague "recovery" number, and the collapse hides two facts that matter. They are bought with different mechanisms, and one of them can be excellent while the other is terrible. A managed store whose standby trails by a second has a tiny recovery point; if promoting that standby and repointing every client takes a working day, its recovery time is dreadful. A store captured once a night that can be brought back in ten minutes is the opposite shape. ## Reading the pair off a design | Question | Recovery point objective | Recovery time objective | |---|---|---| | Direction from the failure | backwards | forwards | | Unit | a duration of work, or a count of transactions | a duration of unavailability | | What missing it costs | work that no longer exists | revenue, trust, a contractual penalty | | Bought with | more frequent capture, continuous capture, a closely following replica | faster restore, capacity already running, a rehearsed cutover | | Usually limited by | how often state is captured, and replication lag | how long copying, replaying and cutting over take | Two practical readings follow: - The recovery point is set mainly by **capture frequency**, not by restore speed. If state is captured once a day and nothing else is kept, the worst case is close to a day of work however fast the restore runs. - The recovery time is set by the **whole procedure**, not by the copy step. Provisioning the replacement, replaying to the chosen moment, reconfiguring clients and verifying the data all sit inside it. ## They are business decisions that engineering prices An objective is not a measurement of what the platform happens to do; it is a statement of what the business tolerates, which engineering then prices. The order matters, and reversing it is the most common way these numbers become meaningless: 1. The service owner states the loss and the outage the business can absorb, in business terms — "we cannot lose a posted payment", "an hour of a stale reporting dashboard is survivable". 2. Engineering costs the mechanisms that would deliver those numbers, including the operational burden of running them. 3. The two sides settle on a number that is written down, owned and dated. 4. Somebody measures the real procedure and compares the measurement with the number. Skip the first step and you get "our recovery point objective is twenty-four hours", which is not an objective at all — it is a description of the capture schedule with the word *objective* attached. Skip the last and you get a number with no evidence behind it. ## Why the pair is set per service Recovery numbers attach to a service, not to an estate. An original record that cannot be reconstructed from anything else — a ledger of payments, a regulated archive — carries a much tighter recovery point than a copy derived from it, because the derived copy can be rebuilt. Providers reinforce this by making recovery a per-service purchase: how often state is captured, how long those restore points are retained, and whether a second copy is kept running are all per-service choices with per-service bills. When an auditor asks for "the recovery objectives", the answer is a table with one row per service, not a single number for the company. ## Where candidates go wrong - **Quoting capability as objective.** The platform can restore in two hours, so the objective becomes two hours — which means the business never stated what it actually needs. - **One number for both.** "Our recovery objective is four hours" is unanswerable: four hours of lost work, or four hours of downtime? - **Forgetting the tail of the procedure.** The copy finished inside the hour, but the service was down far longer because everything after the copy was improvised. - **Assuming the pair is cheap to tighten.** The cost curve steepens sharply as either number approaches zero, and the last increment is usually the most expensive part of the whole design. At a first screen, an interviewer wants the two definitions stated in the right direction, plus one sentence showing you know they are bought separately.

  • A managed store captures state once every twenty-four hours and can copy it back in twenty minutes. What pair of numbers does that actually support?
    The recovery point is up to roughly twenty-four hours of work, because that is the gap between captures and a failure can land just before the next one. The recovery time is the twenty minutes of copying plus everything after it — provisioning, replaying, repointing clients and verifying — so it is meaningfully longer than twenty minutes and only measurement tells you by how much.
  • Who owns these two numbers — the platform team or the service owner?
    The service owner states the tolerance, because the loss is a business loss. The platform team prices the mechanisms that would deliver it and reports what each increment costs to build and to operate. The number is then agreed, written down with an owner and a review date, and measured. A number produced by the platform team alone is a description of current capability wearing the word objective.
  • Why does an auditor ask for these objectives service by service rather than once for the system?
    Because the consequence of losing work differs sharply per service, and so does the cost of protecting it. An original record that nothing can reconstruct needs a far tighter recovery point than a copy that can be rebuilt from it, and recovery is purchased per service anyway — capture frequency, retention and whether a second copy runs are per-service choices with per-service bills.

saying these in an interview costs you the question

  • Treating the two objectives as a single recovery number
  • Saying the recovery point objective is how long a restore takes
  • Quoting the platform's capture schedule as if it were the objective
  • Assuming a fast restore also gives a small recovery point
  • Believing the numbers are technical limits rather than business decisions
  • Counting only the copy step and ignoring the cutover in the recovery time