skip to content

Reliability & Failure Domains

Which failures the platform absorbs for you and which it hands back: a lost zone against a lost region, a degraded management API, recovery targets. Asked because designs assume more than is promised.

on this pageshow

questions

25

What does a recovery point objective promise about a managed service, and what does a recovery time objective promise instead?

level: juniorimportance: must knowfreq 70%

answer

  1. two numbers, one shared instant
  2. one looks back, one looks forward
  3. how much work against how long down
  4. capture frequency against procedure duration
  5. business states it, engineering prices it

basics

~20 s

A recovery point objective caps how much recent work the business accepts losing, measured backwards from the failure. A recovery time objective caps how long the service may stay unavailable, measured forwards from the same moment.

solid answer

~40 s

Both numbers are measured from the instant a service is declared lost, and they point in opposite directions. The **recovery point objective** looks backwards: how much recent work may be lost, usually expressed as a duration of writes. The **recovery time objective** looks forwards: how long the service may be unavailable before it is serving correct traffic again. They are bought with different mechanisms, so one can be excellent while the other is dreadful — a replica that trails by a second gives a tiny recovery point, but if promoting it and repointing every client takes a day, the recovery time is awful. Both are business statements that engineering prices, set per service rather than once for the whole estate.

go deeper

for a junior

Memorise the two directions: the recovery point looks backwards at lost work, the recovery time looks forwards at lost availability. Being able to say which is which, without mixing them, is the whole first-screen answer.

for a middle

Explain what sets each number in practice — capture frequency and replication lag for one, the full procedure duration for the other — and give an example where one is excellent and the other is bad.

for a senior

Show that you have measured a real procedure against its stated numbers, and that you count provisioning, replay, cutover and verification inside the recovery time rather than just the copy.

for a principal

Frame the pair as a purchase: state the cost curve as either number approaches zero, and who signs for the increment. Explain why different services in one system carry different numbers.

## The two numbers, and the instant they share Recovery objectives always come in a pair, and both are measured from the same instant: the moment a service is declared lost. A **recovery point objective** points backwards from that instant to the most recent state you can actually get back — it is the width of the window of work you accept losing. A **recovery time objective** points forwards from that same instant to the moment the service is serving correct traffic again — it is the length of outage you accept. Because they share a starting point and point in opposite directions, they get collapsed into one vague "recovery" number, and the collapse hides two facts that matter. They are bought with different mechanisms, and one of them can be excellent while the other is terrible. A managed store whose standby trails by a second has a tiny recovery point; if promoting that standby and repointing every client takes a working day, its recovery time is dreadful. A store captured once a night that can be brought back in ten minutes is the opposite shape. ## Reading the pair off a design | Question | Recovery point objective | Recovery time objective | |---|---|---| | Direction from the failure | backwards | forwards | | Unit | a duration of work, or a count of transactions | a duration of unavailability | | What missing it costs | work that no longer exists | revenue, trust, a contractual penalty | | Bought with | more frequent capture, continuous capture, a closely following replica | faster restore, capacity already running, a rehearsed cutover | | Usually limited by | how often state is captured, and replication lag | how long copying, replaying and cutting over take | Two practical readings follow: - The recovery point is set mainly by **capture frequency**, not by restore speed. If state is captured once a day and nothing else is kept, the worst case is close to a day of work however fast the restore runs. - The recovery time is set by the **whole procedure**, not by the copy step. Provisioning the replacement, replaying to the chosen moment, reconfiguring clients and verifying the data all sit inside it. ## They are business decisions that engineering prices An objective is not a measurement of what the platform happens to do; it is a statement of what the business tolerates, which engineering then prices. The order matters, and reversing it is the most common way these numbers become meaningless: 1. The service owner states the loss and the outage the business can absorb, in business terms — "we cannot lose a posted payment", "an hour of a stale reporting dashboard is survivable". 2. Engineering costs the mechanisms that would deliver those numbers, including the operational burden of running them. 3. The two sides settle on a number that is written down, owned and dated. 4. Somebody measures the real procedure and compares the measurement with the number. Skip the first step and you get "our recovery point objective is twenty-four hours", which is not an objective at all — it is a description of the capture schedule with the word *objective* attached. Skip the last and you get a number with no evidence behind it. ## Why the pair is set per service Recovery numbers attach to a service, not to an estate. An original record that cannot be reconstructed from anything else — a ledger of payments, a regulated archive — carries a much tighter recovery point than a copy derived from it, because the derived copy can be rebuilt. Providers reinforce this by making recovery a per-service purchase: how often state is captured, how long those restore points are retained, and whether a second copy is kept running are all per-service choices with per-service bills. When an auditor asks for "the recovery objectives", the answer is a table with one row per service, not a single number for the company. ## Where candidates go wrong - **Quoting capability as objective.** The platform can restore in two hours, so the objective becomes two hours — which means the business never stated what it actually needs. - **One number for both.** "Our recovery objective is four hours" is unanswerable: four hours of lost work, or four hours of downtime? - **Forgetting the tail of the procedure.** The copy finished inside the hour, but the service was down far longer because everything after the copy was improvised. - **Assuming the pair is cheap to tighten.** The cost curve steepens sharply as either number approaches zero, and the last increment is usually the most expensive part of the whole design. At a first screen, an interviewer wants the two definitions stated in the right direction, plus one sentence showing you know they are bought separately.

  • A managed store captures state once every twenty-four hours and can copy it back in twenty minutes. What pair of numbers does that actually support?
    The recovery point is up to roughly twenty-four hours of work, because that is the gap between captures and a failure can land just before the next one. The recovery time is the twenty minutes of copying plus everything after it — provisioning, replaying, repointing clients and verifying — so it is meaningfully longer than twenty minutes and only measurement tells you by how much.
  • Who owns these two numbers — the platform team or the service owner?
    The service owner states the tolerance, because the loss is a business loss. The platform team prices the mechanisms that would deliver it and reports what each increment costs to build and to operate. The number is then agreed, written down with an owner and a review date, and measured. A number produced by the platform team alone is a description of current capability wearing the word objective.
  • Why does an auditor ask for these objectives service by service rather than once for the system?
    Because the consequence of losing work differs sharply per service, and so does the cost of protecting it. An original record that nothing can reconstruct needs a far tighter recovery point than a copy that can be rebuilt from it, and recovery is purchased per service anyway — capture frequency, retention and whether a second copy runs are per-service choices with per-service bills.

saying these in an interview costs you the question

  • Treating the two objectives as a single recovery number
  • Saying the recovery point objective is how long a restore takes
  • Quoting the platform's capture schedule as if it were the objective
  • Assuming a fast restore also gives a small recovery point
  • Believing the numbers are technical limits rather than business decisions
  • Counting only the copy step and ignoring the cutover in the recovery time
open as a page

A provider publishes an availability commitment for a managed service — is that a guarantee, and what does it pay you when missed?

level: juniorimportance: must knowfreq 72%

basics

~20 s

An availability commitment is a contract term, not a promise the service stays up. If the provider's own measurement falls short, the remedy is a service credit against your bill for that service, which you normally have to claim.

open as a page

A checkout service runs in three availability zones and one zone goes dark at peak — what keeps serving, and what had to be true beforehand?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Instances in the two healthy zones keep serving, but only if the traffic entry point health-checks the dead ones out, the writable data copy is not stranded in the lost zone, and the survivors already had spare capacity.

open as a page

During a provider incident, running workloads keep serving but no new instance launches - which half of the platform is degraded, and what stops?

level: middleimportance: must knowfreq 62%

basics

~20 s

The control plane - the platform's management API that creates, changes and deletes resources - is degraded, while the data plane that carries request traffic keeps serving. Launching, scaling, replacing and failing over stop; already-running capacity does not.

open as a page

Why does a standby replica of a managed store give a near-zero recovery point yet fail to protect against a mistaken bulk delete?

level: middleimportance: must knowfreq 62%

basics

~20 s

A replica copies committed writes, so it reproduces the mistaken delete as faithfully as any other write. Only a point-in-time restore rewinds the data to a moment before the mistake, at the cost of a much longer recovery time.

open as a page

The platform's own identity service is down, so credential refresh fails — why do your services keep serving at first, and what ends that?

level: middleimportance: must knowfreq 62%

basics

~20 s

Only the refresh path is broken. A workload still holding an unexpired short-lived credential keeps calling platform services normally. Serving stops at the expiry cliff, when that credential's remaining lifetime runs out — often for many workloads within the same few minutes.

open as a page

An availability commitment quotes a percentage over a measurement window — what exactly is measured, and what does the window hide?

level: middleimportance: must knowfreq 58%

basics

~20 s

The figure is a ratio of eligible time or eligible requests the provider's own instrumentation judged available, over the eligible total in a fixed window — usually a calendar month, per service and per region. Averaging over that window hides short, severe outages.

open as a page

Three zones each run an order service at 70% of capacity at peak; one zone is lost — what happens, and what sizing rule prevents it?

level: middleimportance: must knowfreq 61%

basics

~20 s

Each survivor is asked for about 105% of its own capacity, so the service queues, slows and sheds load at the worst moment. Surviving one zone loss means capping each zone's peak utilisation at (N-1)/N — roughly 66% with three zones.

open as a page

Your managed store runs as one instance in a single zone — why might the headline availability commitment not apply to it at all?

level: middleimportance: should knowfreq 46%

basics

~20 s

Availability commitments attach to a deployment configuration, not to a service name. The headline figure normally requires redundant instances spread across separate availability zones, so a single instance in one zone earns a lower commitment or none at all.

open as a page

Your order database has a synchronous standby in a second zone and an asynchronous copy in a second region — what differs when you promote each?

level: middleimportance: should knowfreq 57%

basics

~20 s

Promoting the in-zone synchronous standby loses no acknowledged writes, because every commit already waited for it. Promoting the cross-region asynchronous copy loses whatever the lag held — seconds or minutes of confirmed orders that exist nowhere else.

open as a page

Your transcoding fleet's backlog is growing but every request to add a worker times out while running workers transcode normally - what now?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Accept a frozen fleet and manage demand instead of capacity: keep the durable queue absorbing arrivals, shed or defer the lowest-value work, freeze anything that could shrink or replace workers, and communicate delay rather than failure.

open as a page

A point-in-time restore of a payments ledger finished and the rows are correct, yet nothing works — what did the restore not bring back?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A restore returns rows, not a running service. The name clients resolve still points at the old endpoint, the restored instance has its own credentials, and in-flight messages, caches, derived indexes and registrations elsewhere are untouched by it.

open as a page

The platform's name-resolution and instance metadata services are degraded, yet the databases behind them are healthy — why does your service still fail?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Health and reachability are separate facts. Name resolution turns a dependency's name into an address and the metadata service hands a workload its identity and placement, so both sit in front of everything — a healthy database nobody can find or authenticate to is unreachable, and neither hop is usually drawn.

open as a page

Your dashboards are flat and green during a provider incident, yet customers report errors and the provider's status page says all clear — which signal do you trust?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Trust the customers. Telemetry ingestion is itself a shared platform service, so a flat chart is missing data rather than healthy traffic, and a status page is gated on human confirmation. The signal to believe is the one sharing fewest dependencies with the thing it measures.

open as a page

Which downtime does a provider's availability commitment typically exclude from its own measurement, and why does that matter to your contract?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Announced maintenance, downtime caused by tenant configuration or code, preview and trial tiers, suspension for policy or non-payment, and causes outside the provider's control are normally carved out — so measured downtime is a subset of what your customers actually felt.

open as a page

Everything in your checkout path is deployed across three zones, yet losing one zone took checkout down — what class of dependency explains that?

level: seniorimportance: should knowfreq 49%

basics

~20 s

Something in the path had exactly one home in the lost zone — a single writable database, a broker node, a lock or leader, a session cache, an outbound path, or a fleet whose spread had quietly drifted into one zone. The path is only as multi-zone as its least-spread hop.

open as a page

For a cross-region failover of a checkout stack, would you automate the decision or require a human to declare it, and why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Automate inside a region, where a synchronous standby loses nothing and the move is cheap; require a declaration across regions, where promoting a lagging copy accepts real data loss and the failure signal is often ambiguous. Automation is only safe with independent observers and fencing.

open as a page

How much pre-provisioned headroom should a platform team mandate for control-plane outages, and how do you justify paying for it?

level: principalimportance: should knowfreq 36%

basics

~20 s

Size headroom to the demand a workload must absorb while it can create nothing, not to average utilisation, and set it per tier rather than fleet-wide. Cheaper than any percentage: mandate that every failover path completes without creating a resource.

open as a page

An auditor demands a recovery point and a recovery time per service, so how would you set different numbers for a payments ledger and for the reporting copy derived from it?

level: principalimportance: should knowfreq 38%

basics

~20 s

Derive each number from the consequence of losing that service, not from what the platform already does. A ledger holds originals nothing can reconstruct, so it buys a tight recovery point; a reporting copy is derived data and can be rebuilt, so it buys a loose one cheaply.

open as a page

A dependency review finds several platform services with a single global home sitting under all three of your regions — what posture do you set?

level: principalimportance: should knowfreq 38%

basics

~20 s

Independence is capped by the least independent thing underneath, so decide per dependency: accept and document it, move it off the request path, define a deliberate degraded mode, or genuinely replace it. The lead's deliverable is that decision plus an honest availability story, not the removal of a dependency that is not yours.

open as a page

You are drafting an availability commitment for a reporting service built on three committed platform services in series — what can you honestly promise?

level: principalimportance: should knowfreq 36%

basics

~20 s

Less than the weakest link. Serial dependencies multiply, so a request path through three committed services promises the product of their figures — and the only ways to raise it are removing a link from the path, masking one with redundancy, or degrading without it.

open as a page

Leadership asks whether checkout should survive losing an entire region — how would you frame the choice against a second zone, and what does each posture cost?

level: principalimportance: should knowfreq 38%

basics

~20 s

Frame it as recovery time and acceptable data loss priced against the cost of each posture. Multi-zone is cheap and lossless; a second region costs duplicated capacity, duplicated data and continuous engineering to keep both true, and rises steeply from cold standby to active-active.

open as a page

During a management-API outage, which of your own automated loops can shrink a healthy transcoding fleet, and how do you stop them?

level: seniorimportance: nice to knowfreq 30%

basics

~10 s

Scale-in on a falling backlog, terminate-and-replace supervision, a deployment already in flight, and scheduled teardown can each destroy capacity you cannot rebuild. Suspend scale-in, switch replacement to drain-not-terminate, and pause deployments until creates succeed.

open as a page

Your recovery time objective was signed when the ledger held a tenth of today's data, so what silently changed and what did not?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

The restore's real duration grew roughly with the data, so the recovery time objective is now missed although nobody edited it. The recovery point is usually unchanged, because it follows capture frequency rather than volume.

open as a page

The provider's status page still shows everything normal thirty minutes into an outage you can measure — why, and what do you use instead?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

A status page is a published statement, not a live feed: the provider must detect, confirm and scope the impact before anyone posts, and pages are per service and per region rather than per tenant. Use your own out-of-band measurement as the trigger instead.

open as a page