skip to content

How do you decide what recovery point and recovery time objectives a given database should have, when every tightening of those numbers costs money and adds write latency or operational complexity? Walk through how you would tier a portfolio of databases.

level: principalimportance: should knowfreq 30%

answer

  1. impact per unit time drives the numbers, not engineering taste
  2. reconstructible data is the biggest recovery-point lever
  3. 3-4 priced, drilled tiers beat bespoke targets
  4. sync commit taxes every write; automation can cause outages
  5. tier the capability, not the single store

basics

~20 s

Derive the numbers from business impact, not engineering taste: what a minute of downtime and a minute of lost data cost, and whether the data is reconstructible elsewhere. Then define three or four standard tiers with fixed architectures and prices, and make owners choose a tier and fund it.

solid answer

~1 min

I start from **impact per unit time**, asked of the business owner: what does an hour of unavailability cost in revenue, regulatory exposure, safety, and reputation, and what does losing the last N minutes of writes cost - is the data reconstructible from an upstream system, a client-side queue, or paper? Those answers usually collapse a portfolio into three or four **tiers**, each a fixed, pre-engineered architecture with a known bill: - Tier 1: synchronous cross-zone commit, automated failover, minutes of recovery - expensive, adds per-commit latency. - Tier 2: asynchronous standby plus continuous log archiving, semi-automated failover, tens of minutes. - Tier 3: backups and archiving only, restore-based recovery, hours. - Tier 4: rebuild from an upstream source; backups are a convenience. Tiers beat bespoke design because they are testable, staffable, and priceable, and because they force an explicit choice: an owner who wants tier 1 pays for tier 1. Then I make it honest - drill each tier, publish measured versus declared numbers, and take the political hit early rather than during an incident. The most valuable move is often **reducing the need**: idempotent writes, client-side buffering, and reconstructible data cut the required recovery point far more cheaply than synchronous replication does.

go deeper

for a junior

Recognise that these numbers come from the business and that stricter targets cost more money and add latency; do not try to own the decision.

for a middle

Connect each candidate target to the architecture that delivers it and name the concrete price of each - round trip per commit, standby capacity, drill effort.

for a senior

Drive the tiering: a small set of pre-engineered, drilled, priced tiers, per-tier drill cadence, and published measured-versus-declared numbers with a growth forecast.

for a principal

Own the trade explicitly - expected loss versus permanent cost, availability coupling introduced by zero-loss guarantees, capability-level rather than store-level tiering, and reducing the requirement through idempotency and upstream replay before buying the expensive architecture.

## The numbers are business inputs, not engineering outputs Engineers cannot choose a recovery point objective, because the objective is a statement about acceptable loss - a business risk position. The engineering job is to (a) elicit the impact, (b) price the architectures that meet candidate targets, and (c) make the trade-off visible enough that a decision gets made deliberately rather than by default. The questions that actually generate the numbers: - **What does downtime cost per unit time?** Sometimes it is linear (lost transactions per minute), sometimes it is a step function (a market close, a regulatory reporting deadline, a nightly batch window that cannot slip), and sometimes it is non-financial (safety, patient care, contractual penalties). - **What does lost data cost?** Crucially, *is it reconstructible?* If clients retry idempotently, if an event bus retains the last day, or if the upstream system can replay, then losing five minutes of database writes costs a replay job, not the data. Reconstructibility is the single biggest lever on recovery-point cost. - **What is the failure frequency you are protecting against?** Paying a permanent per-commit latency tax on every transaction to protect against an event with a multi-year return period is a real trade, not an obvious one. - **What are the external obligations?** Regulatory rules, contractual service levels, and insurance requirements sometimes set floors regardless of internal economics. ## Costs of tightening, stated plainly - Tightening the **recovery point** toward zero means synchronous commit - a network round trip on every write, which is sub-millisecond within a zone and tens of milliseconds across regions. That tax lands on all traffic all the time. It also couples availability to the standby: strict synchronous means the primary blocks when the standby is unreachable, so a zero-loss guarantee can *reduce* uptime. - Tightening the **recovery time** means running warm standby capacity that is idle for its purpose (mitigated by serving reads), automated failover with a quorum coordinator, and the operational maturity to be safe about fencing. Automation itself carries risk: a spurious failover is an outage you caused. - Both mean more environments, more drills, more runbooks, more on-call surface, and more people who must understand the topology. Complexity is a recurring cost paid in incidents caused by the disaster-recovery machinery itself. ## Tiering the portfolio Bespoke targets per database do not scale - they cannot be staffed, tested, or budgeted. Define a small number of tiers, each a **pre-engineered, drilled, priced** package: | Tier | Recovery point | Recovery time | Architecture | Relative cost | |---|---|---|---|---| | 1 | ~0 | minutes | synchronous cross-zone commit, async cross-region standby, automated failover, proxy-based repointing | highest, plus write latency | | 2 | seconds | tens of minutes | async standby, continuous archiving, semi-automated failover | moderate | | 3 | minutes | hours | backups plus archiving, restore-based recovery | low | | 4 | hours-days | hours-days | periodic backups; primary path is rebuild from upstream | minimal | Governance that makes tiers work: every data store has a named business owner; the owner selects a tier and the cost is charged to them; each tier has a mandatory drill cadence and a published measured-versus-declared record; a store with no tier defaults to the cheapest, which is what actually forces owners to engage. ## Cross-store consistency A subtlety that separates strong candidates: recovery objectives apply per data store, but a business transaction usually spans several. A tier-1 database recovered to the failure instant alongside a tier-3 object store recovered to six hours earlier produces rows pointing at files that do not exist. So tiering must be applied to a **business capability**, not to isolated databases, and the recovery plan needs a reconciliation step - or the capability's weakest store defines its real objective. ## Reduce the requirement before you buy it The cheapest way to meet a hard recovery point is often to need less of it: - **Idempotent, retryable writes** with client-side or gateway-side buffering mean a few minutes of lost database writes are replayed rather than lost. - **An event log upstream** with retention longer than the worst recovery point turns data loss into a replay job. - **Degraded-mode operation** - read-only serving from a replica, or queueing writes at the edge - converts a hard outage into partial service and relaxes the recovery-time requirement dramatically. Proposing these before proposing cross-region synchronous replication is the mark of someone who has actually paid for the latter. ## Closing the loop Whatever is chosen must be measured and reported: drill results per tier, measured versus declared, and a forecast of when growth will breach a tier's promise. When the measurement crosses the promise, the choice is explicit - fund the next tier or restate the objective. Carrying an undeclared gap is the failure mode that turns a bad day into a company-defining one.

  • An owner insists on a zero recovery point objective for a system that today runs asynchronous replication. What do you put in front of them?
    The full price: a network round trip added to every commit, a measured latency impact on the write path, the availability coupling that makes the primary block or degrade when the standby is unreachable, and the ongoing cost of the extra node and drills. Then the alternatives that may deliver the same business outcome for less - idempotent retryable writes, an upstream event log with sufficient retention, or client-side buffering. The decision is theirs, but it must be made against real numbers rather than a preference for the word 'zero'.
  • How do differing recovery objectives across stores in one business capability cause problems?
    Recovery points are per store, so a database restored to the failure moment and an object store or search index restored to hours earlier leave dangling references - rows pointing at missing files, or an index describing records that no longer exist. The capability's honest objective is set by its weakest store unless the recovery plan includes an explicit reconciliation or rebuild step. Tier the capability as a unit and design the reconciliation before you need it.
  • When is it correct to loosen a recovery objective rather than invest to meet it?
    When measurement shows the current architecture cannot meet the declared number and the business impact does not justify the next tier's cost - for example an internal reporting store whose users tolerate a day of unavailability. Restating the objective honestly is far better than carrying an undeclared gap, because it lets dependent teams plan and stops the organisation from believing in a guarantee that does not exist.

Insurance underwriting: you do not buy the maximum cover for every asset, you price expected loss against premium, group assets into standard policies, and re-underwrite when the asset's value changes.

saying these in an interview costs you the question

  • Setting recovery objectives from engineering preference or vendor defaults rather than from business impact.
  • Applying the strictest tier everywhere 'to be safe', paying permanent write latency and cost for rarely occurring events.
  • Ignoring that strict synchronous replication can reduce availability by blocking writes when the standby is unreachable.
  • Treating each database in isolation and missing that a business capability's real objective is set by its weakest store.
  • Leaving a measured gap between declared and achieved objectives undeclared, so the organisation believes in a guarantee it does not have.

context