An auditor demands a recovery point and a recovery time per service, so how would you set different numbers for a payments ledger and for the reporting copy derived from it?
answer
- consequence first, configuration last
- originals against derivatives
- what could rebuild this, and how fast
- the curve steepens near zero
- a number without a measurement is a guess
basics
~20 sDerive each number from the consequence of losing that service, not from what the platform already does. A ledger holds originals nothing can reconstruct, so it buys a tight recovery point; a reporting copy is derived data and can be rebuilt, so it buys a loose one cheaply.
solid answer
~50 sThe question to ask of each service is what is destroyed if it is lost, and what can rebuild it. A payments ledger holds original records — money movements that exist nowhere else — so its recovery point must be small, which means a closely following standby plus restore points with a retention window long enough to cover realistic detection of a bad write. A reporting copy is **derived**: it can be regenerated from the ledger, so its own recovery point barely matters and its recovery time is whatever regeneration takes. The long-term archive is different again — recovery point is effectively irrelevant because it is append-only, while retrieval time may be hours by design. Then set the numbers in restore order, because a downstream service cannot beat the recovery time of what it depends on, and write each row with an owner, a date and the measurement behind it.
go deeper
Take away the core split: original data that nothing can reconstruct needs tight numbers, and data that can be rebuilt from it does not.
Be able to justify a number per service by naming what would rebuild that service and roughly how long the rebuild takes.
Show the ordering constraint — a downstream service cannot beat its dependency's recovery time — and insist on a measurement behind every quoted number.
Frame it as a portfolio with a steep cost curve near zero, name who accepts each number, and defend spending the budget on the few rows that cannot be rebuilt.
## Start from consequence, not from capability The failure mode of an audit exercise is a table filled in from what the platform currently happens to do. The number is then perfectly accurate and completely meaningless, because it describes configuration rather than tolerance. The discipline is to ask two questions of each service before looking at any setting: 1. **If this service's data is lost, what is destroyed?** Money that cannot be reconstructed, a regulatory obligation, a dashboard, a cache. 2. **What could rebuild it, and how long would that take?** Another system, a recomputation, a re-ingestion, or nothing at all. The second question does most of the work, because it splits an estate into originals and derivatives, and those two classes deserve wildly different numbers. ## The three rows in this system | Service | What it holds | Recovery point | Recovery time | Why | |---|---|---|---|---| | Payments ledger | original records nothing can reconstruct | minutes at most | short — the business stops without it | loss is permanent and externally visible | | Reporting copy | data derived from the ledger | loose — it can be regenerated | as long as regeneration takes | the ledger is the source of truth | | Long-term archive | append-only retained records | effectively irrelevant | can be hours by design | nothing is written that a rebuild could not re-send | The interesting row is the middle one. A team that copies the ledger's numbers onto the reporting copy "for consistency" has bought continuous capture and a second running copy for data it could rebuild by rerunning a job. That is the single most common way a recovery budget is spent on the wrong service. ## The cost curve is the argument Both numbers get expensive in the same shape: cheap to improve while they are large, sharply more expensive as they approach zero. - Moving a recovery point from a day to an hour is usually a change of capture frequency and retention — a bill, but a small one. - Moving it from an hour to seconds normally requires a continuously running second copy, which is a second store to pay for, patch, monitor and fail over. - Moving a recovery time from a day to hours is procedure work: writing the runbook, removing the manual steps, controlling the endpoint indirection. - Moving it from hours to minutes normally requires capacity already running and a cutover that is automatic, which is a different architecture and a standing cost. Stating that curve is what lets a leader say no. "We can give the reporting copy the ledger's numbers; it costs a second running store forever and saves a rebuild job that takes twenty minutes" is a sentence that ends the conversation honestly. ## Recovery is ordered, so the numbers are not independent A service cannot be recovered before what it depends on. If the reporting copy is regenerated from the ledger, its achievable recovery time is the ledger's recovery time **plus** regeneration — promising a shorter one on paper is arithmetic that has already failed. Two practical rules follow: 1. Write the recovery order first, then the numbers, and check each downstream row against the sum of everything above it. 2. Distinguish what must be back before the business resumes from what may return afterwards. Declaring the whole estate equally critical produces a plan nobody can execute under pressure, because everything is first. ## What the auditor is actually asking for A defensible row has five fields, and the last one is the one that is usually missing: - The service, named the way the business names it. - The recovery point and the recovery time, as numbers. - The named owner who accepted them. - The date they were agreed, and when they are next reviewed. - **The measurement**: when the procedure was last performed end to end, and what it actually took. Without the last field the table is a statement of intent. A procedure nobody has ever run is a wish, not a plan, and the parts that have never been run are reliably the parts that fail — the cutover, the secret distribution, the reconciliation of work accepted after the chosen moment. A number with a measurement behind it is a commitment; a number without one is a guess that has been typed into a document and signed. ## What to say when pressed The principled position is that recovery numbers are a portfolio decision: a small number of services carry tight numbers and the standing cost that buys them, most carry loose ones, and every row is justified by what cannot be rebuilt rather than by how important the service feels. Then commit to measuring the few tight rows on a schedule, because those are the only ones whose numbers anybody will ever be held to.
- A team asks for the ledger's recovery numbers on every service in its estate. How do you respond?By pricing it rather than refusing it. Show the standing cost of a continuously running second copy and of continuous capture per service, set against what each one could be rebuilt from and how long that rebuild takes. Uniform tight numbers also dilute attention: when everything is first in the recovery order, nothing is, and the rows that genuinely matter stop being measured.
- How do the numbers change for the long-term archive a regulated business keeps?Its recovery point is close to irrelevant because it is append-only and anything missing can be re-sent from the source. What matters instead is retention length, protection against deletion, and retrieval time — which may legitimately be hours, since nobody serves traffic from it. Writing a tight recovery time there is a common way to buy speed no one needs.
saying these in an interview costs you the question
- Fills the table in from what the platform currently does
- Gives every service the strongest numbers for consistency
- Sets a derived copy's recovery point tighter than its source's
- Promises a downstream recovery time shorter than its dependency's
- Records objectives with no owner, no date and no measurement
- Treats the archive's retrieval time as an availability problem