A point-in-time restore of a payments ledger finished and the rows are correct, yet nothing works — what did the restore not bring back?
answer
- rows are not a service
- the fresh instance sits at a new address
- secrets, queues, caches, registrations
- cutover is the unmeasured half
- writes after the chosen moment need reconciling
basics
~20 sA restore returns rows, not a running service. The name clients resolve still points at the old endpoint, the restored instance has its own credentials, and in-flight messages, caches, derived indexes and registrations elsewhere are untouched by it.
solid answer
~50 sA restore is a data operation, and the recovery time objective covers the whole service. What comes back is a fresh store holding the chosen state; what does not come back is everything that pointed at the old one. The name clients resolve still resolves to the previous endpoint, so nothing reaches the restored data until somebody repoints it. The new instance usually carries its own credentials, so secrets held by callers no longer match. Messages that were in flight when the failure hit are not in the store and never were. Caches and derived indexes still hold pre-restore state, and anything registered with a third party against the old endpoint still names it. Rebuilding the surrounding environment itself is a separate discipline; this is the checklist that stands between a finished restore and a working service.
go deeper
Remember that a restore returns data, not a running service, and that something still has to point clients at the new copy before anyone notices an improvement.
List the categories that do not come back — endpoint name, credentials, in-flight messages, caches and derived indexes, external registrations — and explain why each one blocks traffic.
Show that you count verification, cutover, cache rebuild and reconciliation inside the recovery time, and that you have a plan for the work accepted after the chosen moment.
Push the design change upstream: indirection at the endpoint, secret distribution that survives an instance swap, and a rule that any step never performed end to end does not count as capability.
## The restore finished; the outage did not The recovery time objective is a promise about a **service**, and a restore is an operation on **data**. The gap between those two is where real recoveries lose hours, and it is the part almost never written down, because the person who tested the restore tested the copy and stopped there. What a successful restore hands you is a fresh store holding the state as of the moment you chose. Everything else in the system still refers to the arrangement that existed before the failure. ## What a restore typically does not bring back - **The address clients use.** The restored store is normally a new instance at a new endpoint. The name clients resolve still resolves to the old one, so until somebody repoints it — or repoints the layer in front — traffic keeps arriving at a dead or damaged store. This single item is the most common cause of "the restore worked but the service is still down". - **Credentials and secrets.** A new instance usually has its own credentials, and the values held by callers no longer match. Recovery therefore includes distributing new secrets to every caller, which is a deployment, not a database operation. - **In-flight work.** Messages sitting in queues, requests mid-flight, and work accepted but not yet committed were never in the store. Some are lost, some will be redelivered, and some will be redelivered against data that has been rewound — so the same request may be processed twice, or against a state it has already acted on. - **Caches and derived state.** Caches, search indexes and materialised aggregates still hold pre-restore values. Serving from them after a rewind is how corrected data looks corrupted again, so they need invalidation or rebuilding, and rebuilding takes time that belongs inside the recovery time. - **Registrations elsewhere.** Callback targets, allow-listed addresses, monitoring targets and scheduled jobs that named the old endpoint still name it. Each one is a small item and together they are the long tail of the cutover. - **Instance-level configuration.** A restored instance may come up with default settings rather than the tuned ones — connection ceilings, parameters, extensions — so it can be correct and still too small to carry production load. - **Everything written to the damaged store after the chosen moment.** This is the reconciliation problem: the ledger kept taking payments after the bad write, and rewinding discards those too. Somebody has to extract and replay them, and until that is decided the restored store is not authoritative. Rebuilding the surrounding environment itself — recreating the machines, networks and configuration around the data from source — is a separate discipline with its own owner; the list above is what stands between a completed restore and a service that answers. ## Why the checklist is the recovery time Put the two halves side by side and the shape of a real recovery becomes obvious: | Phase | What happens | Usually measured? | |---|---|---| | Copy and replay | data rebuilt to the chosen moment | yes — this is the number people quote | | Verify | someone confirms the state is the right one | rarely | | Cut over | name repointed, secrets distributed, clients restarted | almost never | | Rebuild derived state | caches invalidated, indexes rebuilt | almost never | | Reconcile | work accepted after the chosen moment replayed | almost never | The quoted recovery time is normally the first row, and the real one is the sum of all five. That is why a team can report a two-hour restore and a six-hour outage in the same post-mortem without anyone lying. ## What a good answer proposes 1. Write the recovery procedure as the whole path to serving traffic, not as the restore command, with the cutover steps named individually. 2. Make the endpoint indirection explicit — clients should reach the store through a name you control and can repoint, so cutover is one deliberate change rather than a redeployment of every caller. 3. Decide in advance what happens to in-flight and post-moment work, because that decision is a business decision about money and cannot be made calmly during the incident. 4. Treat cache and index rebuild time as part of the recovery time, and measure it. 5. Record which steps have ever been performed end to end. A procedure that has never been run through the cutover is a wish, not a plan, and the parts nobody has run are exactly the parts that will fail. Interviewers ask this because it separates candidates who have restored a database in an exercise from candidates who have recovered a service in anger.
- Why do the payments accepted after the bad write make the restore harder rather than easier?Because rewinding to a moment before the bad write also discards them. They are legitimate money movements that exist only in the damaged store, so somebody has to extract them, decide which are still valid, and replay them into the restored ledger. Until that decision is made the restored store is correct but not authoritative, and that reconciliation usually dominates the outage.
- What single design change most shortens the cutover after a restore?Having clients reach the store through an indirection you control — a name or a connection endpoint that can be repointed centrally — rather than through an address baked into each caller's configuration. Then cutover is one deliberate change plus a verification, instead of a coordinated redeployment of every service that talks to the store.
saying these in an interview costs you the question
- Declares the incident over when the restore command returns
- Assumes clients automatically find the restored store
- Forgets that the restored instance has its own credentials
- Serves from caches and indexes holding pre-restore state
- Never plans for work accepted after the chosen moment
- Quotes the copy duration as the recovery time of the service