skip to content

Your secret store's recovery procedure is written down but has never been run — what do you commit to as its owner, and what counts as a pass?

level: principalimportance: nice to knowfreq 26%

answer

  1. a written plan is an assumption
  2. evidence, not intent
  3. a pass is a served value, not a copy
  4. the runbook must not need the store
  5. measure the time, do not assert it

basics

~20 s

A written recovery procedure is an untested assumption; a rehearsal turns it into evidence. Commit to a cadence, a named owner and an isolated environment, and define a pass as a restored store serving a known value within the promised time.

solid answer

~40 s

Documented means someone believed it would work. Rehearsed means it did, at a measured time, with the people who would actually be on call. The commitment has four parts: a cadence tied to change rather than to the calendar, a named owner who may declare it failed, an isolated environment with its own consumers so a restored copy cannot serve production or issue into live systems, and a written pass condition. A pass is not a restored archive — it is a store that serves a known value to a real caller within the recovery time the estate was promised, plus the reconciliation list the restore made necessary. The first rehearsal almost always fails on the custody step or on a circular dependency, and that is its value.

go deeper

for a junior

Recall the distinction: a backup that has never been restored is a hope, not a capability. Someone has to have opened it and used it.

for a middle

Explain what a rehearsal tests beyond the copy: that the archive opens, that the material needed to unwrap it can be produced, and how long the whole sequence actually takes.

for a senior

Show the operating detail: an isolated environment so a restored copy cannot serve production or reach live downstream systems, a measured clock that includes the custody wait, and a written finding list.

for a principal

Own the commitment — cadence tied to change, a named owner who may declare failure, a pass condition written before the run, and a recovery number the dependent teams can plan against.

## Documented is an assumption; rehearsed is evidence A recovery procedure that has never been run records what someone believed would work on the day they wrote it. Between then and now the store changed version, the custody of its protecting key moved, two of the three named people left, and the estate that starts against it grew. None of those changes announce themselves in the document. A rehearsal is the only mechanism that converts the belief into evidence, and it is the only thing that produces a *number* for how long recovery takes. A green backup job is not that evidence. It proves bytes were written. It says nothing about whether they can be opened, by whom, or how quickly. ## What a first rehearsal reliably discovers - **Who can actually produce the unwrap material, out of hours.** The document names a role; the rehearsal finds out whether a human answers, how long approval takes, and whether the device or service involved is itself reachable. - **Whether the archive opens at all.** Format, compression, and the software able to read it all drift, and an archive nobody has opened is an untested file. - **How long the whole path takes**, measured rather than estimated — almost always dominated by the custody step and by the reconciliation, not by copying bytes. - **The circular dependency.** The credentials for the archive location, the access to the key custody, and the contact details for the people involved are frequently held in the store being recovered. That dependency is invisible while the store is up and fatal the moment it is not. - **The reconciliation nobody budgeted for** — the accounts issued after the restore point, the values rotated since, the callers registered since. ## Defining a pass before you run it Write the pass condition down first, or the rehearsal will be graded on relief. A usable definition has four clauses: 1. A store restored from a real archive of the same age as a real one would be, not a purpose-built copy. 2. Serving a **known value to a real caller** — an actual authenticated read, not a status page reporting healthy. 3. Inside the recovery time the estate was told to expect, with the clock running from the declaration, including the wait for custody. 4. With the reconciliation list produced: what the restore did not bring back, enumerated rather than assumed. Anything short of clause 2 is a file-restore drill, and a file-restore drill is what teams accidentally rehearse for years. ## Where it runs, and what that cannot prove The store is the dependency the rest of the estate starts against, so the rehearsal runs in an isolated environment with its own name and its own consumers. Two failures make that non-negotiable: a restored copy reachable by production callers can serve values that have since been rotated, and a restored store allowed to reach live downstream systems can issue or withdraw real accounts. Be honest about the limit. An isolated rehearsal does not prove the production cutover — the moment the restored store takes over from the failed one, with real traffic and real callers. That is why the runbook states the cutover separately, and why the rehearsal's report names it as untested rather than quietly implying it works. ## What the owner is signing up to - **A cadence tied to change**, not only to the calendar: after the custody arrangement changes, after the store's deployment shape changes, and on a floor interval regardless. - **A named owner** with the authority to declare a rehearsal failed. Without that, a rehearsal that half-worked is recorded as a pass. - **A record** of the measured time, what failed, and what was changed as a result — the only artefact that makes the next rehearsal cheaper. - **A standard for the teams that depend on it**, since a store's recovery time is an input to everyone else's: if the store takes four hours to come back, no service that fails closed on it can promise less. ## The number is negotiated, not technical How much recent change you accept losing and how long you accept being down are the two dials, and they are business positions, not properties of the tooling. More frequent archives shrink the first; warm capacity and pre-arranged custody shrink the second, both at a cost. The lead's job is to make the estate state which number it is buying and then to prove, by rehearsal, that the number is real. A recovery time asserted from what the tooling *could* do, and never measured end to end, is the specific failure this whole exercise exists to prevent.

  • What does a rehearsal establish that a green nightly backup job does not?
    That the archive opens, that whoever must supply the material to unwrap it is reachable and correct, that the restored store serves a real read, how long the whole path takes end to end, and how much reconciliation follows. The backup job proves only that bytes were written somewhere.
  • Where do you run the rehearsal, given the store is the estate's start-up dependency?
    In an isolated environment with its own name and its own consumers, so a restored copy cannot be reached by production callers or issue into live downstream systems. The one thing that cannot be proven there is the production cutover, which is why the runbook records it as untested rather than implying it works.

saying these in an interview costs you the question

  • Counts a successful backup job as proof that recovery works
  • Defines the rehearsal as restoring the archive, with nothing served afterwards
  • Keeps the recovery runbook and its access behind the store being recovered
  • Sets a recovery time from what the tooling could do rather than measuring it
  • Rehearses once at build time and never again as the estate changes