skip to content

Your recovery time objective was signed when the ledger held a tenth of today's data, so what silently changed and what did not?

level: seniorimportance: nice to knowfreq 30%

answer

  1. the promise is a constant, the procedure is not
  2. the copy term dominates once data is large
  3. fixed steps stay fixed
  4. capture frequency sets the recovery point
  5. measure with a date and a size

basics

~20 s

The restore's real duration grew roughly with the data, so the recovery time objective is now missed although nobody edited it. The recovery point is usually unchanged, because it follows capture frequency rather than volume.

solid answer

~50 s

An objective is a promise; a restore is a procedure whose duration is a function of how much data has to be copied and replayed. Grow the dataset tenfold and the copy term grows with it, while the fixed steps — provisioning, repointing, verifying — stay roughly flat, so the total moved without anyone touching the document. The **recovery point** usually did not move, because it is set by how often state is captured and by replication lag, not by volume; the exception is when the capture itself starts taking longer than the interval between captures, which is worth checking. Nothing signals this drift, because the number is only tested when it is used or when somebody deliberately measures it. That is why a stated recovery time without a dated measurement behind it decays quietly.

code

pseudocode · 20 lines
pseudocode
// restore duration, decomposed
function estimateRestoreDuration(datasetSize, writeRate, distanceToTarget):
    provisionFixed   = 0.5          // time units, independent of data
    copyTerm         = datasetSize / restoreThroughput
    replayTerm       = (writeRate * distanceToTarget) / replayThroughput
    cutoverFixed     = 0.5          // repoint name, distribute credentials, verify
    return provisionFixed + copyTerm + replayTerm + cutoverFixed

// when the objective was signed: datasetSize = 1 unit
// restoreThroughput = 1 unit per time unit, replayTerm ~ 0.2
// => 0.5 + 1.0 + 0.2 + 0.5 = 2.2  -> comfortably inside a 4 unit objective

// today: datasetSize = 10 units, everything else unchanged
// => 0.5 + 10.0 + 0.2 + 0.5 = 11.2 -> the objective was never edited

// the recovery point, by contrast, is not a function of datasetSize
function worstCaseRecoveryPoint(captureInterval, captureDuration):
    if captureDuration >= captureInterval:
        return captureDuration      // captures now overlap, the gap stretches
    return captureInterval

go deeper

for a junior

Take away that a restore takes longer when there is more data, so a recovery time written down years ago may no longer be true even though nobody changed it.

for a middle

Decompose the procedure into fixed steps and size-dependent steps, and say which of the two recovery numbers each one affects.

for a senior

Show that you record measurements with a date and a dataset size, re-measure on growth rather than on the calendar, and attack the dominant term when the number stops fitting.

for a principal

Own the governance angle: a documented objective with no dated measurement behind it is an unfunded promise, and partial recovery of the recent slice is often the only honest way to keep a short number.

## An objective does not move; a procedure does A recovery time objective is a sentence in a document. The thing it describes is a procedure with real steps, and those steps have durations that depend on the system as it is today, not as it was when the sentence was written. Nothing connects the two, so the document stays accurate-looking while the procedure drifts out from under it. Decompose the procedure and the drift becomes predictable rather than mysterious: - **Provisioning the replacement store** — roughly fixed; it depends on the platform, not on your data. - **Copying the base state** — grows close to linearly with the amount of data, bounded by restore throughput. - **Replaying the change record** to the chosen moment — grows with the write rate and with the distance from the base capture to the target moment, not with total size. - **Cutover** — repointing the name, distributing credentials, restarting clients: roughly fixed, but usually the least measured. - **Rebuilding derived state** — caches and indexes, which grow with data size again. - **Verification and reconciliation** — grows with how much work was accepted after the chosen moment. ## What grows, and what does not | Ingredient | Scales with dataset size? | Notes | |---|---|---| | Provisioning | no | a platform constant | | Base copy | yes, close to linearly | the dominant term once data is large | | Change replay | no — scales with write rate and distance | can dominate if the target is far from the capture | | Cutover steps | no | fixed, and usually unmeasured | | Index and cache rebuild | yes | frequently forgotten entirely | | Recovery point | no | set by capture frequency and replication lag | The last row is the "what did not change" half of the question, and it is the part candidates usually miss. Capturing state every fifteen minutes still means at most fifteen minutes of work at risk whether the store holds a little or a lot. The honest exception: if the capture itself now runs longer than the interval between captures, the effective interval stretches and the recovery point does move — so tenfold growth is a reason to check the capture duration, not only the restore duration. ## Why the drift is silent Three properties of recovery work conspire here: 1. **The number is only tested when it is used.** Ordinary operation produces no signal about restore duration at all, unlike latency or error rate, which are watched continuously. 2. **The document has no input.** Nothing in the system reads the dataset size and re-derives the promise, so the objective is a constant in a world of growing variables. 3. **The last real measurement ages invisibly.** "We tested this" carries no date in most runbooks, and a measurement taken against a tenth of today's data is not evidence about today. The combination is why an unrehearsed procedure is a wish rather than a plan: not because people are careless, but because there is no feedback path from the system back into the promise. ## What to do about it 1. Record the measurement with its **date and the dataset size it was taken at**, so the row can be marked stale automatically rather than by memory. 2. Model the duration rather than quoting a single figure — a stated formula with a copy term, a replay term and a fixed term lets anyone re-estimate from today's size without running anything. 3. Re-measure when the data crosses a threshold, not on a calendar alone; growth, not time, is what invalidates the number. 4. Attack the dominant term when the number stops fitting. If the copy term dominates, the answers are keeping a second copy already running so no bulk copy is needed, splitting the data so recovery is partial and parallel, or renegotiating the objective. Optimising the cutover when the copy takes hours changes nothing. 5. Separate the promise you can meet for **part** of the service from the one you can meet for all of it. Restoring the recent slice of a ledger first, and the deep history afterwards, is often the only way a large store meets a short objective at all. ## The interview point The answer that lands has two halves stated explicitly: the recovery time silently degraded because its dominant term scales with data, and the recovery point did not, because it is set by capture frequency instead. A candidate who says "both got worse" has not understood which mechanism sets which number — and one who says "nothing changed, the objective is still four hours" has confused a promise with a capability.

  • Which part of the restore would you attack first once the number stops fitting?
    Whichever term dominates at today's size, which is almost always the bulk copy. The real answers are keeping a second copy already running so no bulk copy is needed, partitioning the data so the recent slice is recovered first and deep history follows, or renegotiating the objective honestly. Streamlining the cutover is worthwhile but cannot pay for a copy measured in hours.
  • Under what circumstance does dataset growth actually move the recovery point as well?
    When the capture itself begins to take longer than the interval between captures. The effective gap then stretches to the capture duration rather than the configured interval, so the worst-case loss window grows even though the schedule was never changed. It is worth checking capture duration alongside restore duration after any large growth.

saying these in an interview costs you the question

  • Assumes a stated objective stays true as data grows
  • Claims both recovery numbers degrade with dataset size
  • Cites a measurement without the date or the size behind it
  • Counts only the copy step in the restore duration
  • Optimises the cutover while the bulk copy dominates
  • Treats a documented objective as evidence of capability