Your plan states a recovery point and a recovery time for every stream but no switch has ever been rehearsed, so what are those numbers worth and how would you make them real?
answer
- a number is not a capability
- the rehearsal is the instrument
- stated beside measured, per stream
- last-rehearsed date is the signal
- synchronous writes fix only one number
basics
~20 sAn unrehearsed target is a stated intention, not a capability. Only a real switch measures the stages nobody budgeted — access on the standby, re-establishing reader resume points, the decision latency, the catch-up — so rehearse it, record the measured numbers beside the stated ones, and revise whichever is wrong.
solid answer
~50 sUntil a switch has been performed, the two numbers describe what somebody hoped, and the gap between hope and reality on this subject is reliably large. A rehearsal is the only instrument that measures the parts a document cannot: whether credentials and access rules exist on the **standby cluster**, whether reader groups can be re-established there at all, how long a human actually takes to declare the switch, whether copy lag on a bad day resembles the ceiling you published, and how long catch-up really runs. The fix is procedural rather than technical: rehearse on a schedule, keep a per-stream register holding the stated ceiling *and* what the last rehearsal measured, treat the last-rehearsed date as the health signal, and where the measured number is worse, either invest or change the published number. A target you have never tested is a number you will discover is wrong at the worst possible moment.
code
yaml · 20 lines- stream: order-capture
owner: payments
stated:
ceiling_records_lost: 50000
ceiling_seconds_of_copy_lag: 30
ceiling_minutes_until_readers_current: 15
last_rehearsed_switch: 2026-04-11
measured_on_that_rehearsal:
seconds_of_copy_lag_at_switch: 47
minutes_until_readers_current: 96
notes: readers could not be re-established for 22 minutes; access rules absent on standby
- stream: device-telemetry
owner: platform
stated:
ceiling_records_lost: 5000000
ceiling_seconds_of_copy_lag: 300
ceiling_minutes_until_readers_current: 120
last_rehearsed_switch: null
measured_on_that_rehearsal: nullgo deeper
Remember the distinction: a recovery target written down is what someone intends, and only performing a switch shows whether it can be met. Ask when the last one was run.
Explain what a rehearsal measures that a document cannot — access on the standby, re-establishing reader resume points, the decision latency, the catch-up — and why each is invisible until tried.
Show the register: stated ceilings beside measured results, per stream, with a last-rehearsed date. Say what you do when they disagree, and be specific about which stage you would fund first.
Argue the posture across the estate. Make the case for accepting a defended loss window over a synchronous cross-site write, name which few streams earn the expensive treatment, and own the cadence that keeps the numbers from ageing.
## A target is a claim; a rehearsal is the evidence Two numbers get written into a continuity plan for a stream: the **recovery point** — the stated ceiling on how much you may lose — and the **recovery time** — the stated ceiling on how long until readers are working again. Both are usually derived from what the business would like rather than from anything observed. That is fine as a starting point and useless as a commitment. The question a principal is being asked here is not "how do I lower the numbers". It is: *what converts a pair of numbers in a document into something the organisation can actually do?* The answer is a rehearsed switch, plus the discipline of writing down what it measured. ## What a rehearsal measures that a document cannot Each of these is a stage that is invisible until someone tries it, and each has taken a real plan apart: - **Whether clients can connect at all.** Credentials and access rules are part of the bookkeeping that a copy hop does not carry. A standby holding every record and admitting no one is the most common failure a first rehearsal finds. - **Whether reader groups can be re-established.** Where readers own a stored read position, that position is meaningful only on the cluster that issued it, and how you resume on the other side is a decision — replay some records, or skip some. A plan that has not made the decision discovers it live. - **How long the decision takes.** If a human declares the switch, the clock includes detection, paging, and the minutes spent being sure the source cluster is not returning. This is routinely the largest stage and is never in the document. - **What copy lag looks like on a bad day.** A published ceiling of thirty seconds is worth nothing if the measured copy lag between the two clusters during a real disturbance is two minutes. Rehearsing under load is the only way to see it. - **How long catch-up runs.** The pile that built up during the decision has to be worked off, and the drain rate is a property of the readers, not the cluster. - **Whether the companion services came along.** Anything a producer cannot publish without has to be reachable from the other site too, or the switch completes and nothing flows. ## Per-stream targets, not one estate number A single organisation-wide pair of numbers is easy to publish and impossible to honour, because what a loss costs differs enormously per stream. A register makes the difference visible and gives each owning team something to accept or refuse: | Column | Why it is there | |---|---| | Stated ceiling on records lost | The promise, in the unit the owning team understands | | Stated ceiling on seconds of copy lag | The same promise in the unit an operator can observe | | Stated ceiling on time until readers are current | The recovery time, ending at a working stream | | Date of the last rehearsed switch | The health signal — a stale date is the finding | | What that rehearsal measured | The only honest evidence either number is achievable | The register's value is in the last two columns. A gap between stated and measured is either a budget request or an admission that the published number should change, and forcing that choice in daylight is most of the work. ## The case for accepting a stated loss window The instinct when the measured numbers look bad is to buy a synchronous cross-site write and promise zero loss. Usually that is the wrong purchase: 1. It puts a cross-site round trip inside every write, on every stream that uses it, permanently — a cost paid every day for an event that may never happen. 2. It creates a new failure mode: when the far site is unreachable, you must either refuse writes or silently drop back to asynchronous, which restores the loss window without telling anyone. 3. It solves only one of the two numbers. The recovery time — decision, repointing, re-establishing readers, catch-up — is untouched, and that is generally the number the business actually feels. The mature posture is usually the opposite: state a loss window you can defend, prove it by rehearsal, make sure the streams that genuinely cannot afford it are the few that get the expensive treatment, and spend the rest of the money on shortening the recovery time. ## Keeping the numbers honest over time A rehearsal ages. Volumes grow, streams are added to the hop, readers get slower, and a number proved eighteen months ago is a number about a smaller system. Treat the register as living: rehearse on a fixed cadence rather than when someone remembers, re-measure after any material change to the estate, and review the stated ceilings with owning teams at the same cadence. When a rehearsal cannot be run against production, run it against a stream that matters and accept a narrower result — a partial rehearsal that produced a real measurement beats a complete plan that produced none.
- How do you rehearse a switch when the business will not accept downtime on the production estate?Rehearse the stages that do not require it: prove clients can authenticate and read on the standby, prove reader groups can be re-established there, and time the decision path with a real page. Then rehearse a full switch on one genuinely important stream rather than a toy one. A narrow real measurement is worth more than a broad plan with none.
- A rehearsal measured a recovery time six times the stated ceiling. What do you do with the number?Change one of them, in the open. Either fund the specific stage that consumed the time — usually access on the standby or reader re-establishment — and re-measure, or publish the measured figure as the new ceiling and let the owning team decide whether that is acceptable. Leaving an unachievable number in the document is the one option that is not available.
- Which companion dependency most often invalidates an otherwise complete rehearsal?The contract store: producers cannot publish without the contracts and identifiers it hands out, so if it is reachable only from the lost site, a switch completes with a healthy cluster and no traffic. Anything on that critical path belongs in the rehearsal explicitly, not as an assumption.
saying these in an interview costs you the question
- Treats an untested number in a document as the capability itself
- Publishes one recovery target for an entire estate of streams
- Reaches for a synchronous cross-site write before measuring anything
- Assumes a rehearsal proved eighteen months ago still holds
- Leaves the human decision latency out of the stated recovery time
- Keeps an unachievable published number rather than revising it