skip to content

After a switch onto a standby cluster, what dominates a stream's recovery time, and why is it not the standby starting up?

level: middleimportance: should knowfreq 62%

answer

  1. the standby is already running
  2. the clock is the reader side
  3. positions rarely cross the hop
  4. count the pile that accumulated
  5. decision latency is part of the number

basics

~20 s

Readers dominate the recovery time, not the cluster. A standby fed by an ongoing copy is already running; the clock is spent repointing readers and writers, re-establishing where each reader resumes, and working off everything that piled up while nothing was consuming.

solid answer

~50 s

The recovery time is the span from the moment the source cluster is gone to the moment readers are current again on the **standby cluster** — and a standby fed by an ongoing copy is already up, so its start-up is rarely on the clock at all. What is on the clock is the reader side: the human decision to switch, repointing writers and readers at the other cluster, re-establishing where each reader group resumes there, and then burning down whatever accumulated while nobody was consuming. That last stage is the one teams routinely leave out of the number, and it is usually the largest whenever the decision and the switch took real minutes. A recovery time written as "how long the standby takes to accept connections" is a number about a cluster, not about a stream.

go deeper

for a junior

Remember that the second cluster is already running and already holds the records, so it is not what you are waiting for. What takes the time is getting the readers working again on it.

for a middle

Walk the stages: decide, repoint writers and readers, re-establish where each reader resumes, then work off the pile that built up. Say which stage dominates and why it is the last one.

for a senior

Show the arithmetic: inbound rate times outage span gives the pile, drain rate gives the minutes, and the ratio decides whether a fifteen-minute promise survives contact. Name the parallelism ceiling your platform shape imposes on that drain.

for a principal

Decide what the organisation is buying. Automating the declaration cuts the largest stage but risks switching on a transient loss; per-stream numbers cost more to maintain than one estate number but are the only ones anybody can honour.

## What the recovery-time clock actually measures A **recovery time** is the second of the two numbers you state about a stream: how long until the stream is doing its job again. The only useful definition puts both ends of the clock where the business feels them — it starts when the source cluster stops serving, and it stops when readers on the **standby cluster** are current, not when something accepts a connection. That definition matters because it moves the expensive part of the number out of infrastructure and into the reader side. ## Why the standby starting up is rarely the clock A standby cluster in this arrangement is not a cold thing you boot on the day. It is a running cluster, carrying no writers, kept fed by an ongoing **cross-cluster copier**. It already holds the records, already has its storage attached, already has its own in-cluster copies. Pointing traffic at it is a configuration act, not a provisioning act. (A cluster that is only *built* after the loss is a different and much longer story, and a plan that assumes one should say so.) So the honest clock is made of the stages either side of that: 1. **Detection and decision.** Somebody has to conclude the source cluster is not coming back soon and that switching is better than waiting. On most plans this is a human, and it is measured in minutes that no runbook shortens. 2. **Repointing writers and readers.** Every client has to be told to use the other cluster, and every client has to be allowed to — credentials and access rules on the standby are part of this, and are a common place the clock stalls. 3. **Re-establishing where each reader resumes.** This is the stage with the most variance across platforms, and it is discussed below. 4. **Burning down what accumulated.** While the decision and the switch were happening, writers were failing or buffering and readers were not consuming. Once everything is pointed at the standby, there is a pile to work through before the stream is *current*. This is the stage most often missing from the stated number. ## Where platforms differ on stage 3 | Platform shape | What has to be re-established | What that costs on the clock | |---|---|---| | Readers own a rewindable stored read position | Each reader group's resume point on the target cluster, since position numbers there are the target's own | Real minutes, plus whatever replay or gap the chosen resume point implies | | Records are removed on acknowledgement, consumers compete | Nothing to translate — readers simply begin taking what is present | Reconnection and authorisation dominate instead | | Position is expressed as a timestamp rather than a number | A time to resume from, which travels across clusters more naturally than a number | Cheaper, but only as accurate as the clocks involved | The common thread is that the *records* crossed the hop and the *bookkeeping around them* usually did not. Whatever your platform's answer is, the time it takes belongs in the number. ## Budgeting the number honestly A recovery time that is genuinely defensible is built bottom-up: - Write down each stage above with a measured duration, not an estimate. - Include the decision latency. If the plan says a human declares the switch and that human is on a pager, the pager response time is part of the recovery time. - Size the catch-up stage from arithmetic, not hope: the inbound rate times the outage span gives the pile, and the readers' drain rate gives the minutes. If readers can only consume at roughly the rate producers write, a ten-minute outage takes far longer than ten minutes to clear. - Count the stages you do not own. If producers buffer and then release everything at once, the catch-up burden is larger than the outage; if they drop, it is smaller but you have lost more than the uncopied tail. - State the number per stream. A stream whose readers are a handful of stateless workers recovers on a different clock from one whose reader has to rebuild derived state. ## The classic ways the number is wrong - **Measuring the cluster instead of the stream.** "The standby answers in ninety seconds" is true and irrelevant. - **Omitting the catch-up.** The most common single omission, and the one that turns a fifteen-minute promise into two hours in practice. - **Assuming parallelism you do not have.** On platforms that split a stream into a fixed number of parts, catch-up speed is capped by that count; on platforms where consumers simply compete for work, you can often add readers to drain faster. These are genuinely different recovery curves and the plan should say which one it is on. - **Ignoring the decision.** An automated switch and a human-declared one produce recovery times that differ by more than any technical stage in the list.

  • Why is the catch-up stage so often missing from a stated recovery time?
    Because it is the only stage that is not a task someone performs — it happens on its own after the visible work is done, so a runbook timed end to end never sees it. It is also the stage that scales with the outage: the longer the decision took, the bigger the pile, which is exactly the opposite of how people assume a recovery time behaves.
  • Does a faster cross-cluster copier improve the recovery time as well as the recovery point?
    Only marginally. A lower copy lag means less of the tail was lost and slightly fewer records are waiting, but the recovery time is dominated by the decision, the repointing and the catch-up — none of which the copier touches. The two numbers respond to different investments, which is why they are stated separately.
  • How does the recovery time change if the plan requires a human to declare the switch?
    It gains the whole detection-and-decision span: time to alert, time for a human to respond, and time to be confident the source cluster is not returning. That is routinely the largest single stage. Teams accept it anyway, because an automatic switch onto a standby can fire on a transient loss and create the split-brain problem it was meant to avoid.

saying these in an interview costs you the question

  • Defines recovery time as how long the standby takes to accept connections
  • Leaves the catch-up burn-down out of the stated number
  • Assumes reader bookkeeping crosses the hop with the records
  • Ignores the human decision latency before any switch begins
  • Assumes readers can always be scaled out to drain faster