skip to content

The source cluster stopped answering ten minutes ago — what do you want to establish before calling the switch to the standby?

level: seniorimportance: should knowfreq 52%

answer

  1. unusable, or only unreachable?
  2. how stale is the standby right now
  3. can producers be kept off the source
  4. readers need a defensible start position
  5. the switch is not an undo

basics

~20 s

Establish that the cluster is genuinely unusable rather than merely unreachable, how far behind the standby is now, that producers can be kept off the source, and that every reader group has a defensible start position on the standby.

solid answer

~40 s

Four things. First, is the source cluster down or merely unreachable from where I am standing? A path failure means producers elsewhere may still be writing into it, and switching then creates two writing sites. Second, how far behind is the standby right now — because that distance is the records I am choosing to leave behind. Third, can I actually stop the source from accepting writes, or at least be confident nothing is reaching it? Fourth, does every reader group have a start position on the target that I can defend, and has anyone ever done it? The switch is a deliberate call because it is not an undo: once traffic runs on the standby, coming back is a larger operation than going.

go deeper

for a junior

Recall that switching is a deliberate call, not an automatic reaction, and that the first thing to establish is whether the cluster is really down or just unreachable from where you are looking.

for a middle

Explain the inputs: current staleness of the standby, whether producers can be kept off the source, and whether reader groups have a start position on the target. Say why each one changes the answer.

for a senior

Demonstrate that you weigh the switch's own cost against how long the source is likely to stay unusable, that you insist on a second vantage point, and that you treat copier health as evidence rather than background.

for a principal

Argue for the decision being written before the event: named owner, stated conditions, per-stream restart rules, and an accepted answer for the records stranded on the original cluster.

## The call is the interesting part, not the mechanics The mechanics of a switch are short: repoint producers, restart reader groups on the standby, keep one site writing. The judgement is in deciding to do it at all, because the switch trades a known cost — the records the copier had not carried, plus the reprocessing every reader group does at restart — against an unknown one, how long the source cluster stays unusable. ## The four things to establish 1. **Unusable, or unreachable?** This is the question that most often gets skipped. A broken path between your operators and the cluster looks identical, from a dashboard, to a cluster that has stopped. If it is a path failure, producers in other places may still be writing happily into the source. Switching in that state gives you two sites taking writes and two histories that will not reconcile. Seek evidence from more than one vantage point, and prefer evidence produced by the cluster itself over evidence produced by something watching it. 2. **How far behind is the standby right now?** The **copy lag** at the moment of the switch is the size of what you are agreeing to leave behind. It also tells you something else: if the copier stopped an hour before the cluster did, the standby is an hour stale and the switch is a much worse trade than you think. The health of the copier is part of the evidence, not a background detail. 3. **Can writes to the source be stopped?** Either the cluster is genuinely unreachable to everyone, or you have a way to keep producers off it — network, credentials, or the deployment that points them elsewhere. If the answer is "we think nothing is writing to it", that is not the same answer. 4. **Does every reader group have a start position you can defend on the target?** Restarting readers is the part of the switch that takes the clock. Knowing in advance whether a **position map** exists, or whether you will restart from a timestamp, or whether some groups simply begin at the oldest record they can find, is the difference between a forty-minute switch and a four-hour one. ## What is not a reason to switch | observation | why it is not the signal | |---|---| | one unhealthy indicator on a dashboard | indicators lie in both directions; confirm with a second, independent view | | producers are erroring | they may be failing for their own reasons, or against a path, not against the cluster | | the stream contents are wrong | the copier carried the wrong records across as well | | a reader group is far behind | that is a reader problem and follows the group onto the standby | | the cluster is slow | a slow cluster usually recovers sooner than a switch completes | That last row deserves weight. A switch has a floor of tens of minutes once every producing service and every reader group is counted. If the honest estimate for the source coming back is under that floor, switching costs more than waiting. ## Why it stays a human call with a stated cost The switch is not symmetric and it is not an undo. The moment writers land on the standby, that cluster holds records the original does not, so the return is a second operation with its own preparation. It also spends something that cannot be unspent: the records still sitting on the original cluster are stranded until that cluster is back, and if it is gone for good they are gone with it. That is why the useful form of the decision is written down before the incident: a short list of conditions, an owner who is allowed to say yes, and a per-stream note of what readers do at restart. During the event you want to be **confirming a prepared decision**, not designing one. The questions above are that list. ## Where platforms differ - Where the reader owns a rewindable numeric position, point 4 is a choice about where to land and can be made quickly with a prepared rule. - On **destructive-read** designs, where acknowledged messages are removed and there is no position to rewind, point 4 becomes "which messages had been carried", and anything unacknowledged on the stranded cluster generally has to be regenerated by its producer rather than recovered. - Where the cluster is rented, the mechanics may be one action in a console, but the four questions above are unchanged — and the vantage-point problem is often worse, because you are reading the provider's view of its own health.

  • Why is an unreachable source cluster more dangerous to switch on than a confirmed dead one?
    Because unreachable is only a statement about your vantage point. If the cluster is alive behind a broken path, producers that can still see it keep writing while you start writing to the standby. You then hold two histories with no way to merge them, and every record written to the stranded side is either abandoned or re-injected out of order.
  • The copier stopped two hours before the cluster failed. How does that change the call?
    Substantially. The standby is two hours stale, so the switch now discards two hours of records rather than seconds. That may still be right if the cluster is gone, but it moves the decision from routine to expensive, and it usually justifies spending longer trying to recover the source or to drain whatever remains reachable on it first.
  • Who should be allowed to make this call?
    One named role, decided before the incident, with the conditions written down and a per-stream note of how readers restart. The value is not seniority — it is that someone is confirming a prepared decision instead of designing one under pressure, and that nobody is waiting for a consensus that will not form.

saying these in an interview costs you the question

  • Switches as soon as one monitoring indicator turns unhealthy
  • Treats unreachable from my laptop as the cluster is gone
  • Assumes the switch can simply be reversed later
  • Never checks how stale the standby is before landing on it
  • Expects the switch to fix wrong or poisoned records
  • Switches for a slow cluster that would have recovered sooner