skip to content

A standby cluster is kept fed by an ongoing copy and carries no writers — what happens when an operator switches onto it?

level: juniorimportance: must knowfreq 58%

answer

  1. a second cluster, no writers on it
  2. writers move, readers restart
  3. records crossed, bookkeeping did not
  4. asynchronous copy, so a tail is missing
  5. one site accepting writes at a time

basics

~20 s

Writers are pointed at the standby, readers restarted there, and only one site keeps accepting writes. The standby holds only what the copier had carried, and a reader's stored position from the source names a different record there.

solid answer

~40 s

A standby cluster is a second cluster at another site, kept current by a cross-cluster copier that reads streams from the source cluster and writes the same records into the target cluster. Switching onto it is three separate movements: point producers at the standby, restart each reader group there at a position that means something on that cluster, and make sure the source stops accepting writes so only one site is live. Because the copier is an ordinary client of both clusters, the copy is asynchronous — records the copier had not yet carried are simply not on the standby. The records that did cross arrive without the bookkeeping around them, so a reader group's stored read position from the source does not identify the same record on the target.

go deeper

for a junior

Recall the three moves: producers repointed, reader groups restarted, one site accepting writes. And recall why the standby is slightly behind — the copy is asynchronous, so the newest records have not arrived.

for a middle

Explain why a reader group cannot simply carry on: its stored read position belongs to the source cluster's own numbering, and the copier moves records rather than that bookkeeping. Be able to say what is on the standby at the moment of the switch.

for a senior

Show that you treat the switch as a deliberate call with a cost, that you know the reader restart dominates the clock, and that you know a standby answers a lost cluster rather than bad contents.

for a principal

Frame what the organisation is really buying: a second running cluster, a rehearsed restart procedure for every reader group, and the acceptance that the return trip is a bigger operation than the switch itself.

## What a standby cluster is A **standby cluster** is a complete second cluster, normally at another **site**, kept fed by a **cross-cluster copier**: a process that continuously reads the streams you care about from the **source cluster** and writes the same records into the **target cluster**. It runs the same kind of broker and holds the same stream names, but no producer writes to it while it is a standby, and usually nothing reads from it either. It exists for exactly one purpose — to be somewhere to go when the cluster you run is gone, wrong or unreachable. The copy is **asynchronous by construction**. The copier is an ordinary client of both ends: it reads a record from one cluster, writes it into the other, and only then moves on. So the newest records on the source are always the ones that have not arrived yet. That distance is **copy lag**, and while it may be small, you cannot assume it is zero at the moment you need it. ## The three moves the switch makes "The switch" is not a button. It is three movements, and each can fail on its own. 1. **Writers are pointed at the standby.** Producers must resolve a different address set and reconnect. Whether that is a configuration push, a redeploy or a name change, it takes real minutes and it touches every producing service, not just the ones you remember. 2. **Readers are restarted on the standby.** This is the hard one, because a **reader group**'s **stored read position** is a number in the source cluster's own space. The copier moves records; it does not move the group's bookkeeping, and even where a position can be approximated on the target it lands on the nearest carried record rather than the identical one. 3. **Only one site keeps accepting writes.** If the source is merely unreachable from you and not actually down, producers elsewhere may still be writing into it. Two live sites is a different arrangement with its own problems; during a switch you want exactly one. | | before the switch | after the switch | |---|---|---| | producers | write to the source cluster | write to the standby | | the copier | source to standby | stopped, or later turned around | | reader groups | stored positions on the source | restarted against the standby | | newest records | on the source | whatever the copier had carried | ## What arrives, and what does not What the standby has is the records the copier managed to carry before the source became unusable. What it does not have is the tail the copier had not reached — those records exist only on a cluster you cannot currently use, and whether they ever come back depends on whether that cluster does. The second, less obvious gap is that a stream's records and a reader's progress through them are different objects. The copier is a client writing records; a group's stored read position is state the source cluster kept about its consumers. Records cross the hop; that state does not follow automatically, which is why "where does each reader start" is the question the switch really turns on. ## Why this is a decision, not a reflex The switch has a cost and it is not symmetric. Once traffic runs on the standby, that cluster becomes the only holder of the newest records, so going back is a second, larger operation rather than an undo. That asymmetry is the reason the call is made deliberately, on evidence that the source is genuinely unusable, rather than fired by a single unhealthy indicator. It is also worth being clear about what a standby protects against. It protects against **losing a cluster**. It does not protect against bad records: the copier faithfully carried whatever was written, so a stream poisoned on the source is poisoned on the standby too. ## Where platforms differ The shape of step 2 depends on the platform class: - On platforms that **split a stream into parts and let the reader own a numeric position**, restarting means choosing a position per part on the target, and there is a defensible answer either side of the true one. - On **destructive-read** designs, where a message is removed once acknowledged and there is no stored position to rewind, the equivalent question is which not-yet-acknowledged messages the copier had carried, and whether the ones still sitting on the unreachable cluster have to be regenerated. - Some **managed** offerings present a paired arrangement where the switch is one control-plane action; the movements above still happen, they are simply performed for you and the reader restart is still yours. The common core across all three: a second cluster fed asynchronously, a deliberate human call, and readers that have to be told where to begin.

  • Why does a standby carry no writers while it is a standby?
    Because the moment producers write to both clusters, it stops being a standby and becomes a two-accepting-site arrangement, which is a different design with its own hazards. Keeping writers on one side means there is a single authoritative order of records, and the switch is a clean handover of that role rather than a merge of two divergent histories.
  • Does switching to the standby help when the records on the source are wrong rather than missing?
    No. The copier carries whatever was written, so bad records reach the standby along with the good ones — usually within seconds. A standby answers "the cluster is unusable", not "the contents are wrong". Recovering from bad contents is a different exercise, and switching sites during one mostly wastes the switch.
  • How long does the switch actually take?
    Rarely the cluster part. The standby is already running, so the clock is spent on reconnecting every producing service, deciding and applying a start position for each reader group, and then waiting while those groups work through whatever they reprocess. The restart of the readers, not the start of the cluster, is what dominates.

It is like moving a shop to a second premises that has been receiving a daily delivery of the same stock. The shelves look right, but today's delivery has not arrived, and the stocktaking notes stayed at the old shop — so staff have to work out where they had got to before they can serve anyone.

saying these in an interview costs you the question

  • Thinks the standby is always exactly in step with the source
  • Assumes readers resume where they were because the records are present
  • Treats the switch as something an unhealthy indicator should trigger by itself
  • Believes switching back afterwards is simply the same operation in reverse
  • Expects a standby to protect against wrong or poisoned records
  • Assumes the two clusters number the same record identically