skip to content

Cross-Cluster Continuity

Keeping something worth switching to when a cluster is gone: an asynchronous copy, a standby readers restart against, and what a restore cannot return. Asked because the switch is never rehearsed.

part ofBroker & streaming operationsoverview, primer and where to startread it →
on this pageshow

questions

22

A standby cluster is kept fed by an ongoing copy and carries no writers — what happens when an operator switches onto it?

level: juniorimportance: must knowfreq 58%

answer

  1. a second cluster, no writers on it
  2. writers move, readers restart
  3. records crossed, bookkeeping did not
  4. asynchronous copy, so a tail is missing
  5. one site accepting writes at a time

basics

~20 s

Writers are pointed at the standby, readers restarted there, and only one site keeps accepting writes. The standby holds only what the copier had carried, and a reader's stored position from the source names a different record there.

solid answer

~40 s

A standby cluster is a second cluster at another site, kept current by a cross-cluster copier that reads streams from the source cluster and writes the same records into the target cluster. Switching onto it is three separate movements: point producers at the standby, restart each reader group there at a position that means something on that cluster, and make sure the source stops accepting writes so only one site is live. Because the copier is an ordinary client of both clusters, the copy is asynchronous — records the copier had not yet carried are simply not on the standby. The records that did cross arrive without the bookkeeping around them, so a reader group's stored read position from the source does not identify the same record on the target.

go deeper

for a junior

Recall the three moves: producers repointed, reader groups restarted, one site accepting writes. And recall why the standby is slightly behind — the copy is asynchronous, so the newest records have not arrived.

for a middle

Explain why a reader group cannot simply carry on: its stored read position belongs to the source cluster's own numbering, and the copier moves records rather than that bookkeeping. Be able to say what is on the standby at the moment of the switch.

for a senior

Show that you treat the switch as a deliberate call with a cost, that you know the reader restart dominates the clock, and that you know a standby answers a lost cluster rather than bad contents.

for a principal

Frame what the organisation is really buying: a second running cluster, a rehearsed restart procedure for every reader group, and the acceptance that the return trip is a bigger operation than the switch itself.

## What a standby cluster is A **standby cluster** is a complete second cluster, normally at another **site**, kept fed by a **cross-cluster copier**: a process that continuously reads the streams you care about from the **source cluster** and writes the same records into the **target cluster**. It runs the same kind of broker and holds the same stream names, but no producer writes to it while it is a standby, and usually nothing reads from it either. It exists for exactly one purpose — to be somewhere to go when the cluster you run is gone, wrong or unreachable. The copy is **asynchronous by construction**. The copier is an ordinary client of both ends: it reads a record from one cluster, writes it into the other, and only then moves on. So the newest records on the source are always the ones that have not arrived yet. That distance is **copy lag**, and while it may be small, you cannot assume it is zero at the moment you need it. ## The three moves the switch makes "The switch" is not a button. It is three movements, and each can fail on its own. 1. **Writers are pointed at the standby.** Producers must resolve a different address set and reconnect. Whether that is a configuration push, a redeploy or a name change, it takes real minutes and it touches every producing service, not just the ones you remember. 2. **Readers are restarted on the standby.** This is the hard one, because a **reader group**'s **stored read position** is a number in the source cluster's own space. The copier moves records; it does not move the group's bookkeeping, and even where a position can be approximated on the target it lands on the nearest carried record rather than the identical one. 3. **Only one site keeps accepting writes.** If the source is merely unreachable from you and not actually down, producers elsewhere may still be writing into it. Two live sites is a different arrangement with its own problems; during a switch you want exactly one. | | before the switch | after the switch | |---|---|---| | producers | write to the source cluster | write to the standby | | the copier | source to standby | stopped, or later turned around | | reader groups | stored positions on the source | restarted against the standby | | newest records | on the source | whatever the copier had carried | ## What arrives, and what does not What the standby has is the records the copier managed to carry before the source became unusable. What it does not have is the tail the copier had not reached — those records exist only on a cluster you cannot currently use, and whether they ever come back depends on whether that cluster does. The second, less obvious gap is that a stream's records and a reader's progress through them are different objects. The copier is a client writing records; a group's stored read position is state the source cluster kept about its consumers. Records cross the hop; that state does not follow automatically, which is why "where does each reader start" is the question the switch really turns on. ## Why this is a decision, not a reflex The switch has a cost and it is not symmetric. Once traffic runs on the standby, that cluster becomes the only holder of the newest records, so going back is a second, larger operation rather than an undo. That asymmetry is the reason the call is made deliberately, on evidence that the source is genuinely unusable, rather than fired by a single unhealthy indicator. It is also worth being clear about what a standby protects against. It protects against **losing a cluster**. It does not protect against bad records: the copier faithfully carried whatever was written, so a stream poisoned on the source is poisoned on the standby too. ## Where platforms differ The shape of step 2 depends on the platform class: - On platforms that **split a stream into parts and let the reader own a numeric position**, restarting means choosing a position per part on the target, and there is a defensible answer either side of the true one. - On **destructive-read** designs, where a message is removed once acknowledged and there is no stored position to rewind, the equivalent question is which not-yet-acknowledged messages the copier had carried, and whether the ones still sitting on the unreachable cluster have to be regenerated. - Some **managed** offerings present a paired arrangement where the switch is one control-plane action; the movements above still happen, they are simply performed for you and the reader restart is still yours. The common core across all three: a second cluster fed asynchronously, a deliberate human call, and readers that have to be told where to begin.

  • Why does a standby carry no writers while it is a standby?
    Because the moment producers write to both clusters, it stops being a standby and becomes a two-accepting-site arrangement, which is a different design with its own hazards. Keeping writers on one side means there is a single authoritative order of records, and the switch is a clean handover of that role rather than a merge of two divergent histories.
  • Does switching to the standby help when the records on the source are wrong rather than missing?
    No. The copier carries whatever was written, so bad records reach the standby along with the good ones — usually within seconds. A standby answers "the cluster is unusable", not "the contents are wrong". Recovering from bad contents is a different exercise, and switching sites during one mostly wastes the switch.
  • How long does the switch actually take?
    Rarely the cluster part. The standby is already running, so the clock is spent on reconnecting every producing service, deciding and applying a start position for each reader group, and then waiting while those groups work through whatever they reprocess. The restart of the readers, not the start of the cluster, is what dominates.

It is like moving a shop to a second premises that has been receiving a daily delivery of the same stock. The shelves look right, but today's delivery has not arrived, and the stocktaking notes stayed at the old shop — so staff have to work out where they had got to before they can serve anyone.

saying these in an interview costs you the question

  • Thinks the standby is always exactly in step with the source
  • Assumes readers resume where they were because the records are present
  • Treats the switch as something an unhealthy indicator should trigger by itself
  • Believes switching back afterwards is simply the same operation in reverse
  • Expects a standby to protect against wrong or poisoned records
  • Assumes the two clusters number the same record identically
open as a page

Your only continuity plan for a live stream is last night's file backup of the broker data volumes. What has that already cost you by morning?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A file backup fixes a stream at the instant it was taken, so every record written since is gone, along with every reader's progress. Streams keep moving while files do not, which is why the gap is counted in hours.

open as a page

Why is a cross-cluster copy of a stream always behind the source cluster that feeds it?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A cross-cluster copier is an ordinary client of both clusters: the source stores a record and answers the writer before the copier has even read it. The record therefore exists on the source first and on the target some time later.

open as a page

Your stream is copied asynchronously to a second cluster — why can the recovery point you state for it never be smaller than the copy lag?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The asynchronous copy hop sets the floor. Records acknowledged on the source cluster but not yet carried to the target exist in one place only, so losing the source loses them — the recovery point is at least the copy lag.

open as a page

Two sites accept writes to a stream of the same name, with a cross-cluster copier running each way — what stops a record circulating forever?

level: middleimportance: must knowfreq 52%

basics

~20 s

An origin stamp: a marker on the record, or an origin-qualified stream name, that tells the copier running the other way this record came from elsewhere. Without one, each copier treats the other's output as new local traffic and the loop never ends.

open as a page

Why does a reader group's stored read position from the source cluster not identify the same record on the standby, and how does it resume there?

level: middleimportance: must knowfreq 62%

basics

~20 s

Each cluster numbers records itself, so the number names a different record on the other one. A group resumes from a position map the copier maintains, or from a timestamp; both land near the old place, not on it.

open as a page

A cluster's brokers are all healthy yet every producer's publish fails - which service beside the cluster do you suspect, and how do you confirm it?

level: middleimportance: must knowfreq 50%

basics

~20 s

Suspect a companion service on the publish path - most often the contract store a producer must reach before it may publish. Broker metrics cannot see it, so the cluster reads green while every write is refused.

open as a page

When a cross-cluster copier writes a record onto the target cluster, which properties of the record survive the hop?

level: middleimportance: must knowfreq 58%

basics

~20 s

A cross-cluster copier carries a record's payload and key unchanged, and preserves order only where one part of the source stream feeds one part of the target in a single stream of writes. The position number is assigned fresh by the target cluster.

open as a page

A broker cluster is restored from a file backup and the records are present, yet no reader group progresses and clients are refused. What did the restore not return?

level: seniorimportance: must knowfreq 53%

basics

~20 s

The bookkeeping around the records: stored read positions, the entries saying who may connect and act on which stream, per-stream settings that differed from the cluster defaults, and the contract store the payload identifiers resolve against. The records are the easy part.

open as a page

Your cluster already parks closed history in cheap object storage, and the continuity plan calls that the backup. What is that remote-storage copy genuinely for?

level: middleimportance: should knowfreq 41%

basics

~20 s

It exists to make long history affordable and readable through the cluster, not to be restored from. The parked objects are addressed through the cluster's own bookkeeping, so on their own they are opaque files with no reader positions, permissions or settings attached.

open as a page

A copied record gets a new position number on the target cluster, so what relates the two clusters' position spaces?

level: middleimportance: should knowfreq 48%

basics

~20 s

Only a position map written by the cross-cluster copier while it ran: sampled pairs of a source position and the nearest position on the target. Without one, a number recorded against the source names no particular record on the target.

open as a page

After a switch onto a standby cluster, what dominates a stream's recovery time, and why is it not the standby starting up?

level: middleimportance: should knowfreq 62%

basics

~20 s

Readers dominate the recovery time, not the cluster. A standby fed by an ongoing copy is already running; the clock is spent repointing readers and writers, re-establishing where each reader resumes, and working off everything that piled up while nothing was consuming.

open as a page

Both sites accepted writes to a stream of the same name during an hour-long network split. Why can no merge afterwards recover the true order across them?

level: seniorimportance: should knowfreq 45%

basics

~20 s

There was never one order to recover: each site sequenced only its own writes, and nothing anywhere recorded how the two interleaved. Any merge has to invent a rule — write timestamps, site priority — and different rules produce different, equally unfounded answers.

open as a page

The source cluster stopped answering ten minutes ago — what do you want to establish before calling the switch to the standby?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Establish that the cluster is genuinely unusable rather than merely unreachable, how far behind the standby is now, that producers can be kept off the source, and that every reader group has a defensible start position on the standby.

open as a page

After a day of traffic on the standby, why is returning to the original cluster harder than the switch onto the standby was?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The roles have reversed: the standby now holds the only copy of a day's records, while the original holds a stale history plus a leftover tail. Returning means running the copy the other way first, then translating reader positions again.

open as a page

Restarting a reader group on the standby means either reprocessing records it already handled or skipping some — how do you choose?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Choose per stream, by what a repeat and a miss each cost downstream. Default to landing early and reprocessing, since a duplicate is usually recoverable and an unread record is not, then make the overlap wider than the restart's uncertainty.

open as a page

Where does the contract store beside a cluster keep its own state, and why does that decide whether a restore of the brokers is usable?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A contract store keeps state of its own, and where it lives decides the restore: on the cluster it serves, in another team's database, or in a checked-in definition. Only the last comes back for free.

open as a page

Why can't copy lag between two clusters be measured by subtracting the target stream's newest position number from the source's?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The two clusters number their streams independently, so that subtraction is arithmetic on unrelated counters and is meaningless even when the hop is perfectly caught up. Measure inside the source's own space, or in seconds of record age.

open as a page

The source cluster is lost and its uncopied tail dies with it, so what does that gap cost the systems downstream?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Those records were acknowledged to their writers, so upstream believes the facts happened while every system derived from the stream is permanently short of them — and the gap is silent, because readers resuming on the target cluster see an unbroken stream with no hole in it.

open as a page

A team wants one shared stream writable at both sites — why do operators propose one stream per site instead, and what does that cost readers?

level: principalimportance: should knowfreq 38%

basics

~20 s

One stream written at two sites has no defined order across them and needs loop prevention on every hop just to exist. Giving each site its own stream with one owning writer, read everywhere, removes both problems and moves the combining work onto readers.

open as a page

Which services standing beside your clusters belong in the continuity plan, and what do you do about the ones you exclude?

level: principalimportance: should knowfreq 32%

basics

~20 s

Inventory everything a publish and a read touch besides the brokers, then give each companion a posture: one shared instance, a copy per site, or rebuild on demand. Anything excluded should be excluded on purpose, in writing.

open as a page

Your plan states a recovery point and a recovery time for every stream but no switch has ever been rehearsed, so what are those numbers worth and how would you make them real?

level: principalimportance: should knowfreq 45%

basics

~20 s

An unrehearsed target is a stated intention, not a capability. Only a real switch measures the stages nobody budgeted — access on the standby, re-establishing reader resume points, the decision latency, the catch-up — so rehearse it, record the measured numbers beside the stated ones, and revise whichever is wrong.

open as a page