After a day of traffic on the standby, why is returning to the original cluster harder than the switch onto the standby was?
answer
- the roles have reversed
- the standby now holds the newest records
- the copy must run the other way first
- a stranded tail that went nowhere
- often the right answer is do not return
basics
~20 sThe roles have reversed: the standby now holds the only copy of a day's records, while the original holds a stale history plus a leftover tail. Returning means running the copy the other way first, then translating reader positions again.
solid answer
~50 sFailing over moves you onto a cluster that was being kept current for exactly that purpose. Failing back has no such preparation. The standby is now the only holder of a day of new records, so before anyone can return, a copy hop must be created in the opposite direction and allowed to carry that day across — which takes real time and must be caught up before the return, not during it. Meanwhile the original cluster still holds its old history plus the records that were written just before it failed and never reached anywhere, so you must decide whether to abandon that tail or re-inject it, knowing re-injection places those records at the end of the order rather than where they belong. And every reader group needs its positions translated again, in the other direction.
go deeper
Recall that after traffic runs on the standby, that cluster holds records the original has never seen, so going back is a new operation rather than undoing the first one.
Explain the required sequence: create a copy hop in the opposite direction, let it catch up fully, then move producers and translate reader positions again — and why doing those out of order reintroduces loss.
Demonstrate judgment about the stranded tail on the original and about the original's leftover history, and be able to argue that not failing back is frequently the correct answer.
Decide the posture in advance: whether the estate is designed for two interchangeable sites or one live and one spare, since that single choice determines whether returning is a routine swap or a bespoke project each time.
## The asymmetry The switch onto a standby works because someone prepared for it: a **cross-cluster copier** had been running for months, the target held the streams, and the only novelty was where readers should start. **Failing back** reverses every one of those conditions. After a day of production traffic on the standby: - the **standby is now the only cluster holding that day's records**, and the original has none of them; - the original holds an **older history of the same streams**, plus whatever was written into it in the minutes before it became unusable and never left — a tail that exists nowhere else; - no copier is running in the direction you now need; - the two clusters' numbering diverged further while they ran apart. So the return is not an undo. It is a fresh cross-cluster arrangement, built in the opposite direction, with the added complication that the destination is not empty. ## What the return actually requires 1. **Turn the copy around.** A hop from the standby back to the original has to be created and allowed to run until it has carried everything the day produced. This is elapsed time proportional to the volume, and it must finish *before* the return, not during it — returning while the original is still behind reintroduces exactly the loss you were trying to avoid. 2. **Decide what happens to the leftover tail on the original.** Those records were written, acknowledged to their producers, and never copied anywhere. Three answers exist and all are uncomfortable: abandon them; re-inject them, which appends them after a day of newer records rather than in their original place; or reconcile them out of band, outside the stream, which is usually what actually happens for anything financially meaningful. 3. **Deal with the original's old history.** The original's streams still contain the records from before the failure, and the copier is about to write copies of some of the same records back into it. Whether you empty and reseed those streams or accept a doubled span is a decision to make explicitly, because nothing in the mechanism makes it for you. 4. **Translate reader positions again**, now from the standby's numbering into the original's, with the same early-or-late choice as before — and this time the whole exercise is voluntary, so a mistake has no excuse. 5. **Repoint every producer a second time**, with the same reach and the same minutes as the first move. ## The conclusion this usually leads to Weighed honestly, the most common good answer is **do not fail back**. If the standby can carry production indefinitely, the cheaper move is to keep running there and rebuild the *standby role* on the original cluster — start a copier from the current live site to the old one and leave it as the new standby. The roles swap permanently; nothing has to be merged; no second voluntary outage is spent. The cases that genuinely argue for returning are structural rather than sentimental: - the original site is where the data is legally required to live; - the standby is deliberately smaller or cheaper and cannot hold the load or the retention long-term; - the standby lacks the neighbouring services the workload depends on, so latency or cost is materially worse there; - the arrangement is licensed, contracted or budgeted as one live site and one spare. Where those apply, schedule the return as a **planned change** with the copy caught up first, a quiet period for writers, and the reader positions worked out beforehand — not as an emergency reversal. ## Where platforms differ - Where the reader owns a rewindable numeric position, the translation step exists in both directions and is symmetric in difficulty. - On **destructive-read** designs there are no positions to translate, but the leftover-tail problem is sharper: unacknowledged messages stranded on the original are either re-driven by their producers or lost, since there is no history to re-read them from. - Where clusters are rented, turning the hop around may be a control-plane action rather than a deployment, but the elapsed catch-up time and the decision about the stranded tail are unchanged. - Some arrangements are designed from the start so that either site can lead, with the copier's direction a setting rather than a rebuild. That removes step 1's engineering work; it does not remove the catch-up wait or the tail.
- What do you do with the records written to the original cluster just before it failed and never copied anywhere?Decide explicitly rather than by default. Re-injecting them puts them after a day of newer records, so any consumer sensitive to order sees them out of place; abandoning them loses acknowledged work. For anything financially meaningful the usual answer is neither — reconcile them outside the stream, against the downstream systems that should have seen them.
- When is keeping production on the standby permanently the better answer?Whenever the standby can carry the load and retention indefinitely and nothing structural requires the original site. You then rebuild the standby role in the other direction and swap the two roles for good. That avoids a second voluntary outage, a catch-up wait, and a second reader-position translation, none of which buy anything.
- Why must the reversed copy be caught up before the return rather than during it?Because returning while the original still trails means producers start writing there while a day of records is still in flight behind them. You reintroduce a loss window on a change you chose to make, and you create an interleaving no consumer expects. Catch up fully, pause writers briefly, then move.
It is like evacuating to a second warehouse for a day. Going out was easy — that warehouse had been receiving deliveries all along. Coming back means first shipping a day of accumulated stock the other way, and deciding what to do with the pallets left on the old loading dock that were never on anybody's manifest.
saying these in an interview costs you the question
- Describes failing back as simply switching in the opposite direction
- Assumes the original cluster still holds everything it needs
- Starts writing to the original before the reversed copy is caught up
- Ignores the records stranded on the original that were never copied out
- Re-injects the stranded tail without considering where it lands in the order
- Returns out of habit when the standby could carry production permanently