A failover test passes. What does verifying failback to the primary catch that the failover leg does not?
answer
- Two switches, only one rehearsed
- The window creates its own state
- Which way do changes flow now
- Clients keep their old address for a while
- Someone edited a setting during the incident
basics
~20 sFailback catches everything the switch itself created: writes accepted on the standby that must reach the primary, a replication path that has to reverse, stale reads right after the return, clients still pinned to the old node, and configuration edited during the incident.
solid answer
~50 sFailover starts from a known state; failback starts from a state the incident invented, which is why it is the leg that loses data. Verifying it catches divergence created during the window — every write the standby accepted has to reach the primary before the primary serves reads — plus the reversal of the replication direction, promotion of a primary that has not caught up, stale reads where a caller re-reads its own write and gets an older version, clients whose pools and cached lookups still point at the previous node, and configuration hand-edited during the incident and never reverted. It also exposes cold performance: the recovered primary has empty caches and can miss its latency target under real load while passing a functional check. Assert with a marker stream running across both switches, comparing acknowledged records against stored ones in both directions.
code
pseudocode · 14 lineswriter.start(rate_per_second = 37) # each marker carries a sequence number
failover_to(standby)
wait(minutes = 26)
failback_to(primary)
writer.stop()
acked = writer.acknowledged_sequences()
stored = primary.stored_sequences()
assert acked - stored == empty # accepted, then lost
assert stored.duplicates() == empty # applied twice on reversal
assert primary.read(acked.max()) == writer.value_for(acked.max())
assert connection_census(old_node) == 0 # nobody still pinned
assert config_diff(primary, standby) == baseline_diffgo deeper
Know the two words and the order: failover moves traffic to the standby, failback returns it to the primary once it is healthy. Be ready to say that returning is a separate step that needs its own checks, not an automatic undo.
Explain the mechanics: writes taken during the window must reach the primary, replication has to reverse before promotion, and a caller can read its own write and get an older value. Be ready to name the assertions you would run rather than describing the switch narratively.
Show you would run a marker stream across both switches and reconcile the acknowledged set against the stored set, plus a connection census, a configuration diff and a latency check under load. Expect to be asked what you assert when both nodes accepted writes during a partition.
Own the tradeoff of when to return at all, how long a system may safely run on a standby, and how you fund a rehearsal of the return path across many services without pretending an untested failback is a neutral cost.
## Two switches, one of them rehearsed **Failover** is the move away from a failed primary onto a standby. **Failback** is the return to the primary once it is healthy again. Teams rehearse the first and improvise the second, and then discover that the second is the one that loses data — because failover starts from a known state and failback starts from a state that the incident invented. A failover test proves a genuinely useful but narrow set of things: that the switch is detected and triggered, that the standby holds enough data to serve, that clients survive the interruption, and that the feature works on the other side. Everything created *by* the switch is out of its scope. ## What only the failback leg can prove **Divergence created during the window.** While the standby was authoritative it accepted writes. Failing back means those writes must reach the primary before the primary serves reads, or they are lost. This is the single most common way a successful failover ends with missing data, and no failover assertion can see it, because at the moment failover is declared green the divergence has not been created yet. **Replication direction.** After failover the flow of changes has to reverse: the old primary must become a follower of the standby, catch up, and only then be promoted back. That reversal is often a manual, rarely used procedure, and a stale primary that is promoted before it has caught up silently rolls the system back in time. **Stale reads immediately after the switch.** Right after failback, a caller can read a value it wrote during the failover window and get an older version, because the recovered primary is behind or a read replica is behind it. The test must include a read-your-writes check across the switch, not merely a "the endpoint responds" check. **Client stickiness.** Connection pools, cached name lookups and long-lived sessions keep clients pointed at the previous node for as long as their own caches live. A census that counts connections still landing on the old node is a real assertion; an eyeball on a dashboard is not. **Configuration drift.** During the incident someone raised a limit, disabled a job, repointed a setting. Failback is when those hand edits must be reverted, and nobody has a list. Diff the two nodes' effective configuration before the drill and after it, and assert that any scheduled job runs in exactly one place — a job left enabled on both sides is a duplicate-processing bug waiting for the next quiet night. **Cold performance.** The recovered primary comes back with empty caches and empty pools. It may pass a functional check and still miss its latency target under real load for the first several minutes. Assert against the target with load applied, not against an idle smoke check. ## Split-brain Split-brain is the state in which **both** nodes accept writes and their data diverges — not a lagging standby, and not a standby serving reads. It is worth exercising deliberately during the switch: partition the link in both directions while a writer is running, and assert one of two outcomes, decided in advance. Either the system fences the loser so that exactly one node accepts writes, in which case the assertion is that writes to the fenced node are refused; or the system does allow both, in which case the assertion is that the divergence is **detected and surfaced** for reconciliation rather than resolved silently by whichever write happened last. A silent last-write-wins merge is the failure mode, because it destroys evidence. ## How to assert all of it Run a continuous marker stream across both switches: a writer emits sequence-numbered records at a steady rate, records exactly which ones were acknowledged, and keeps running through failover, the outage window and failback. Afterwards, three set comparisons decide the drill: - acknowledged minus stored must be empty — nothing was accepted and then lost; - stored must contain no duplicates — nothing was applied twice through the reversal; - the last acknowledged value must be readable from the primary immediately after failback. Add the connection census, the configuration diff, and a latency check under load. On one playlist-service exercise, that stream showed 4,118 edits accepted by the standby during a 26-minute window and 63 of them missing after failback — a functional check on either side of the switch would have reported success both times. ## The judgement call Failback is usually the more disruptive of the two switches, so teams postpone it and run for weeks on the standby. That is a defensible operational choice, but it must be a choice: an untested failback is an unbounded liability, and the sensible compromise is to rehearse the return path in a lower environment on a schedule and to keep the marker-stream assertions ready so that the real return is measured rather than hoped for.
- During the switch you partition the link so both nodes accept writes. What should the test assert?One of two outcomes, chosen in advance. Either the system fences the loser so exactly one node takes writes, and the assertion is that writes to the fenced node are refused with a specific error; or both are allowed to write, and the assertion is that the resulting divergence is detected and surfaced for reconciliation. The failure mode is a silent merge that keeps whichever write landed last, because it destroys the evidence that divergence happened.
- Why is the recovered primary often slower than the standby immediately after failback?It returns with cold caches, empty connection pools and unwarmed data, so the first minutes of real traffic run against the slow path. A functional check will pass throughout. Assert latency targets with production-shaped load applied after the switch, and consider warming or a staged traffic ramp so the return does not read as a second outage.
- How would you decide whether failback is worth rehearsing at all, given it is disruptive?Treat it as a choice rather than an omission. Rehearse the return path in a lower environment on a schedule, keep the reconciliation assertions ready so the real return is measured rather than hoped for, and record how long the system may stay on the standby. Indefinitely deferring the return is defensible only while someone owns the risk that the primary's path back has never been executed.
A detour is easy to drive into; the mess happens when the road reopens and everyone has to merge back.
saying these in an interview costs you the question
- Calls the test green as soon as traffic reaches the standby
- Assumes replication reverses itself automatically
- Never checks whether clients are still pinned to the old node
- Dismisses a stale read after the switch as normal eventual consistency
- Leaves configuration edited during the outage unreverted and unchecked
- Thinks split-brain just means the standby lagged behind