For a cross-region failover of a checkout stack, would you automate the decision or require a human to declare it, and why?
answer
- cost of being wrong decides it
- cannot see it is not is gone
- observers outside both failure domains
- fence before you promote
- declared, but with the trigger pre-agreed
basics
~20 sAutomate inside a region, where a synchronous standby loses nothing and the move is cheap; require a declaration across regions, where promoting a lagging copy accepts real data loss and the failure signal is often ambiguous. Automation is only safe with independent observers and fencing.
solid answer
~50 sThe two decisions differ because their costs differ. Within a region, a standby that is already current can be promoted with no data loss, the failure signal is relatively clear, and the alternative is minutes of downtime waiting for a person — so automate it. Across regions, the copy is behind by its lag, so promoting it throws away confirmed orders, callers must be repointed, and failing back afterwards is a planned migration rather than an undo. Worse, you often cannot tell 'the region is gone' from 'I lost my view of it', which is exactly how both sides end up writable at once. If you do automate a cross-region promotion, it needs observers outside both regions, a majority rule so a partitioned side stands down, fencing of the old primary, and a check that the current lag is inside what the business has agreed to lose.
code
pseudocode · 23 lineson primaryUnreachableFor(observedSeconds) where observedSeconds >= failoverDelay:
agreeing = countObserversReportingUnreachable()
if agreeing < observerMajority:
alert "only one vantage point lost sight of the primary; not promoting"
return
if not fence(oldPrimary): // revoke its write credential and route
alert "old primary could not be fenced; escalate to the on-call decider"
return
lag = replicationLagAtLastAcknowledgedWrite(standby)
if lag > acceptedDataLoss:
alert "promotion would discard more than the agreed window; escalate"
return
if promotedWithin(cooldownWindow):
alert "recent promotion; refusing to flap"
return
promote(standby)
repointClients(standby)
record("promoted", lagDiscarded = lag)go deeper
Take away that failing over across regions is not free: the far copy is behind, so promoting it loses recent work. That is why somebody usually has to decide rather than letting it happen automatically.
Explain why 'I cannot reach the primary' and 'the primary is gone' are different statements, and what happens when a system confuses them — two writable copies and data that has to be reconciled by hand.
Name the controls that make automation safe: observers outside both domains, a majority rule, fencing, a lag gate, a cooldown. Say explicitly what you do when fencing cannot be confirmed.
Own the policy: how much confirmed data may be discarded, who holds the authority on each shift, and what evidence would convince you to move the decision from a person to a machine.
## Two failovers, two different bets The question is not 'is automation good'. It is **what does a wrong decision cost, and how confident is the signal**. | | Within a region | Across regions | |---|---|---| | Data loss on promotion | none, with a synchronous standby | whatever the lag held | | Cost of a false trigger | small; promote back later | large; divergent data and a planned failback | | Signal quality | usually unambiguous | frequently ambiguous | | Reversibility | close to reversible | a second migration | | Usual answer | automate it | declare it, with automation-assisted execution | Within a region the arithmetic favours automation easily. A human takes minutes to be paged, to log in, to confirm and to act; an automated promotion takes seconds; and if it was wrong, nothing was lost. Across regions the same arithmetic reverses: acting wrongly discards confirmed customer work, so the value of a few extra minutes of certainty is high. ## Why the cross-region signal is ambiguous A system deciding to fail over is really deciding *'the other side is gone'*, but what it can actually observe is *'I cannot see the other side'*. Those are the same observation for two very different worlds: - the primary region really is unavailable, and promoting is correct; or - the network between the observer and the primary is broken while the primary happily keeps serving customers, and promoting creates a second writable copy. The second case is **split brain**: two sides accept writes, both are internally consistent, both are wrong about the world, and reconciling them afterwards is manual work over customer data. This is why the decision is so often kept with a person for cross-region moves — a human can look at signals the automation cannot, including whether customers are actually complaining. ## What makes an automatic promotion safe If you do automate it, these are the parts that turn it from a hazard into a control: 1. **Observers outside both failure domains.** A watcher living in the standby's region is not neutral; it sees exactly the partition that would mislead it. Use several vantage points and require agreement. 2. **A majority rule.** With an odd number of participants or an external tie-breaker, a partitioned minority stands down rather than promoting itself. 3. **Fencing.** Before the new primary accepts a write, the old one must be prevented from accepting any — by revoking its credential, withdrawing its route, or stopping it. If fencing cannot be confirmed, the safe action is to stop and escalate, not to promote anyway. 4. **A lag gate.** Compare the standby's current lag with the data loss the business has agreed to accept. If it is worse, a person should be making that call. 5. **Hysteresis.** Require the failure to persist for a defined interval, and refuse to fail over again within a cooldown, so a flapping signal does not move the writable copy repeatedly. ## What a declared failover needs to be quick 'A human decides' is only a good answer when the human can decide fast. That requires three things prepared in advance: - **A pre-agreed trigger**, written as a threshold and a duration rather than a feeling, so the responder is confirming a condition rather than inventing a policy at three in the morning. - **One named decision-maker** for the shift, with an explicit deputy, so nobody spends the outage looking for authority. - **A procedure that does not depend on the failed side** — not a document stored only in the failing region, not a step that needs a credential issued there, and not an instruction that assumes someone can log into it. A procedure that has never been exercised is a plan, not a capability; teams routinely find that the first real attempt takes several times as long as the number they published. ## Failing back is its own decision Once the original region returns, the promoted side holds writes the old primary never saw. Going back means replicating forwards to the old side, catching up, and cutting over at a chosen quiet moment — with the same fencing question in the other direction. Because this is a deliberate migration rather than an emergency, it is usually scheduled, and the common mistake is to treat it as an undo and rush it into the tail of the incident. ## How to answer this out loud Say that the blast radius and the reversibility decide it, not a preference for automation. Automate the cheap, lossless, in-region move; put a person on the expensive, lossy, hard-to-reverse one, and make that person fast by pre-agreeing the trigger and the authority. Then name fencing, because a candidate who talks about failover without mentioning how the old primary is stopped has not run one.
- Where should the observers that decide a failover actually run?Outside both the primary's and the standby's failure domains, and in more than one place. An observer inside the standby's region sees precisely the partition that would mislead it, so it reports 'primary unreachable' for a network fault as readily as for a real outage. Several independent vantage points with a majority rule are what make the signal trustworthy.
- Why is failing back usually riskier than the failover was?Because the promoted side has accepted writes the original never saw, so returning means replicating in the new direction, waiting for catch-up, and cutting over deliberately — with the same fencing question reversed. It is a planned migration, best scheduled at a quiet moment, and the common error is rushing it as though it were an undo at the end of the incident.
- If a person declares the failover, what should already be decided before the page arrives?The trigger, as a threshold plus a duration; who holds the decision on that shift and who deputises; how much data loss is acceptable; and a procedure that needs nothing from the failed side. The responder should be confirming a condition, not inventing policy during the outage.
saying these in an interview costs you the question
- Treating a cross-region failover as the same decision as a zone-level one
- Promoting on a single observer's view that the primary is unreachable
- Promoting without first stopping the old primary from accepting writes
- Assuming a failback is just the failover run backwards
- Believing a documented procedure that has never been run will hold its stated time
- Leaving the decision to a person without pre-agreeing the trigger or the authority