Would you enable fully automatic failover for a production OLTP relational database, and how would you set detection thresholds, promotion criteria, and failback policy? Justify the tradeoffs.
answer
- automate the common case, bound the worst case
- quorum + confirmed fence + max-lag rule
- thresholds from measured p99, not folklore
- rate-limit, cooldown, maintenance pause
- failback = planned switchover, never automatic
basics
~20 sUsually yes, because humans are slower than the outage budget, but only with quorum-gated promotion, confirmed fencing, a maximum-lag rule, hysteresis against flapping, and automated rejoin. Failback should be a separate, scheduled switchover, never automatic.
solid answer
~60 sI default to automatic failover, because most incidents are single-node failures and a paged human takes minutes while automation takes tens of seconds. But automation is only safe with preconditions: leadership decided by quorum, fencing confirmed before promotion, and a candidate chosen by replay position. Thresholds: detection long enough to survive p99 network jitter, garbage collection pauses, and checkpoint stalls, typically tens of seconds, tuned from measured distributions rather than guessed. Too short and you buy spurious failovers, each costing a connection storm, aborted transactions, and a rewind. I also require several consecutive failed checks from more than one observer, and a maximum acceptable replication lag for the candidate; beyond it, escalate to a human rather than silently discard minutes of data. Hysteresis matters: rate-limit failovers, refuse to promote a node that was recently demoted, and support a maintenance mode so patching does not trigger promotions. Failback is not automatic. The old primary rejoins as a replica via rewind or rebuild, and returning it to primary is a planned switchover during a quiet window. Finally, automatic failover only counts as working if it is rehearsed on a schedule and the application retries idempotently.
go deeper
Know that automatic failover exists, that it needs safeguards, and that a failover breaks client connections.
Explain the tradeoff between fast detection and spurious failovers, and name the basic safeguards: quorum, fencing, lag limits.
Derive thresholds from measured latency and stall distributions, cover anti-flapping and maintenance mode, and describe rejoin and rehearsal in operational detail.
Take a position and justify it as a bounded-risk argument: automation may cost availability but must never cause divergence, with failback planned and drills scheduled.
## The decision The question is what the system does in the ninety seconds after a primary stops responding at 03:00. Manual failover means a page, a wake-up, context, verification, and a runbook: realistically five to twenty minutes. Automatic means tens of seconds. For an OLTP system with a tight RTO, automation wins on the common case (a host, disk, or process failure) as long as it cannot cause the uncommon catastrophe (two primaries or an avoidable data loss). So I frame it as: automate the decision, but constrain it so the worst outcome it can produce is bounded unavailability, never divergence. ## Preconditions I require before enabling it 1. **Quorum-gated leadership.** Promotion requires a majority of a fixed member set, held in a consensus store, so a minority partition can never elect. 2. **Confirmed fencing.** The old primary is provably unable to write, and if that cannot be confirmed, the automation aborts and pages rather than promoting. 3. **Candidate selection by position**, with a maximum-lag rule: if the best candidate is behind by more than the configured threshold, escalate instead of promoting. 4. **Automated rejoin.** The returning old primary is checked for divergence and rewound or reinitialised automatically, never allowed to accept writes first. 5. **A routing layer that follows leadership**, so promotion actually redirects clients and stale routes cannot keep writing to the old node. 6. **Application-side retry** on connection loss with idempotent operations, because every failover aborts in-flight transactions. Without the first three I would not enable automation; the failure mode is worse than the outage it prevents. ## Setting the thresholds Detection is a tradeoff between false positives and RTO. Set it from data, not folklore: look at the p99 and p99.9 of health-check latency, network round trip, checkpoint and GC pause duration, and the consensus store's own latency. The detection budget must exceed the worst benign stall you actually observe, with margin. In lease-based systems this is the TTL, commonly around 30 seconds, with an agent loop several times faster. I require multiple consecutive failures and, where the tooling supports it, corroboration from more than one observer, so a single flaky monitor cannot trigger a promotion. I also prefer checks that exercise the real path (a trivial query on a real connection) over a TCP-port ping, which stays healthy while the database is unusable. Cost of getting it wrong in each direction: too aggressive means spurious failovers, and each one costs aborted transactions, a client reconnect storm, cache warm-up on the new primary, and a rewind of the demoted node. Too conservative means paying extra seconds on every real failure. In most OLTP shops the spurious-failover cost is higher than people expect, so I start conservative and tighten with evidence. ## Anti-flapping and operator control - Rate-limit: no more than one automatic failover per cluster per configured window; beyond it, freeze and page. - Cooldown on a recently demoted node before it is eligible again. - Maintenance/pause mode that suspends automation during patching, restores, or bulk loads, with an expiry so nobody leaves it off. - Loud, structured alerting on every automated action, including refusals, since a refusal is often the more important signal. ## Failback Failback should be manual and planned. After the incident the old primary rejoins as a replica; whether it should be primary again is a question of hardware, site, and load, and it is another write interruption. Doing it automatically doubles the disruption from one incident and can ping-pong if the underlying fault is intermittent. The right shape is a scheduled switchover during a low-traffic window, once the node has proven stable, and only if there is a real reason to prefer it. ## Proving it works Automation that has never been exercised is a hypothesis. I schedule regular switchovers in production (they are lossless) and periodic true failover drills in a production-like environment, measure the actual RTO and RPO, and treat every drill finding as a defect. I also verify the boring parts that break silently: fencing device credentials, replication slots on the new primary, backup jobs following the writer, and monitoring that survives the role change. ## Interview framing Give a clear answer (yes, with conditions), list the preconditions that make it safe, derive thresholds from measured distributions rather than defaults, cover flapping controls, state that failback is planned rather than automatic, and end on rehearsal and application retry behaviour.
- What is the actual cost of a spurious failover on a busy OLTP database?In-flight transactions abort and every client reconnects at once, which often hits the new primary as a thundering herd against a cold cache and empty plan caches, so latency degrades for minutes even after connectivity returns. The demoted node then needs a rewind or rebuild before it can serve again, temporarily reducing redundancy. If the trigger was transient, you have paid all of that for nothing, which is why detection thresholds should be derived from observed benign stalls.
- When would you deliberately keep failover manual?When the preconditions are missing, for example no quorum store, no reliable fencing path, or a two-site topology where promotion crosses sites and the loss window is large. Also for databases where a wrong promotion is more expensive than downtime, such as a system of record with strict zero-loss requirements, or where the application cannot retry safely. In those cases I would still automate detection and preparation, and leave only the final promotion decision to a human with a rehearsed runbook.
saying these in an interview costs you the question
- Enabling automatic failover without quorum or fencing because the tool supports it
- Choosing detection timeouts by copying a default instead of measuring benign stalls
- Configuring automatic failback so the cluster ping-pongs during an intermittent fault
- Ignoring that every failover aborts in-flight transactions and requires client retry
- Claiming automation works when it has never been rehearsed