What is the difference between a planned switchover and an unplanned failover in a primary/replica relational database setup, and how does each affect the risk of losing committed data?
answer
- switchover: drain, catch up, swap, lossless
- failover: primary gone, promote furthest replica
- loss window = commit acked but not shipped
- failover leaves a diverged old primary
- rehearse with switchovers
basics
~20 sA switchover is a controlled role swap: writes are stopped, the replica catches up fully, then roles change with no data loss and the old primary can rejoin as a replica. A failover is reactive after a crash, so unreplicated commits can be lost.
solid answer
~50 sA **switchover** is planned. You quiesce writes on the primary, wait until the chosen replica has replayed everything, then demote the old primary and promote the replica. Because the old primary participates, the handover is lossless and it can immediately follow the new primary, so the operation is symmetric and reversible. Typical uses: OS or engine patching, hardware moves, rebalancing after an earlier incident. A **failover** is unplanned. The primary is gone or unreachable, so nobody can flush the remaining changes. You promote the most advanced replica and accept that anything committed but not yet shipped is lost, which is the recovery point objective. Extra work follows: the other replicas must be repointed at the new primary, and the old primary, when it returns, is potentially diverged and must be rewound or rebuilt. The practical implication is that you should rehearse the machinery with switchovers regularly, but never assume the two have the same cost: a switchover costs a short write pause, a failover costs data and cleanup.
go deeper
Give the two definitions and the key contrast: planned and lossless because the replica catches up first, versus reactive with a possible loss window.
Walk the switchover sequence, explain where the failover loss window comes from, and mention repointing the remaining replicas.
Add the aftermath: diverged old primary, replicas ahead of the new primary, client retry behaviour, and using scheduled switchovers as rehearsal.
Frame it as an RTO/RPO policy question: what loss is acceptable, what is automated, and how often the machinery is exercised.
## The two operations Both end with a different node accepting writes, but they differ in whether the outgoing primary cooperates. **Switchover (planned role swap).** The rough sequence: 1. Stop or drain application writes, or put the primary into a state where no new transactions start. 2. Wait for the target replica's replay position to equal the primary's flush position, i.e. lag reaches zero. 3. Cleanly shut down or demote the primary so its final changes are shipped and durable on the target. 4. Promote the target; it becomes writable. 5. Repoint the remaining replicas and the old primary at the new primary. 6. Repoint client traffic through the proxy or service layer. Because step 2 completes before step 4, no committed transaction is lost. The cost is a brief write outage, usually seconds, and it must be brief enough that connection pools and retry logic tolerate it. **Failover (reactive promotion).** The primary crashed, its host died, or it became unreachable. You cannot drain or wait for catch-up because the source of truth is gone. The manager or operator picks the replica with the furthest replay position, promotes it, and repoints everything else. Anything the dead primary had committed locally but not yet shipped is lost. ## Where data loss comes from Under asynchronous replication a transaction commits locally, returns success to the client, and only then are its changes streamed. The window between local durability and remote receipt is the exposure. If the primary dies inside that window, those commits exist only on its disk, so promotion loses them; if the disk survives you may be able to recover them later by hand, but the cluster has already moved on without them. Switchover eliminates the window by construction. Failover does not; it merely bounds it by how small you kept the lag and which promotion policy you use. ## What else differs operationally - **Other replicas.** In a switchover, they can be repointed calmly. After a failover, any replica that had received changes the new primary never got is ahead of it and cannot simply follow; it must be rewound or rebuilt. - **The old node.** After a switchover it rejoins as a replica immediately. After a failover it may have committed transactions the new primary does not have, so it has diverged and needs a rewind or full rebuild. - **Client impact.** A switchover can be sequenced with the routing layer, so clients see a short pause. A failover means in-flight transactions are aborted, connections are broken, and applications must reconnect and retry. - **Sequences and identity.** In both cases the new primary continues issuing surrogate keys, but after a failover an id range may have been consumed by lost transactions, so gaps appear. Gaps are normal; duplicates would mean split-brain. ## Confusion to avoid Candidates often use the words interchangeably or claim failover is lossless because replication is enabled. It is lossless only when the replica had every change at the moment of failure, which asynchronous replication cannot promise. Synchronous commit narrows this but changes write latency and introduces its own availability question when the standby is down. Similarly, a switchover is not merely a failover you happened to schedule. The difference is that the old primary is alive and cooperating, which is what allows the catch-up step and makes the old node reusable afterwards. ## Interview framing Define both in one sentence each, name the catch-up step as the reason a switchover is lossless, state that a failover's loss window comes from asynchronous replication, and mention the follow-up work: repointing other replicas and dealing with the returning old primary. Add that switchovers are the standard way to rehearse the machinery without waiting for an outage.
- How do you choose which replica to promote during a failover?Pick the one with the furthest applied position (highest LSN or most complete relay log), because it minimises lost transactions, subject to it being healthy and not disqualified by policy, for example a delayed replica or one in the wrong site. Automated managers compare positions from cluster state and often refuse candidates lagging beyond a configured threshold. If two are equal, secondary criteria like site locality or hardware class decide.
- Can a switchover ever lose data?It should not, if the procedure waits for the target replica to reach the primary's final flush position before promotion and the old primary is shut down cleanly. It can lose data if someone forces promotion without confirming catch-up, or if writes keep arriving at the old primary through a route that was not drained, which is really an unfenced failover wearing a switchover's name.
saying these in an interview costs you the question
- Using the words switchover and failover interchangeably
- Claiming failover is lossless simply because replication is configured
- Forgetting that the other replicas must be repointed at the new primary
- Assuming the old primary can always rejoin by just starting it as a replica
- Never practising switchover and expecting failover to work first time in an outage