Beyond electing a new leader correctly, what concrete mechanisms stop a still-running, previously-elected leader from continuing to write to shared storage after a network partition has caused a new leader to be chosen, and why is preventing the election alone not enough?
answer
- election agreement != stopping the old leader's writes
- control plane (who's leader) vs data plane (who writes)
- fencing token: monotonic number, storage rejects lower tokens
- STONITH: forcibly kill/power-off the old node instead
- storage-enforced > cooperative self-stepping-down
basics
~20 sEven after everyone agrees on a new leader, the old leader may not know it's been replaced and could keep writing. Fix: storage refuses writes from anyone but the current leader, via a growing 'ticket number' that rejects old ones.
solid answer
~50 sElecting a new leader only changes what the rest of the cluster believes; it does nothing to physically stop the old leader's process from still issuing writes if it hasn't yet noticed it's been superseded, particularly under a partial partition where the old leader can't see the new leader or quorum but can still reach the shared storage. The real defense has to live at the point of shared-resource access, not just in the election protocol: fencing tokens are the standard mechanism, where each new leadership term gets a monotonically increasing number that must accompany every write, and the storage layer rejects any write carrying a token lower than the highest one already accepted. STONITH (Shoot The Other Node In The Head) is a blunter, infrastructure-level alternative used in some HA clustering software, forcibly powering off or fencing the old leader's machine rather than relying on cooperative token checks.
go deeper
Should grasp that just picking a new leader doesn't magically stop the old one from doing anything, in plain terms.
Should describe fencing tokens at a basic level: a growing number attached to writes, rejected if too old.
Should clearly separate the control-plane/data-plane distinction, explain why the check must live at the storage layer rather than be self-enforced by the leader, and name at least fencing tokens plus one alternative (STONITH).
Should reason about which mechanism (fencing tokens vs. STONITH) fits which operational context (owned storage vs. legacy/third-party resources), and cite concrete production examples of each pattern being enforced.
## The control plane and the data plane A subtle but critical gap exists in almost every leader-election protocol taught in isolation: the protocol guarantees that the cluster reaches agreement on who the current leader is, but it says nothing about forcibly stopping a previous leader that hasn't yet learned it's been replaced. Election algorithms like Bully, Ring, Raft's `RequestVote`, or ZooKeeper's ephemeral-znode recipe all operate over the **control plane** - the messages nodes exchange to agree on leadership - but a leader's actual dangerous action is on the **data plane**: writing to a shared database, appending to a replicated log, or granting an exclusive lock. A partition can very easily separate these two planes: the old leader might be cut off from the rest of the cluster (so it can't see the new leader's heartbeats or a quorum) while remaining fully connected to the shared storage system it writes to. From the old leader's own local point of view, nothing looks wrong yet - it just can't reach its peers, which it may interpret as everyone else being down rather than itself being isolated - so it keeps right on writing. ## Why a correct election is not a claim about writes This is precisely why 'we ran a correct election' is not the same claim as 'writes are safe.' Even Raft, which is provably safe about log replication among the nodes that participate in its own quorum, relies on the assumption that only quorum-participating writes matter; if some external system (a downstream cache, an external API the leader calls, a file the leader writes directly) is written to by the leader outside of Raft's own replicated log, Raft's internal safety guarantees don't automatically extend to that side effect - the leader still needs some way to prove, to that external system, that its leadership claim is current at the moment of the write. ## Fencing tokens move the check to the resource Fencing tokens solve this by moving the safety check to the one place that actually matters: the shared resource itself. 1. Every time leadership changes hands - whether via lease renewal, a new Raft term, a new ZooKeeper ephemeral znode, or any other election mechanism - **the new leadership grant carries a monotonically increasing number** (a term number, a znode's transaction ID, a dedicated counter from the coordination service). 2. **The leader is required to attach this token to every write** it sends to shared storage. 3. Critically, **the storage system itself must track the highest token it has ever accepted** and reject any incoming write whose token is lower. This means even if a zombie old leader wakes up and issues a write, that write simply gets rejected the moment a legitimate newer leader has performed even a single write with its higher token - the old leader's write is denied not because anyone told it to stop, but because the resource itself refuses stale credentials. This shifts the safety property from 'every process behaves correctly and stops promptly when told' (a liveness-dependent, cooperative assumption) to 'the resource enforces correctness regardless of what any client believes' (a much stronger, storage-enforced guarantee). ## STONITH, the blunter alternative A blunter alternative used in classic high-availability clustering software (Pacemaker, older SAN-based failover clusters) is **STONITH** - Shoot The Other Node In The Head - where, instead of relying on cooperative token checks at the storage layer, the cluster's fencing mechanism actively and forcibly powers off, reboots, or cuts network/power access to the presumed-dead old leader's physical or virtual machine before allowing a new leader to take over shared resources like a SAN volume or a floating IP. This is more disruptive (it genuinely kills the old node rather than just rejecting its writes) and requires out-of-band control (IPMI, a cloud provider API, a managed power strip) that's independent of the primary network path the cluster uses for its own coordination messages, precisely so that a network partition affecting the coordination path doesn't also block the fencing action. STONITH trades operational complexity and blast radius for a guarantee that doesn't depend on every downstream storage system correctly implementing token checks. ## Which one fits your constraints Which approach a team picks depends heavily on what they control: | The constraint | What fits | |---|---| | **If you own the storage layer** (a custom service, a database you can add a token column and a rejection check to) | fencing tokens are strictly better - no forced reboots, no destructive action, and correctness holds even if the old leader keeps running harmlessly forever, since its writes are simply always rejected. | | **If the shared resource is something you can't modify to add token checks** (legacy hardware, a third-party SAN, a physical device) | STONITH-style forcible fencing may be the only option available, accepting the operational cost of actually killing a node that might otherwise have recovered gracefully. | Production systems that care about this distinction explicitly document it: - **Kubernetes' controller-manager leader election** combined with resource version checks on the Kubernetes API server effectively implements token-based fencing (`resourceVersion` acts as the token). - **Pacemaker/Corosync clusters** for traditional shared-storage HA setups document STONITH as a hard requirement, refusing to promote a new primary at all until fencing of the old one is confirmed, specifically to avoid the exact split-brain-writes-corrupt-data scenario this question is about.
- Why doesn't the old leader just check in with the coordination service before every write to confirm it's still leader?It can, and some systems do add this check, but it doesn't fully close the gap: there's still a race window between the confirmation check and the actual write where leadership could change, and worse, a stop-the-world pause can happen after a successful check but before the write is actually sent, during which the lease or term can expire - a fencing token check performed by the storage layer itself, atomically with the write, avoids this race entirely.
- Can fencing tokens be implemented without a central coordination service issuing them?The token still needs to come from somewhere that guarantees monotonicity and uniqueness across leadership changes, which in practice is almost always the same coordination mechanism used for election itself - Raft's term number, ZooKeeper's zxid, or an explicit counter in the lease-granting service - because independently generated tokens (like leader-local counters) could collide or fail to be strictly increasing across a leadership change.
- Why is STONITH considered riskier or more disruptive than fencing tokens even though both prevent split-brain?STONITH forcibly terminates a node that might actually still be healthy and merely network-partitioned, destroying any in-flight work it was doing and requiring it to fully restart and rejoin the cluster from scratch, whereas fencing tokens let a zombie leader keep running harmlessly - its writes are simply rejected - so the node can rejoin gracefully once connectivity is restored without ever needing to be killed.
Firing an employee doesn't erase their old keycard automatically - if the building's door system doesn't actually deactivate the old badge, the fired employee can still walk in and touch things even though HR has 'elected' their replacement. Fencing tokens are like the door system checking a badge's issue date and rejecting anything older than the newest badge issued, while STONITH is more like physically changing the locks the moment someone new is hired.
saying these in an interview costs you the question
- Believes a correct election protocol alone is sufficient to prevent split-brain writes
- Cannot distinguish the control-plane (leadership agreement) from the data-plane (actual writes to shared storage)
- Thinks fencing tokens need to be checked by the leader itself rather than enforced by the storage/resource layer
- Confuses STONITH with fencing tokens as the same mechanism
- Assumes a partitioned old leader will always notice and stop on its own without any enforcement mechanism