Suppose a network partition isolates the current Raft leader from the rest of the cluster, but that leader doesn't crash - it just can't reach anyone. The majority side elects a new leader. When the partition heals, how does the protocol prevent both leaders from being treated as valid, and what stops the old leader from having accepted writes in the meantime?
answer
- term = election epoch, monotonic
- one vote per term per node
- higher term seen -> step down to follower
- stale leader can't get majority acks -> nothing commits
- no partition detection needed, staleness self-corrects
basics
~20 sEvery election bumps a counter called a term. The old leader's term is now lower than the new one's, so once they reconnect, everyone - including the old leader - recognizes the higher term and steps down; and the old leader couldn't commit any writes alone anyway since it needs majority acks it never got.
solid answer
~50 sRaft attaches a monotonically increasing term number to every leader and every message. When the majority side times out on the old leader, it elects a new leader with a strictly higher term. The isolated old leader keeps thinking it's leader and may keep accepting client writes locally, but it can never commit them because AppendEntries to the unreachable majority never succeed - those writes stay uncommitted and are never acknowledged to clients as durable. When the partition heals, the old leader receives an RPC (or response) carrying the higher term number; per Raft's rule ('if a request or response contains term T > currentTerm, set currentTerm = T and convert to follower'), it immediately steps down to follower and its uncommitted tail is overwritten to match the new leader's log. There's no window where two leaders are both treated as authoritative - only the higher-term leader can ever get majority acknowledgment.
go deeper
Should know the basic idea: only one node can really be 'leader' at a time because it needs agreement from most of the cluster, and the old one gets corrected once reconnected.
Should explain that terms increase on every election and that a node adopts the higher term it sees and steps down, in plain terms.
Should walk through the full mechanism: one-vote-per-term preventing two same-term leaders, quorum requirement preventing the stale leader from committing, and the 'higher term wins, no detection needed' rule.
Should discuss why partition detection is fundamentally avoided rather than attempted, the operational monitoring implications (healthy-looking but non-committing leader), and persistence-of-term-to-stable-storage as a hardening detail.
## The two cooperating mechanisms Raft prevents split-brain using two cooperating mechanisms: 1. a strictly increasing **term** number, and 2. the **majority-quorum requirement** for every consequential action (both winning an election and committing a log entry). Neither mechanism alone is sufficient - it's their combination that closes the gap. ## Terms as a logical clock Every node tracks a `currentTerm`, starting at 0 and incrementing every time an election is called. A term acts like a logical clock: within a single term there can be at most one leader, guaranteed because a candidate needs a majority of votes to win, and Raft's voting rule says a node casts at most one vote per term, so two different candidates can't both collect a majority in the same term - their vote sets would have to overlap, and an overlapping voter can't vote twice. When a partition isolates the current leader from the majority, the followers on the majority side each start their own randomized election timeout (since they stop receiving heartbeats), and one of them - say node B - times out first, increments its term from, say, 5 to 6, and requests votes. Because the old leader (still on term 5) is unreachable to vote and everyone else on the majority side is also on term 5 and hasn't voted yet in term 6, node B collects a majority of the reachable nodes and becomes leader for term 6. ## What the isolated leader can and cannot do Meanwhile, the old leader (still calling itself leader on term 5) is completely unaware of this. - It keeps sending heartbeats and `AppendEntries` into the void, and if it receives client writes, it appends them locally as uncommitted entries. - Crucially, it can never mark them committed, because committing requires a majority of `AppendEntries` acknowledgments, and it can't reach the followers who form that majority. - So from the client's point of view, assuming the client waits for a commit acknowledgment as it should, those writes never succeed - they just hang or eventually time out. This is the second half of the mechanism: term numbers prevent two leaders from ever winning the same election, and the quorum requirement independently prevents the stale leader from making its writes durable even while it's mistakenly still acting as leader. ## When the partition heals When the partition heals, the reconciliation is almost mechanical. Raft's rule is simple and applied uniformly to every RPC: whenever a node - leader, candidate, or follower - sees a term number higher than its own `currentTerm` in any incoming request or response, it immediately updates `currentTerm` to that higher value and reverts to follower state, abandoning any leader or candidate role it held. So the moment the old leader (term 5) either sends a heartbeat to node B (now term 6, and B's response carries term 6) or receives a `RequestVote/AppendEntries` from B, it sees term 6 > its own term 5, steps down, and becomes a follower of B. Its uncommitted log entries from the partition window get resolved by the standard log-matching-and-backtrack process - B's log wins for any conflicting positions, and the old leader's uncommitted entries are discarded, exactly as if they had never happened, which is safe because they were never acknowledged to clients as committed. ## Why detection is sidestepped The reason this design exists is that a network partition, from any single node's local point of view, is indistinguishable from the rest of the cluster being slow, offline, or crashed - there's no reliable way for the isolated leader to detect 'I've been partitioned' versus 'everyone else happens to be busy.' Rather than trying to detect partitions (which is provably impossible in general), Raft sidesteps the detection problem entirely: - it makes staleness self-correcting via term comparison, and - it makes premature commitment structurally impossible via the quorum requirement, so a stale leader is harmless by construction rather than by any explicit check. ## The trade-off and the operational trap The trade-off is the liveness cost already discussed - the isolated leader is wasted for the duration of the partition, unable to do useful work even though it's fully healthy, and clients talking to it experience hung requests rather than fast failures unless they have their own timeout logic. In production, this shows up as a subtle operational trap: monitoring that only checks whether the leader process is up and responding to heartbeats can show green while writes are actually failing, because the leader is up but not committing - which is why healthy Raft-based systems (etcd, Consul) expose commit-index / quorum-health metrics separately from basic liveness checks, and clients are expected to implement request timeouts with leader-redirect retries rather than trusting a single node indefinitely.
- Could the isolated old leader accidentally 'win' by responding to a client and telling it the write succeeded before the partition heals?Only if it incorrectly told the client success without waiting for a majority commit acknowledgment, which is a client-facing implementation bug, not a Raft protocol violation - a correctly implemented Raft leader never reports a write as durable until it has committed, and it structurally cannot commit while isolated from a majority.
- What if the partition heals but the old leader's term counter was somehow corrupted to be higher than the new leader's?Term numbers aren't wall-clock time, so there's no clock-skew risk, but if a term counter were corrupted to be artificially high, that node could win future elections it shouldn't - this is exactly why Raft implementations persist currentTerm to stable storage before responding to any RPC, so a term can only advance through the legitimate election/RPC process, never through external tampering.
- Does this term-based mechanism also work if the partition is a full network split rather than just the leader being cut off (e.g., a true minority/majority split)?Yes - the mechanism is identical; the only difference is which side, if any, can reach a majority. Term numbers plus quorum requirements resolve leadership uniquely on whichever side (if any) can form a majority, and the minority side simply never succeeds in electing anyone.
Like a company where every new CEO announcement comes with a higher-numbered memo, and every employee is instructed to always obey the highest-numbered memo they've seen and ignore older ones - a CEO who got physically cut off from head office and kept issuing orders under an old memo number gets automatically overruled the instant anyone shows them the newer memo, no investigation needed.
saying these in an interview costs you the question
- thinks Raft needs to explicitly detect that a partition occurred
- believes the old leader could commit writes while cut off from the majority
- doesn't mention term numbers or confuses them with timestamps/wall-clock time
- thinks both leaders remain 'leader' simultaneously after the partition heals
- suggests manual/operator intervention is required to resolve the stale leader, rather than the protocol resolving it automatically