skip to content

In Raft, nodes use randomized election timeouts, terms, and majority votes to pick a leader. Walk through exactly how a follower becomes a candidate and wins an election, and explain why the timeout is randomized rather than fixed.

level: middleimportance: must knowfreq 70%

answer

  1. follower -> candidate on timeout, vote for self
  2. RequestVote RPC, majority wins
  3. term number rejects stale messages
  4. randomized timeout avoids split votes
  5. log up-to-date check prevents regressing committed entries

basics

~20 s

Every node waits a random amount of time to hear from the leader. Whoever times out first asks everyone to vote for it. Getting votes from more than half the nodes wins. The random wait stops everyone timing out together.

solid answer

~50 s

Each Raft node is a follower, candidate, or leader, and tracks a monotonically increasing 'term' number. A follower that doesn't hear a heartbeat from the leader within its randomized election timeout (commonly 150-300ms, re-randomized each time) becomes a candidate: it increments its term, votes for itself, and sends RequestVote RPCs to all peers. A peer grants its vote if it hasn't already voted in that term and the candidate's log is at least as up-to-date as its own. Once the candidate collects votes from a strict majority of nodes, it becomes leader for that term and starts sending heartbeats to suppress further elections. The timeout is randomized specifically so that, most of the time, only one follower times out first and cleanly wins before others even start an election; a fixed timeout would make simultaneous candidacies (and repeated split votes) far more likely.

go deeper

for a junior

Should give the basic flow: nodes wait for a heartbeat, whoever times out first asks for votes, majority wins.

for a middle

Should accurately name terms, RequestVote RPCs, majority quorum, and explain in their own words why randomized timeouts reduce split votes.

for a senior

Should explain the log-up-to-date voting condition and why it's necessary for safety, plus discuss real timeout-tuning trade-offs (leader flapping vs. slow recovery).

for a principal

Should reason about heartbeat-to-election-timeout ratios in production systems (e.g. etcd/Consul), the availability cost of mistuned timeouts under real network jitter, and how this interacts with client-facing latency SLOs during re-elections.

## Three states and a term number Raft was designed explicitly to be an understandable alternative to Paxos, and its leader-election sub-protocol is the piece most engineers interact with directly, since it governs how quickly a cluster recovers after a leader failure and how safely it avoids conflicting leadership. Every Raft node is always in exactly one of three states - **follower**, **candidate**, or **leader** - and every decision is stamped with a **term**, a strictly increasing integer that acts like a logical clock for leadership epochs; at most one leader can be elected per term, and any message carrying an older term is rejected, which is the core mechanism that lets nodes cheaply detect and reject stale information. ## From follower to candidate In steady state, the leader sends periodic empty `AppendEntries` RPCs as heartbeats to every follower. Each follower resets an election timeout timer whenever it receives a valid heartbeat or log entry from a leader it recognizes (matching or higher term). If a follower's timer expires before any heartbeat arrives - because the leader crashed, was partitioned away, or is simply too slow - that follower transitions to candidate: it increments its own term by one, votes for itself, resets its own election timer, and sends `RequestVote` RPCs in parallel to every other node, including its candidate term, its ID, and metadata about how up-to-date its own log is (the index and term of its last log entry). A recipient grants a vote only if three conditions hold: 1. the requesting term is at least as high as anything it has seen; 2. it has not already voted for a different candidate in that term; 3. the candidate's log is at least as up-to-date as the recipient's own log by Raft's log-comparison rule (higher last-term wins, ties broken by longer log). That last condition is the leader-election piece's main safety contribution: it guarantees a new leader always already holds every entry that was committed by any previous leader, so leadership never regresses committed history. ## Why a strict majority settles it A candidate becomes leader the moment it collects votes from a **strict majority** of the cluster (including its own self-vote), because in Raft's fault model any two majorities out of N nodes must overlap in at least one node, so only one candidate can ever reach a majority in a given term - this is precisely why Raft is safe even though multiple candidates can start elections concurrently. On becoming leader, the node immediately starts sending heartbeats to establish authority and reset every follower's timer, preventing further elections while it remains healthy. ## Why the timeout is randomized Randomization of the election timeout, typically drawn freshly from a range like 150-300ms on every reset, exists to solve the 'split vote' problem. - **If every follower used the same fixed timeout**, a leader failure would cause many followers to become candidates in the very same instant, each voting for itself and requesting others' votes, with the votes for that term splitting across several candidates so nobody reaches a majority - the term then increments again and the same tie is likely to repeat. - **By making each follower's timeout an independent random draw**, in the overwhelming majority of cases exactly one follower's timer expires meaningfully before the others notice anything is wrong; that follower sends its `RequestVote` RPCs, the rest of the (still-timed-out-but-not-yet-candidate) followers see a valid, higher-term request first and grant their vote instead of becoming candidates themselves, and the election resolves in a single round without contention. If a split vote does still occur - say, two followers' timers happen to fire within the same short window - Raft simply lets the term expire without a winner, and the randomized timeout mechanism makes it statistically very likely the next round resolves cleanly, since the odds of the same tie recurring shrink multiplicatively with each retry. ## Tuning it in production The practical failure modes engineers actually hit in production are less about the algorithm's correctness (Raft's safety proof holds regardless of network delay) and more about tuning. - **An election timeout set too low** relative to real network latency or GC pauses causes unnecessary 'leader flapping' - a healthy leader gets falsely presumed dead under transient load, triggering pointless re-elections and brief write unavailability while a new leader catches up and re-establishes quorum. - **Too high a timeout** means genuinely slow leader-failure recovery, directly hurting availability after a real crash. Systems like etcd (built on Raft) default to timeouts in the hundreds-of-milliseconds range and explicitly document this trade-off, and operators tune heartbeat interval versus election timeout ratios (commonly at least a 1:10 heartbeat-to-timeout ratio) to keep false elections rare while still recovering quickly from genuine failures. CockroachDB and Consul, both Raft-based, similarly expose these knobs because getting them wrong under real-world network jitter is the most common operational headache with Raft leader election, distinct from the underlying consensus safety - which, per this leaf's boundary, is a separate concern from the election mechanism itself.

  • What happens if two followers' randomized timeouts happen to expire close enough together that a split vote occurs?
    Neither candidate reaches a majority in that term, so the term times out unresolved; each candidate's own election timer (also freshly randomized) will eventually expire again, incrementing the term and starting a fresh round. Because the random draws are independent, the probability of the exact same tie recurring drops sharply each round, so the cluster converges quickly in practice even though there's no hard upper bound in theory.
  • Why does a candidate need the recipient's log to be checked, not just a majority of votes, to become leader safely?
    Without the log-up-to-date check, a node that fell behind (missed recently committed entries) could still win an election on raw vote count and then, as leader, overwrite or fail to replicate already-committed data, violating Raft's core safety guarantee. Requiring the candidate's log to be at least as current as each voter's log ensures the new leader already contains every entry any previous leader had committed.
  • How does a returning old leader (that was network-partitioned, not crashed) get demoted once the partition heals?
    The old leader either receives an AppendEntries or RequestVote message carrying a higher term number than its own from the new leader or a candidate, and per Raft's rule that any node seeing a higher term immediately reverts to follower and updates its own term - this single rule is what lets a stale leader step down cleanly without any special-cased demotion protocol.

It's like a group of people in a dark room waiting for a whistle from a leader; if nobody hears the whistle for a random amount of time (everyone's patience runs out at a slightly different moment), whoever gets impatient first shouts 'vote for me' - and because they shout first, everyone else votes for them before their own patience runs out too, avoiding everyone shouting over each other at once.

saying these in an interview costs you the question

  • Says Raft leader election uses a fixed timeout, missing why randomization matters
  • Cannot explain what a 'term' is or why stale-term messages get rejected
  • Believes a candidate can win with fewer than a strict majority of votes
  • Omits the log-up-to-date check when describing vote granting
  • Confuses election timeout with heartbeat interval, or doesn't know they're different, related settings

context