skip to content

Raft is often described as 'Paxos, but designed to be understandable.' Concretely, what problem does Raft's strong-leader approach solve relative to classic Multi-Paxos, and what does that design choice cost?

level: principalimportance: nice to knowfreq 25%

answer

  1. Multi-Paxos: leader election is an add-on, not specified
  2. Raft: leader is first-class, fully specified
  3. Paxos Made Live documented the spec gap
  4. Raft cost: single-leader throughput ceiling + full election on leader loss
  5. etcd/Consul/TiKV chose Raft for implementability

basics

~20 s

Multi-Paxos lets any node propose changes, which gets confusing and can cause competing proposals; Raft forces all changes through one leader at a time, which is much easier to reason about and implement correctly, but makes the leader a bottleneck and needs a full election protocol to replace it.

solid answer

~50 s

Classic Paxos (and naive Multi-Paxos) is specified around independent proposers that can each try to get a value accepted, with no built-in notion of a stable leader - implementers have to bolt on leader election and log-ordering themselves, and the base algorithm's proof is famously hard to map onto a practical, fully-specified implementation (this is why Google's account of implementing Chubby describes significant gaps between the paper and a working system). Raft addresses this by making strong leadership a first-class part of the protocol: only the leader can propose log entries, all writes flow through it in strict order, and leader election, log replication, and safety are specified together as one coherent algorithm with a full proof and reference implementation guidance. The cost is centralization: throughput is bounded by a single leader's capacity, and every leader failure requires a full election round-trip before writes resume, whereas some Paxos variants allow more flexible, less leader-dependent proposing patterns at the cost of that same complexity.

go deeper

for a junior

Should know Raft and Paxos both solve the same core problem (agreeing on a replicated log) and that Raft is generally considered easier to understand and implement.

for a middle

Should articulate that Raft makes leader election and log replication part of one specified algorithm, while classic Multi-Paxos leaves more to the implementer.

for a senior

Should explain the concrete cost - single-leader throughput ceiling and election-triggered unavailability window - and why organizations still choose Raft anyway.

for a principal

Should discuss the historical context (the documented Multi-Paxos implementation gap), name at least one leaderless alternative (e.g., EPaxos) and its trade-offs, and explain how sharding across many Raft groups sidesteps the single-leader throughput ceiling in practice.

## Where classic Paxos leaves off Paxos, as originally published by Leslie Lamport, specifies a two-phase protocol (Prepare/Promise, then Accept/Accepted) for getting a group of nodes to agree on a single value, using ballot numbers to resolve conflicts between competing proposers. It's provably correct and foundational, but the base algorithm ('single-decree Paxos') only agrees on one value ever - to build a real replicated log (needed for something like a database or a coordination service), you need 'Multi-Paxos': running many independent instances of Paxos, one per log slot, plus additional machinery to make that efficient and to decide who proposes for each slot. Critically, the original Paxos papers describe the reduction from single-decree to multi-decree only sketchily, leaving decisions like these as exercises for the implementer: - **leader election**, - **log-slot numbering**, and - **reconfiguration**. Different teams built genuinely different, incompatible Multi-Paxos variants, and even experienced engineers found translating the proof into a correct, complete system difficult - this is the exact gap Google's engineers documented in their well-known account of building Chubby, describing a nontrivial amount of unspecified engineering needed to get from the algorithm to a production system. ## What Raft makes first-class Raft, published by Diego Ongaro and John Ousterhout specifically with an explicit goal of understandability, closes that gap by making **strong leadership** a first-class, fully specified part of the algorithm rather than an implementation detail bolted on afterward. - In Raft, there is always at most one leader per term, and only that leader may append new entries to the log or decide the order of operations. - Followers never independently propose entries, they only replicate what the leader sends them, and clients are redirected to the current leader if they contact a follower. - Leader election, log replication, and the safety argument for why a newly elected leader is guaranteed to have all previously committed entries (**Leader Completeness**, enforced by the rule that a node votes only for a candidate whose log is at least as up-to-date as its own) are all specified as one integrated mechanism with a complete proof, and the original paper deliberately includes a full reference-quality specification, not just a proof sketch. This is why Raft became the default choice for many newer systems (etcd, Consul, CockroachDB's Raft layer, TiKV) - not because it's more powerful than Paxos, but because it's dramatically easier to implement correctly and to reason about when something goes wrong in production. ## The concrete cost, in two parts The concrete cost of centralizing all proposing through a single leader is twofold. 1. **First, throughput and latency are bounded** by that one leader's capacity to send and receive messages with a majority of followers - Raft doesn't allow, say, two different nodes to concurrently propose two different log entries for two different slots the way some Multi-Paxos variants permit (which can pipeline multiple concurrent proposals from different proposers when there's no leader contention). In practice this rarely matters for typical replicated-log workloads because a healthy single leader's network round trips to a handful of followers are fast, but it is a real theoretical ceiling that more flexible leaderless or multi-proposer designs (or protocols like EPaxos, explicitly designed to avoid a single-leader bottleneck) don't share. 2. **Second, every leader failure** - crash, network partition, or even just a slow node losing its position - triggers a full election: followers must individually notice missed heartbeats (after their randomized election timeout, typically on the order of 150-300ms in common implementations), a candidate must campaign and collect a majority of votes, and only then can new entries start committing again. During that window, the cluster is fully unavailable for writes, even though most of it is healthy - a startup cost Paxos-based systems that don't insist on a single stable leader can sometimes avoid, at the cost of the coordination complexity Raft was designed to eliminate. ## The choice teams actually make In production terms, this trade-off shows up as a deliberate choice: teams building a new coordination or metadata service today overwhelmingly reach for Raft (or a Raft-based product like etcd) specifically because getting a correct Multi-Paxos implementation from scratch is a known engineering risk, even though a well-tuned Paxos-family system can theoretically offer more flexible proposing patterns and, in some variants, lower tail latency by not always routing through one node. The understandability Raft optimizes for is not just an academic nicety - it directly reduces the chance of a subtle, hard-to-find bug in the exact code path responsible for a system's core correctness guarantee, which is a cost most engineering organizations are happy to pay the leader-bottleneck price for. ## A practical mitigation A practical mitigation worth knowing: Raft's single-leader throughput ceiling applies **per replicated log**, not per cluster as a whole. Real systems scale write throughput by sharding data across many independent Raft groups, each with its own leader, so aggregate throughput grows with the number of shards even though any single shard's writes still funnel through one leader at a time - this is exactly how systems like CockroachDB and TiKV scale far beyond what a single Raft group could handle.

  • If Raft has a throughput ceiling from routing everything through one leader, why don't more systems use a leaderless alternative like EPaxos instead?
    EPaxos and similar leaderless protocols can offer lower latency in some geo-distributed setups by letting nearby nodes propose without always contacting a distant leader, but they're substantially more complex to implement and reason about - for most workloads, a single leader's round trip to a majority of nearby replicas isn't the bottleneck, so teams accept Raft's simpler mental model over the marginal throughput gains of a leaderless design.
  • Does Raft's single-leader design mean it can't scale to handle high write throughput?
    It means the replicated-log write path itself is bounded by one leader's capacity, but real systems scale around that by sharding data across many independent Raft groups (each with its own leader), so aggregate cluster throughput scales with the number of shards even though each individual shard's writes still funnel through one leader at a time.
  • What specifically makes Raft's leader-election safety easier to verify than Multi-Paxos's equivalent guarantee?
    Raft ties leader eligibility directly to log completeness via a single explicit rule - a candidate can only win votes from nodes whose logs it is at least as up-to-date as - and proves Leader Completeness from that one rule, whereas classic Multi-Paxos descriptions typically treat leader election as a separate liveness optimization layered on top of the core agreement protocol, so the interaction between 'who becomes leader' and 'is the log preserved correctly' isn't argued as one unified proof.

Classic Multi-Paxos is like a recipe that says 'get everyone to agree on each dish, roughly like this' and leaves 'who's actually running the kitchen tonight' up to the restaurant to figure out; Raft is a recipe that also specifies exactly who the head chef is, how a new head chef takes over if the current one steps out, and proves that dishes never get served in the wrong order - at the cost of the kitchen only being able to move as fast as that one head chef can direct it.

saying these in an interview costs you the question

  • claims Raft is 'more powerful' or 'more fault tolerant' than Paxos rather than more understandable/implementable
  • thinks Multi-Paxos has no leader concept at all in practice, rather than an underspecified/optional one
  • doesn't know that a Raft leader failure causes a real (if brief) write-unavailability window
  • believes Raft's single-leader design caps a whole cluster's throughput rather than just one replicated log/shard's throughput
  • unaware of any real system (etcd, Consul, CockroachDB, TiKV) that chose Raft specifically for implementability

context