skip to content

Why is the Kubernetes cluster store etcd almost always deployed with 3 or 5 members rather than 2 or 4, and what happens to the cluster when a majority of those members is unavailable?

level: middleimportance: must knowfreq 60%

answer

  1. Raft: leader + majority commit
  2. tolerance = floor((N-1)/2)
  3. 4 members = same tolerance as 3, more latency
  4. no quorum -> writes fail, Pods keep running
  5. snapshot restore = lose writes since snapshot

basics

~20 s

etcd uses Raft: a write must be replicated to a majority (quorum) before it commits, so an N-member cluster tolerates (N-1)/2 failures. Even sizes add a member without adding fault tolerance. Without quorum etcd rejects writes, so the control plane freezes — but running Pods keep running.

solid answer

~50 s

etcd replicates via **Raft**: one leader, and every write commits only after a **majority** of members persist it. Majority of 3 is 2 (tolerates 1 failure); majority of 4 is 3 (still tolerates only 1); majority of 5 is 3 (tolerates 2). Even-sized clusters cost an extra machine and an extra fsync on the write path for zero extra fault tolerance, and raise the chance of split votes during elections. Hence 3 or 5; beyond 5, write latency grows because every commit waits on more replicas. **Quorum loss** (2 of 3 members down): no leader can be elected, so all writes fail. kube-apiserver errors on create/update/delete; scheduling, scaling, rollouts and controller reconciliation stop. Linearizable reads fail too, though the API server's watch cache can still serve some reads. Crucially the **data plane keeps running**: existing Pods stay up and kube-proxy keeps routing. Recovery means restoring the lost members, or — if they are permanently gone — restoring a snapshot into a fresh single-member cluster and growing it back.

code

bash · 3 lines
bash
etcdctl member list -w table
etcdctl endpoint status --cluster -w table
etcdctl endpoint health --cluster

go deeper

for a junior

Know that etcd replicates with Raft, needs a majority to accept writes, and is therefore deployed in odd numbers — usually 3.

for a middle

Do the arithmetic out loud: floor((N-1)/2), why 4 buys nothing over 3, and what specifically breaks when quorum is lost versus what keeps running.

for a senior

Talk about diagnosis and recovery order — restore members before restoring snapshots, replace one at a time, use learners — plus disk latency as the dominant health signal.

for a principal

Reason about failure-domain placement, region latency on the commit path, RPO set by snapshot cadence, and when a second cluster beats a bigger etcd.

## Raft in one paragraph Raft is a consensus algorithm that keeps a replicated log identical across a set of servers. Members elect a **leader** for a term; all client writes go to the leader, which appends the entry to its log and replicates it to followers. Once a **majority** (including itself) has durably written the entry, the leader marks it committed, applies it to the key-value state machine, and answers the client. Followers that fall behind catch up from the leader; if the leader dies, followers whose logs are at least as up to date start an election and a new leader is chosen once it wins a majority of votes. The majority rule is the whole point: two disjoint majorities cannot exist in the same member set, so a split network can never produce two leaders committing conflicting writes. That is what makes etcd linearizable, and therefore what makes Kubernetes state coherent. ## Why odd member counts Fault tolerance for N members is floor((N-1)/2): | Members | Quorum | Failures tolerated | |---|---|---| | 1 | 1 | 0 | | 2 | 2 | 0 | | 3 | 2 | 1 | | 4 | 3 | 1 | | 5 | 3 | 2 | | 7 | 4 | 3 | Going from 3 to 4 buys nothing: you still survive only one failure, but you have added a machine to maintain and every commit now waits for one more disk fsync. Going from 4 to 5 is what actually buys the second failure. Odd sizes are the only sizes on the efficient frontier, and they avoid tied elections. Why stop at 5? Every additional member adds replication work and, more importantly, the leader must hear from more disks before committing, so p99 write latency rises. Three is standard; five is used when you want to survive losing two control-plane nodes, or to tolerate one being down for maintenance while still surviving a failure. Scaling reads is not a reason to add voting members — use serializable reads or **learner** members, which replicate but do not vote and do not count toward quorum, which also makes them the safe way to stage a replacement member. ## What quorum loss looks like Suppose a 3-member etcd loses 2 members: - `etcdctl endpoint status` on the survivor shows no leader; logs repeat election attempts. - kube-apiserver returns 500s or `etcdserver: request timed out` on writes; `kubectl apply`, scale and delete all fail. - Controllers cannot update status, so rollouts, HPA scaling and Job creation stall. - Nodes cannot renew their Leases, so the node controller would mark them NotReady — except it cannot write that either. - Existing Pods keep running and Service routing keeps working, because kubelet and kube-proxy operate on state they already have. A cluster with no control plane is degraded, not dead. The surviving member deliberately refuses writes rather than serving forked data: availability is traded for consistency. ## Recovering Order of preference: 1. **Bring the failed members back.** If their data directories and disks are intact, restarting them restores quorum with no data loss. Always try this first. 2. **Replace a member.** If one member's data is corrupt, `etcdctl member remove` then `member add` (as a learner first is safer) and start the new node with `--initial-cluster-state existing`. One member at a time, so quorum is never lost. 3. **Restore from snapshot.** If quorum is unrecoverable, restore a snapshot into a fresh single-member cluster, then grow back to 3. You lose every write made after the snapshot — which is why snapshot cadence defines your RPO. ## Placement matters as much as count Quorum math assumes independent failures. Three members on one host, rack or availability zone tolerate one *process* failure but zero infrastructure failures. Spread members across failure domains, and beware two-AZ layouts: 3 members split 2/1 means losing the 2-member zone loses quorum. Avoid stretching etcd across regions — every commit pays the inter-region round trip on the write path.

  • You need to survive a full availability-zone outage. How many etcd members do you deploy and where?
    Spread members across at least three zones so no single zone holds a majority — typically 3 members, one per zone, or 5 as 2/2/1. With only two zones there is no safe layout: whichever zone holds the majority becomes a single point of failure. Keep all members in one region, because Raft commits pay the inter-member round-trip latency on every write.
  • Does adding etcd members improve read performance?
    Not for linearizable reads, which the leader confirms with a quorum round trip, and writes actually get slower because more disks must acknowledge. Serializable reads can be served locally by any member, and non-voting learner members can absorb some load without affecting quorum math. In practice read pressure is relieved at the API server's watch cache, not by growing etcd.

A committee that requires more than half its members to sign off: adding a fourth member to a three-person committee does not let more people be absent, it just means one more signature to collect.

saying these in an interview costs you the question

  • Saying more etcd members means better performance or higher write throughput
  • Thinking a 4-member cluster tolerates 2 failures
  • Claiming running Pods stop when etcd loses quorum
  • Believing a 2-member etcd is highly available
  • Restoring from snapshot as the first response while the failed members' disks are still healthy

context