You are designing the control plane for a self-managed, bare-metal Kubernetes cluster spread across racks. How do you choose between three and five control-plane nodes, and where do you place them?
answer
- budget the failures first
- maintenance uses up the spare
- no domain holds a majority
- two racks is the trap
- latency bounds the spread
basics
~20 sPick the member count from the failures you must survive: three tolerate one loss, and five tolerate two, including one planned. Spread members so no single rack or power domain holds a majority, which needs at least three failure domains with low-latency links.
solid answer
~50 sStart from the failure budget, not the node count. Three control-plane nodes (stacked, or three etcd members) survive one failure. That budget is used up whenever a node is down for patching or an upgrade, so a second fault during maintenance loses quorum. Five nodes survive two failures, so one can be in maintenance and one can fail unexpectedly. The costs: two more hosts, each write waiting on three acknowledgements, and more replication traffic. Placement matters as much as count. With three nodes in two racks, the rack holding two of them takes quorum with it. Five nodes in two racks (3+2) have the same problem. You need a third failure domain, even if it holds only one member, with low, stable latency, because etcd commits wait on the network. On a 9-node shipment-tracking cluster I would run three stacked nodes, one per rack, tainted, and document the maintenance risk. I would move to five nodes or external etcd only if the SLO cannot accept that risk.
go deeper
Remember the basics: three members survive one failure and five survive two, and they should not share a single rack.
Explain the majority arithmetic, why even counts add no tolerance, and why a 2+1 split across two racks fails.
Bring in the operational facts: maintenance uses up the failure budget, etcd needs fast disks and low latency, and the endpoint must be spread too.
Own the trade-off: failure budget, failure domains, hardware cost and team recovery speed together decide the design, and the accepted risk should be written down.
## Start from the failure budget A control plane stays writable while its etcd cluster keeps a **majority**. Choose the size from how many simultaneous losses you must survive: | etcd members | Majority | Tolerated losses | Typical use | |---|---|---|---| | 1 | 1 | 0 | labs, throwaway clusters | | 3 | 2 | 1 | most self-managed production clusters | | 5 | 3 | 2 | strict availability targets, frequent maintenance | | 6 | 4 | 2 | never: same tolerance as 5, more cost | | 7 | 4 | 3 | rarely justified; replication overhead grows | The detail people miss is **maintenance**. Rebooting a control-plane node for kernel patches or a Kubernetes upgrade uses up the single tolerated failure of a three-member cluster for as long as it takes. Any unplanned fault in that window, such as a disk, a PSU or a mispatched cable, loses quorum. Five members keep one failure in reserve during maintenance. ## What five costs - **Two more hosts.** On a small cluster, that capacity comes out of the worker pool. - **Slower writes.** Every commit waits for three members to write to disk instead of two. The difference is small on good disks and noticeable on poor ones. - **More replication and more API servers.** With stacked etcd, five nodes also means five API server watch caches holding cluster state. - **More operational steps.** Every upgrade and certificate rotation touches more machines. Five members do **not** increase throughput. etcd serves writes through a single Raft leader, so adding members adds safety, not write capacity. Control-plane *read* capacity can grow by adding API servers, which is easier with external etcd. ## Placement: count failure domains, not machines A failure domain is anything that takes several machines down together: a rack, a PDU, a top-of-rack switch, a room. 1. **Never let one domain hold a majority.** Three nodes in two racks become 2+1. Losing the 2-node rack loses quorum, so this design is no better than a single rack for that failure. 2. **Five nodes in two domains do not fix it.** 3+2 still loses quorum when the 3-node side fails. 3. **Use three domains.** Place members 1+1+1 or 2+2+1, so any single domain can fail. 4. **Keep latency low and stable between domains.** Each etcd write waits for replication across the network. Stretching members across distant sites trades one failure mode for slow writes and spurious leader elections. 5. **Put the endpoint in the same design.** A virtual IP or load balancer that lives in one rack re-creates the single point of failure the spread was meant to remove. ## Applying it to the shipment-tracking cluster The platform has 9 bare-metal machines in three racks and no cloud load balancer. - **Recommended:** three stacked control-plane nodes, one per rack, keeping the `node-role.kubernetes.io/control-plane` taint so shipment-tracking Pods run on the six workers. The endpoint is a virtual IP that can move to any of the three racks. Each node gets fast local disks for etcd. - **Accepted risk:** during a control-plane node's maintenance, one more fault loses writes. The data plane keeps serving, so this is a window of lost control, not an immediate outage. Short, scheduled maintenance windows keep it small. - **When to move to five nodes:** if the SLO cannot accept that window, or if maintenance is frequent and long. With nine machines, five stacked control-plane nodes leave four workers. The alternatives are removing the taint, which means accepting resource contention with the API, or buying hardware. - **When to use external etcd:** if etcd disk latency suffers from API server load, or if you need more API servers without more etcd members. ## Decision checklist - How many simultaneous faults, including planned ones, must the control plane survive? - How many independent failure domains exist, and does any one hold a majority? - Are inter-domain latency and disk sync times good enough for etcd? - Is the API endpoint itself spread across those domains? - Can the team rebuild a lost member quickly? Fast recovery shortens the risk window as much as a larger count does. There is no universally right answer. Three well-placed nodes with fast recovery often beat five nodes packed into two racks.
- You only have two racks. What are your options?Add a third failure domain for one member: a separate power circuit and switch, a small host in another room, or a nearby site with low latency. That member can be an external etcd host, with no API server. Otherwise accept that losing the rack holding the majority stops the control plane, and plan for a fast rebuild from backups. Five nodes split across two racks do not fix it.
- Would you let workloads run on the control-plane nodes of the 9-node cluster?Only when it is clearly worth it. Removing the taint gains capacity, but a busy shipment-tracking Pod can then compete with etcd for disk and CPU and slow every cluster write. If capacity is that tight, I would rather keep the taint and add workers, or set strict resource limits and keep the Pods off the etcd disk.
saying these in an interview costs you the question
- Five etcd members make writes faster than three
- Three control-plane nodes in two racks survive either rack's failure
- An even member count adds a tie-breaker and more safety
- More control-plane nodes are needed once workers exceed a few dozen
- Stretch etcd members across distant regions for maximum resilience