In a highly available Kubernetes control plane, how do stacked and external etcd topologies differ, and what failures can each tolerate?
answer
- where do the members live
- stacked: one node, two roles
- majority of three is two
- external: separate failure budgets
- disk sync contention
basics
~20 sStacked etcd runs a member on each control-plane node beside kube-apiserver; external etcd uses dedicated hosts. Stacked needs fewer machines but couples failures: three stacked nodes tolerate one loss, and each loss removes an API server and an etcd member.
solid answer
~40 sIn a **stacked** topology, each control-plane node runs `kube-apiserver`, kube-controller-manager, kube-scheduler **and** an etcd member. With kubeadm, each API server's `--etcd-servers` points at its local member on `127.0.0.1`. In an **external** topology, etcd runs on its own hosts and every API server lists all etcd members. Quorum math is the same in both: 3 members need 2 to agree, so they tolerate 1 loss, and 5 members need 3, so they tolerate 2. The difference is coupling. Losing one stacked node costs an API server and an etcd member at once. With external etcd you can lose an API-server host without touching quorum, and etcd's disk I/O no longer competes with the API server. The price is more hosts (typically 6 instead of 3) and a second fleet to patch and monitor.
go deeper
Remember the two layouts: stacked puts etcd on the control-plane nodes, external puts it on its own hosts. Three members survive one loss.
Explain why coupling matters: a stacked node failure costs an API server and an etcd member together, while external etcd keeps those failure budgets separate.
Show the operational trade-off: etcd write latency depends on disk syncs, so a busy control-plane node can slow the cluster. Also note that an unhealthy local member takes its API server out through /readyz.
Decide with evidence: host count, disk quality, control-plane load and team capacity to run a second fleet all go into the choice, not topology fashion.
## What the two topologies are In Kubernetes, **etcd** is the key-value store that holds every API object, and `kube-apiserver` is its only client. An HA control plane replicates both. The question is where the etcd members live. - **Stacked etcd:** each control-plane node runs the full set: `kube-apiserver`, `kube-controller-manager`, `kube-scheduler` and one etcd member. This is kubeadm's default. Each API server talks to the etcd member on its own node, so `--etcd-servers` points at `127.0.0.1`. - **External etcd:** the etcd members run on separate hosts. Every API server's `--etcd-servers` lists all of those members, so any API server can use any healthy member. ## Quorum math is identical etcd replicates writes with the Raft consensus protocol, and a write commits only when a **majority** of members acknowledge it. The member count, not the topology, sets the failure budget: | Members | Majority needed | Members you can lose | |---|---|---| | 1 | 1 | 0 | | 3 | 2 | 1 | | 4 | 3 | 1 | | 5 | 3 | 2 | An even count adds a machine without adding tolerance. That is why control planes use three or five members. ## Where the topologies differ: coupled failures Take the shipment-tracking platform's 9-node bare-metal cluster with three stacked control-plane nodes. 1. **One node fails.** The cluster loses one API server *and* one etcd member. Two etcd members remain, which is still a majority, and two API servers keep serving. The cluster stays read-write. 2. **A second node fails.** One etcd member of three is not a majority. The last API server is still running but cannot commit writes, so the control plane is effectively down even though a process is still answering. 3. **Only the etcd member on a node fails.** That node's API server fails its `/readyz` check, because the check includes etcd readiness and this API server only talks to its local member. A well-configured load balancer stops sending it traffic. With **external** etcd (for example, three API-server hosts and three etcd hosts), the failure budgets are separate. You can lose two of three API servers and still have a working, if thin, control plane, because the API servers keep no state. Independently, you can lose one of three etcd members. One host failure never counts against both budgets. ## The trade-off | Aspect | Stacked | External | |---|---|---| | Hosts for a 3-member setup | 3 | 6 (3 + 3) | | A host failure costs | API server + etcd member | only one of them | | Disk and CPU contention | etcd shares the node with the API server | etcd has the node to itself | | Operational surface | one fleet | two fleets, separate client certificates and upgrades | | Scaling | API servers and etcd grow together | each grows on its own | Key points when choosing: - **etcd is latency-sensitive.** Every write waits for disk syncs on a majority of members. On a busy node, a hungry API server or a noisy neighbour can slow those syncs. That is the strongest argument for external etcd on large clusters. - **Stacked is the pragmatic default** for small and medium self-managed clusters. Three nodes, one role, and quorum loss requires two simultaneous node failures. - **External etcd pays off** when the control plane is large, when etcd needs dedicated fast disks, or when you want to scale API servers for read load without growing the etcd cluster. - **Whichever you choose, the member count stays odd.** Adding a fourth stacked node for extra API capacity also adds a fourth etcd member, which raises the majority to 3 without adding tolerance. ## What this question is not How kubeadm joins a node with either topology, how the etcd certificates are issued, and how to restore etcd from a snapshot are separate subjects. The design question here is only this: which failures take down which part of the control plane, and what does it cost to separate them.
- Why does a fourth stacked control-plane node add no failure tolerance?It adds a fourth etcd member, and a majority of four is three. Four members tolerate one loss, the same as three. You pay for an extra machine and an extra replication target, and every write now waits on three acknowledgements instead of two. Go from three to five if you need to survive two failures.
- In a stacked cluster, what should happen when only one node's etcd member is unhealthy?That node's kube-apiserver only talks to its local member, so its `/readyz` fails on the etcd readiness check. The load balancer in front of the API servers should health-check `/readyz` and remove that backend. The other two API servers keep serving through the remaining etcd majority.
saying these in an interview costs you the question
- External etcd removes the need for an odd member count
- Two surviving stacked nodes out of four keep quorum and add safety
- A lone surviving API server can still accept writes without etcd quorum
- Stacked etcd tolerates more failures because each node is self-contained
- External etcd needs fewer hosts than stacked etcd