skip to content

HA Control Planes

Running three or five control-plane nodes instead of one: where etcd lives, how many members you can lose before writes stop, and how controller-manager and scheduler replicas agree on a single active leader. Self-managed cluster design interviews open here.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

5

If every kube-apiserver in a Kubernetes cluster becomes unreachable, what keeps running on the nodes and what stops working?

level: juniorimportance: must knowfreq 74%

answer

  1. control plane versus data plane
  2. kubelet keeps what it knows
  3. installed rules keep routing
  4. no writes: no scheduling, no scaling
  5. readiness changes never reach endpoints

basics

~20 s

Running Pods keep running: the kubelet keeps their containers alive and kube-proxy's installed Service rules keep routing. Anything that needs a write or fresh state stops working: scheduling, scaling, rollouts, rescheduling off failed nodes and kubectl.

solid answer

~40 s

Kubernetes separates the **control plane** (API server, etcd, scheduler, controllers) from the **data plane** (kubelet, container runtime, kube-proxy). When no `kube-apiserver` answers, each kubelet keeps running the Pods it already knows about and restarts crashed containers according to their `restartPolicy`. kube-proxy leaves its programmed rules in place, so traffic to existing Service backends keeps flowing. What stops is everything that must read new state or write: `kubectl` fails, kube-scheduler cannot bind new Pods, the ReplicaSet and Deployment controllers cannot create replacements, autoscaling stops, and Pods on a node that dies are not rescheduled. Endpoint membership also freezes, so a Pod that turns unready keeps getting traffic. An HA control plane exists to make that window rare. It does not make the outage harmless.

go deeper

for a junior

Remember the split: nodes keep running what they already have, while anything that needs a change made through the API stops. Name the kubelet and kube-proxy as what keeps going.

for a middle

Explain why: the kubelet reconciles a locally held spec and runs probes itself, while scheduling, replacement and endpoint updates all need API writes.

for a senior

Point out the hidden hazards: readiness changes that never reach routing, failed nodes that are handled only after recovery, and the full-disruption guard that prevents mass eviction when the control plane returns.

for a principal

Frame the outage as time-bounded risk. The data plane buys minutes, not days, so HA sizing, endpoint failover and alerting should be judged by how short they keep that window.

## Two halves of a cluster A Kubernetes cluster has two halves that fail differently: - The **control plane** decides what should run: `kube-apiserver` (the only component that reads and writes cluster state), **etcd** (where that state is stored), `kube-scheduler` (picks a node for each new Pod) and `kube-controller-manager` (the reconcile loops that create, scale and replace objects). - The **data plane** runs what was decided: the **kubelet** on every node, the container runtime behind the CRI, and **kube-proxy**, which programs Service routing into the node's packet-filtering rules. The data plane gets its instructions through the API server, but it does not need the API server to *keep doing* what it was last told. That asymmetry is the whole answer. ## What keeps running Picture a shipment-tracking API with six replicas on the six worker nodes of a 9-node bare-metal cluster. The three control-plane nodes lose their uplink at once. - **Containers stay up.** Each kubelet holds the Pod specs it last received and keeps reconciling them locally. If a shipment-tracking container crashes, the kubelet restarts it on the same node according to the Pod's `restartPolicy`. That restart is a node-local decision. - **Probes still run.** Liveness probes are executed by the kubelet, so a failing liveness probe still triggers a local restart. - **Static Pods keep running**, because the kubelet reads them from its local manifest directory. - **Service traffic keeps flowing** between existing Pods. kube-proxy's iptables or nftables rules stay installed, and the cluster DNS server typically keeps answering from its in-memory view. - **External traffic** that enters through a NodePort or an already-configured load balancer still reaches the Pods. ## What stops 1. **Every write and every fresh read.** `kubectl` gets connection errors. CI pipelines cannot deploy. A shipment-tracking rollout started before the outage simply halts. 2. **Scheduling.** kube-scheduler cannot see new Pods or write bindings. 3. **Self-healing across nodes.** If a worker loses power, nothing creates replacement Pods elsewhere, because the ReplicaSet controller cannot write. 4. **Scaling.** The HorizontalPodAutoscaler cannot read metrics through the API or change replica counts. 5. **Endpoint changes.** When a Pod fails its readiness probe, the kubelet cannot report the new status. The EndpointSlice controller never sees it, so kube-proxy keeps sending traffic to that Pod. 6. **Node heartbeats.** Kubelets cannot renew their node Leases or post status. | Keeps working | Stops working | |---|---| | Running containers and local restarts | Creating, updating or deleting any object | | Liveness probes and local restarts | Readiness changes reaching Service routing | | Installed kube-proxy rules | Scheduling new Pods | | Static Pods | Replacing Pods from a failed node | | Existing load balancer paths | Autoscaling, rollouts, `kubectl` | ## When the API server comes back Recovery is its own risk. Every node's heartbeat is now stale, so the node lifecycle controller in kube-controller-manager could conclude that all nodes died. It guards against this: when it sees every node as not ready at once, it enters a **full-disruption** mode and stops evicting Pods, on the assumption that the problem is the control plane rather than the whole fleet. Kubelets then reconnect, re-list their Pods and report status, and controllers resume from the current state. Kubernetes controllers are level-triggered, so nothing that happened during the outage has to be replayed. ## Why this matters for HA design This behaviour is why people say a control-plane outage is not a workload outage, at least for a while. The risk grows with time: the longer the window, the more likely a node failure, a bad Pod or a traffic spike happens while nothing can react. Running three or five control-plane nodes behind a stable endpoint keeps the window short. The data plane's independence is what makes a short window survivable.

  • A worker node loses power during the API outage. What happens to its Pods once the API server returns?
    Nothing happens during the outage, because no controller can write. After recovery, the node's Lease stays stale, so the node lifecycle controller marks the node unreachable and taints it. Pods without a matching toleration are evicted after the default toleration period, and their ReplicaSets create replacements on healthy nodes. The failure is handled late, not lost, because controllers compare current state with desired state instead of replaying events.
  • Can kubectl read anything useful from a node while the API server is down?
    No. kubectl only talks to the API server. On the node itself you can inspect containers through the CRI with `crictl`, read kubelet logs, and look at the static-pod manifests. Those are local views of what the kubelet is running, not cluster state, and they cannot change desired state.

It is like an air-traffic-control outage for planes already in the air: they keep flying their filed routes, but no new flight can take off and nobody can reroute around a storm.

saying these in an interview costs you the question

  • All Pods stop as soon as the API server is unreachable
  • The kubelet needs API server approval to restart a crashed container
  • Controllers will reschedule Pods off failed nodes during the outage
  • Readiness failures still remove Pods from Service routing during the outage
  • An HA control plane means a total API outage can never happen
open as a page

In a highly available Kubernetes control plane, how do stacked and external etcd topologies differ, and what failures can each tolerate?

level: middleimportance: must knowfreq 63%

basics

~20 s

Stacked etcd runs a member on each control-plane node beside kube-apiserver; external etcd uses dedicated hosts. Stacked needs fewer machines but couples failures: three stacked nodes tolerate one loss, and each loss removes an API server and an etcd member.

open as a page

In an HA Kubernetes control plane, why do kubelets and clients reach kube-apiserver through a load balancer, and how should it check backends?

level: middleimportance: should knowfreq 52%

basics

~20 s

kube-apiserver replicas are stateless and all active, so clients need one stable address that survives any single server's loss. Put a TCP pass-through load balancer or virtual IP in front of them, and have it health-check each server's /readyz endpoint.

open as a page

You are designing the control plane for a self-managed, bare-metal Kubernetes cluster spread across racks. How do you choose between three and five control-plane nodes, and where do you place them?

level: principalimportance: should knowfreq 40%

basics

~20 s

Pick the member count from the failures you must survive: three tolerate one loss, and five tolerate two, including one planned. Spread members so no single rack or power domain holds a majority, which needs at least three failure domains with low-latency links.

open as a page

The control-plane node running the active kube-controller-manager loses power, and a shipment-tracking rollout stalls for 47 seconds. What governs Kubernetes leader-election failover, and would you tune it?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

A standby takes over only after the leader's Lease looks expired: 15 seconds by default, plus a jittered 2-second retry. The rest of a 47-second stall usually comes from endpoint failover, stuck connections and cache sync.

open as a page