skip to content

When a Kubernetes cluster's API server becomes unreachable, what happens to pods that are already running, and what stops working?

level: juniorimportance: must knowfreq 60%

answer

  1. control plane vs data plane
  2. kubelet works from local state
  3. kube-proxy rules stay in the kernel
  4. no scheduling, scaling or endpoint updates
  5. missed CronJob and startingDeadlineSeconds

basics

~20 s

Running pods keep running and existing Service traffic keeps flowing, because kubelets and kube-proxy work from state they already have. Anything that needs the API server stops: scheduling, rescheduling, scaling, rollouts, endpoint updates and CronJob runs.

solid answer

~40 s

The data plane mostly keeps going. Each kubelet keeps running the containers it already has and still restarts crashed ones according to `restartPolicy`. kube-proxy keeps the Service rules it last programmed, so traffic to existing pods keeps flowing. What stops is anything that needs the API server. The scheduler cannot place pods. kube-controller-manager cannot replace pods from a dead node, scale a Deployment or advance a rollout. EndpointSlices stop updating, so a pod that dies can stay in the routing rules. `kubectl` gets no answer at all. A CronJob such as a nightly ledger-reconciliation batch does not start while the control plane is down. So a control-plane outage usually shows up first as "nothing changes", not as "everything is down", and it gets riskier the longer it lasts.

go deeper

for a junior

Remember the split: pods and Service routing already in place keep working, while scheduling, scaling, rollouts and kubectl stop until the API server is back.

for a middle

Explain why: the kubelet and kube-proxy act on state they already hold, and every controller is a loop that needs API reads and writes to act.

for a senior

Name the risks that grow with time: stale EndpointSlices sending traffic to dead pods, failed nodes not replaced, missed CronJobs, and the burst of load when everything reconnects.

for a principal

Treat control-plane availability as a limit on how long the cluster can safely run unattended, and design batch schedules, deadlines and recovery load with that in mind.

## Why the cluster does not stop Kubernetes separates the **control plane** from the **data plane**: - The **control plane** is `kube-apiserver`, etcd (the cluster store), `kube-scheduler` and `kube-controller-manager`. It decides what should happen. - The **data plane** is every node's **kubelet**, container runtime and `kube-proxy`, plus the pods themselves. It carries out the decisions. The kubelet does not ask the API server's permission for each action. It **watches** for pod assignments, keeps a local copy of what it was told to run, and drives the container runtime to match. If the watch breaks, the kubelet keeps the containers it already has. It still runs liveness probes and still restarts a crashed container according to the pod's `restartPolicy`, because both are local decisions. `kube-proxy` has already written iptables, nftables or IPVS rules into the node's kernel, and those rules keep forwarding traffic without any connection to the API server. So in the first minutes of an outage, users of already-running services often notice nothing. ## What stops working Everything that needs a **write to or read from the API server** stops: | Function | Owner | Effect while the API server is down | |---|---|---| | Placing new pods | kube-scheduler | Pending pods stay Pending | | Replacing pods from a failed node | kube-controller-manager | Lost capacity is not replaced | | Scaling and rollouts | Deployment and HPA controllers | Replica counts freeze | | Updating EndpointSlices | EndpointSlice controller | A pod that dies can stay in the routing rules | | Starting CronJobs | CronJob controller | Scheduled runs are missed | | Reporting status | kubelet | Pod and node status in the API goes stale | | Anything a human does | kubectl | No answers and no changes | A few of these need a closer look: 1. **Stale routing.** If a pod crashes for good, or its node dies, the EndpointSlice is not updated. kube-proxy on the other nodes keeps sending that pod a share of the connections. 2. **Missed batch work.** Say a nightly ledger-reconciliation CronJob is due at 02:00 and the control plane is down from 01:40 to 02:25. That run does not start on time. When the control plane comes back, the CronJob controller can still start the missed run, but only if the CronJob's `startingDeadlineSeconds` has not passed (or is unset). The run can therefore start late, or be skipped entirely. 3. **No self-healing beyond the node.** The kubelet restarts containers, but a pod whose node dies is not recreated anywhere else. 4. **Cluster DNS.** CoreDNS typically keeps answering from its last-synced view of Services, so existing names keep resolving. New Services do not appear. ## What happens when the control plane comes back Recovery is its own risk. Kubelets reconnect, list their pods and push a burst of status updates. Controllers re-list everything and reconcile a large backlog. The node lifecycle controller in kube-controller-manager protects against one failure mode: if every zone looks fully unhealthy (for example because no node could renew its heartbeat while the API server was down), it enters a **full-disruption** state and stops evicting pods rather than deleting everything at once. ## How to use this during an incident - Check what users actually see before calling it a total outage. The data plane is often still healthy. - Hold off on node reboots and drains, and on anything else that relies on controllers to recover, until the API server is back. - Make a list of the time-sensitive work (CronJobs, autoscaling and deployments in flight) that has to be checked once the control plane returns. - For long outages, remember the data plane slowly drifts from reality: failed nodes are not replaced and the routing to dead pods is not cleaned up.

  • If a node dies while the Kubernetes API server is down, what happens to its pods?
    Nothing is done about them. The node lifecycle controller can't see heartbeats or taint the Node, and even if it could, no controller can create replacements. The EndpointSlices still list those pod IPs, so kube-proxy on the healthy nodes keeps sending them traffic, and those connections fail. Capacity and routing are only repaired after the API server returns and the controllers reconcile.
  • Why can the return of the Kubernetes API server itself cause trouble?
    Every kubelet and controller reconnects at once, re-lists objects and pushes a backlog of status updates. That burst can overload the API server and etcd. The node lifecycle controller also sees a whole cluster of stale heartbeats. It has a full-disruption state that stops evictions when all zones look unhealthy, so the cluster doesn't evict every pod on the way back.

It is like an airport whose control tower loses power. Planes already in the air keep flying their filed routes, but no new flight gets cleared and nobody gets rerouted around trouble.

saying these in an interview costs you the question

  • All pods stop as soon as the API server goes down
  • The kubelet stops restarting crashed containers without the API server
  • Service traffic stops because kube-proxy needs the API server for every packet
  • Controllers keep replacing failed pods from their local cache
  • A missed CronJob run is always started automatically once the control plane returns