skip to content

Midway through upgrading a 48-node kubeadm Kubernetes cluster, kubectl stops answering entirely. How do you diagnose the control plane from a control-plane node using crictl and the static-pod manifests?

level: seniorimportance: should knowfreq 45%

answer

  1. static pods need no API server
  2. kubelet first, then containers
  3. crictl ps -a shows the loop
  4. livez includes the etcd check
  5. three members, quorum of two

basics

~20 s

SSH to a control-plane node. Check the kubelet, then use crictl ps -a and crictl logs on kube-apiserver and etcd, and inspect /etc/kubernetes/manifests. Check etcd quorum first, because an apiserver restart loop is often an etcd symptom.

solid answer

~40 s

With kubectl gone, work on the control-plane node itself. On kubeadm the control plane runs as **static pods**: the kubelet runs whatever manifests are in `/etc/kubernetes/manifests`, with no API server involved. First check that the kubelet is up with `systemctl status kubelet` and `journalctl -u kubelet`. Then `crictl ps -a` shows every container, including exited ones and their restart attempts, and `crictl logs` shows why kube-apiserver or etcd died. The apiserver's `/livez` includes an etcd check, so a lost etcd quorum makes the kubelet restart the apiserver over and over. Check etcd before blaming the apiserver: `etcdctl endpoint health --cluster` and `endpoint status`. In a mid-upgrade incident, compare the manifests with kubeadm's backup copies, and look for a member that was already unhealthy before this node's etcd restarted.

code

bash · 5 lines
bash
sudo systemctl status kubelet --no-pager
sudo journalctl -u kubelet --since '15 min ago' | tail -n 50
sudo crictl ps -a --name 'kube-apiserver|etcd'
sudo crictl logs --tail 40 $(sudo crictl ps -a --name etcd -q | head -n 1)
ls -l /etc/kubernetes/manifests

go deeper

for a junior

Remember that kubeadm control-plane components are static pods defined by files in /etc/kubernetes/manifests, which the kubelet runs on its own.

for a middle

Explain how crictl talks to the container runtime directly, and how a static pod differs from its mirror pod in the API.

for a senior

Walk the sequence: kubelet, container restarts, logs, etcd health and quorum, manifest diff. Show why an apiserver loop can be an etcd symptom.

for a principal

Discuss upgrade guardrails: checking etcd health before each control-plane step, alerting on member health, and practising node-level access before the incident.

## The scenario A 48-node cluster has a **stacked HA control plane**: three control-plane nodes behind a load balancer, each running kube-apiserver, kube-controller-manager, kube-scheduler and an etcd member as **static pods**. During `kubeadm upgrade apply` on the first control-plane node, kubectl suddenly returns connection refused and timeouts through the load balancer. Workloads are still serving, but nobody can see or change anything. ## Why the node is the right place to look A **static pod** is defined by a file in the kubelet's static-pod directory. On kubeadm that is `/etc/kubernetes/manifests`, which holds `kube-apiserver.yaml`, `kube-controller-manager.yaml`, `kube-scheduler.yaml` and `etcd.yaml`. The kubelet re-reads that directory periodically (every 20 seconds by default) and runs whatever it finds **without the API server**. The copies you see through kubectl are only **mirror pods** (they carry the `kubernetes.io/config.mirror` annotation), and without an API server you cannot see them anyway. So the tools are the node's own: | Tool | What it tells you | |---|---| | `systemctl status kubelet`, `journalctl -u kubelet` | Whether the kubelet runs and why it rejects a manifest | | `crictl ps -a` | Every container including exited ones, with attempt counts | | `crictl logs <id>` | The component's own error output, including a previous attempt | | `crictl pods` | Pod sandboxes, showing whether a static pod exists at all | | `/etc/kubernetes/manifests/*.yaml` | The exact spec the kubelet is running | ## A triage sequence 1. **Is the kubelet alive?** If the kubelet is down, nothing in the manifest directory is being run or restarted. A broken kubelet configuration after an upgrade is a classic cause. 2. **Which containers are restarting?** Run `crictl ps -a --name kube-apiserver` and `crictl ps -a --name etcd`. A climbing attempt count shows which component is in a loop. 3. **Read the logs of the last failed attempt.** Look for flag errors (a mistyped flag in an edited manifest), certificate errors, or failures to reach etcd. 4. **Check etcd before the apiserver.** kube-apiserver's `/livez` includes an **etcd health check**. When etcd has no quorum, the apiserver's liveness probe fails and the kubelet keeps restarting it, so the apiserver loop is a **symptom**. Query etcd with the certificates kubeadm created under `/etc/kubernetes/pki/etcd/`. 5. **Do the quorum arithmetic.** Three members need two for quorum. The upgrade restarts this node's etcd member. If a second member was already unhealthy (a full disk, a failing node), only 1 of 3 is available, below quorum, and every apiserver in the cluster fails at once. That is exactly why kubectl through the load balancer went dark. 6. **Compare the manifests.** kubeadm keeps the previous manifests in a timestamped `kubeadm-backup-manifests` directory under `/etc/kubernetes/tmp`, and rolls back automatically if a new component fails to come up during the upgrade. A manual `diff` shows what changed. ## Restoring service - **Fix the unhealthy etcd member first** (disk space, the process, the node) so the cluster gets back to two healthy members. The apiservers then come back without being touched. - **To restart a static pod**, move its manifest out of the directory, wait for the kubelet to stop the pod, then move it back. `crictl stop` alone works too, but the kubelet simply restarts the same spec. - **Never keep editing manifests in place with a backup file in the same directory.** The kubelet runs every manifest file it finds there, so a stray copy can start a second, conflicting pod. - **Large databases are slow.** With a 4.2 GB etcd database, a member that rejoins may need a long snapshot transfer from the leader. Do not read the delay as a new failure. - **Restoring from a snapshot** is a last resort when quorum cannot be recovered. That procedure belongs to the backup-and-restore topic. ```bash ETCDCTL_API=3 etcdctl \ --endpoints=https://127.0.0.1:2379 \ --cacert=/etc/kubernetes/pki/etcd/ca.crt \ --cert=/etc/kubernetes/pki/etcd/healthcheck-client.crt \ --key=/etc/kubernetes/pki/etcd/healthcheck-client.key \ endpoint status --cluster -w table ```

  • The etcd database on this cluster is 4.2 GB and apiserver latency is high after recovery. How would you defragment it safely?
    Defragmentation rewrites a member's database file and blocks that member while it runs, which can take a while at 4.2 GB. Do one member at a time and leave the leader for last, and check endpoint status and health between members so quorum is never at risk. Take a snapshot first, and schedule it away from heavy write periods such as the nightly ledger batch.
  • Why does the kube-apiserver container keep restarting when etcd is the thing that is broken?
    kubeadm gives the apiserver static pod a liveness probe on /livez, and kube-apiserver includes its etcd health check in /livez. Without etcd quorum that check fails, the probe fails, and the kubelet kills and restarts the container. The restart count points at the apiserver, but its logs will show the etcd connection failures.
  • The kubelet itself is running, but the kube-apiserver static pod never appears in crictl. What do you check?
    Check the kubelet journal for errors parsing that manifest: bad YAML, an invalid field, or a duplicate pod name from a stray file in the directory. Confirm the kubelet's staticPodPath still points at /etc/kubernetes/manifests after the upgrade. Then check that the image can be pulled on this node, because a missing control-plane image leaves only a sandbox or nothing at all.

saying these in an interview costs you the question

  • Without kubectl there is no way to inspect the control plane
  • Static pods are created by the API server, so they cannot run while it is down
  • An apiserver restart loop always means the apiserver itself is broken
  • Restart etcd members all at once to recover faster
  • Keep a backup copy of a manifest inside /etc/kubernetes/manifests