skip to content

Cluster PKI & Certificates

The cluster's own PKI: a CA issuing client and serving certificates between the API server, kubelets, etcd and controllers, with leaf certs that expire after a year. Interviewers ask because expiry breaks a healthy cluster overnight and the fix is not obvious.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

4

A kubeadm-built Kubernetes cluster ran untouched for a year and kubectl now fails with 'x509: certificate has expired'. How do you diagnose and recover it?

level: seniorimportance: must knowfreq 62%

answer

  1. clock first, then files
  2. renew on every control-plane node
  3. static pods need a manifest move
  4. admin.conf embeds the old cert
  5. expired kubelets can't self-renew

basics

~20 s

Confirm the expiry with kubeadm certs check-expiration, run kubeadm certs renew all on every control-plane node, restart the control-plane static pods so they load the new files, then copy the renewed admin.conf. Kubelets whose own certificate expired need re-bootstrapping.

solid answer

~40 s

First confirm the cause: `kubeadm certs check-expiration` on each control-plane node, or `openssl x509 -enddate` on `apiserver.crt`, and check the node clocks, since a wrong clock produces the same error. Then run `kubeadm certs renew all` on **every** control-plane node. It re-signs the one-year leaves and the client certs in `admin.conf`, `super-admin.conf`, `controller-manager.conf` and `scheduler.conf` with the existing CAs, so trust does not change. kubeadm tells you to restart kube-apiserver, kube-controller-manager, kube-scheduler and etcd, because not every component reloads every file. They are static pods, so you move each manifest out of `/etc/kubernetes/manifests` and back; deleting the mirror pod does not restart them. Copy the new `admin.conf` to your kubeconfig. Finally check the kubelets: a kubelet whose client certificate also expired cannot renew itself and needs a fresh credential.

code

bash · 7 lines
bash
sudo kubeadm certs check-expiration
sudo kubeadm certs renew all
sudo mv /etc/kubernetes/manifests/kube-apiserver.yaml /root/
sleep 30
sudo crictl ps --name kube-apiserver
sudo mv /root/kube-apiserver.yaml /etc/kubernetes/manifests/
sudo cp /etc/kubernetes/admin.conf "$HOME/.kube/config"

go deeper

for a junior

Know the error string, the command that renews kubeadm certificates, and that the control plane has to be restarted afterwards.

for a middle

Explain which files renewal touches and which it skips, and why static pods need a manifest move rather than a kubectl delete.

for a senior

Run the recovery in order across every control-plane node while keeping etcd quorum, then deal with kubelets holding expired client certificates.

for a principal

Turn the incident into policy: yearly upgrades, expiry alerts per node and a rehearsed runbook, so certificate expiry becomes routine work instead of an outage.

## What actually broke kubeadm signs every control-plane **leaf certificate** for one year, and all of them are created by the same `kubeadm init`. So on a cluster that was never upgraded, they all expire within minutes of each other. Take a 9-node bare-metal cluster running an IoT telemetry ingest gateway: 3 control-plane nodes and 6 workers. After 365 days: - `kubectl` fails: `x509: certificate has expired or is not yet valid`. - kube-controller-manager and kube-scheduler can no longer authenticate to the API server, so no rescheduling and no controller reconciliation. - kube-apiserver may lose its etcd connection (`apiserver-etcd-client`) and fail its own health checks. - Pods already running keep serving. The gateway may even absorb a 3,400 requests-per-second burst, but nothing heals when a pod or node fails. ## Step 1: diagnose 1. **Rule out the clock.** The same error appears when a node's time is wrong (`not yet valid`). Check `timedatectl` or the NTP state on each node. 2. **Read the expiry.** `sudo kubeadm certs check-expiration` works without a healthy API server; it reads the files and warns if it cannot load cluster config. You can also run `openssl x509 -noout -enddate -in /etc/kubernetes/pki/apiserver.crt`. 3. **Check the CAs** in the same output. If a CA expired too, renewal alone will not help (see CA rotation). 4. **Check the kubelets.** Inspect `/var/lib/kubelet/pki/kubelet-client-current.pem` on each node and read the kubelet log. ## Step 2: renew on every control-plane node ```bash sudo kubeadm certs renew all ``` This re-signs, with the **existing** CAs: | Renewed | Not renewed | |---|---| | `apiserver`, `apiserver-kubelet-client`, `front-proxy-client` | `ca`, `front-proxy-ca`, `etcd-ca` | | `etcd-server`, `etcd-peer`, `etcd-healthcheck-client`, `apiserver-etcd-client` | `sa.key` / `sa.pub` (not certificates) | | certs in `admin.conf`, `super-admin.conf`, `controller-manager.conf`, `scheduler.conf` | `kubelet.conf` (the kubelet manages it) | The renewal keeps the existing certificate's SANs and subject, so clients that trust the CA accept the new files immediately. The certificates are per node, so repeat on **each** control-plane node; renewing one leaves the other two broken. Renewal is also not possible for a certificate whose CA is `EXTERNALLY MANAGED`, because kubeadm does not have that CA's key. kubeadm ends with: *"You must restart the kube-apiserver, kube-controller-manager, kube-scheduler and etcd, so that they can use the new certificates."* ## Step 3: restart the static pods correctly These components are **static pods** that the kubelet runs from files in `/etc/kubernetes/manifests`. The API objects you see for them are only **mirror pods**. Deleting a mirror pod with `kubectl` does not restart the container, and `kubectl` does not work yet anyway. Instead: 1. Move a manifest (for example `kube-apiserver.yaml`) out of the directory. 2. Wait longer than the kubelet's `fileCheckFrequency` (default 20s) and confirm with `crictl ps` that the container stopped. 3. Move the manifest back and wait for the container to be running again. 4. Repeat for `kube-controller-manager`, `kube-scheduler` and `etcd`, one node at a time so etcd keeps quorum. ## Step 4: fix client access and kubelets - Copy the renewed `/etc/kubernetes/admin.conf` to `~/.kube/config`. The old file still embeds the expired certificate. - Any kubeconfig handed out earlier and signed by the cluster CA is unaffected by renewal. It expires on its own schedule. - **Kubelets**: worker kubelets normally rotated their client certificate on their own, because `rotateCertificates` is on in kubeadm. A kubelet that could not renew in time, for example because the control plane was already down or the node was off, holds an expired certificate and cannot authenticate its own renewal. Give it a fresh credential: a new kubelet kubeconfig signed by the cluster CA, or a bootstrap kubeconfig so it runs TLS bootstrap again. Then restart the kubelet. ## Step 5: make it not happen again - Upgrade at least yearly: `kubeadm upgrade apply`/`node` renews leaves by default. - Alert on residual certificate time well before expiry, per node. - Put `kubeadm certs renew all` plus the static-pod restart in a tested runbook, and rehearse it on a staging cluster. ## Timeline worth remembering On this cluster the first alert would have been `RESIDUAL TIME` dropping under 30 days, 335 days after `kubeadm init`. Nothing reported the problem after that, because the failure is silent until the moment of expiry. At expiry, the control-plane components fail within seconds of each other. Kubelets that were already rotating on their own keep working until their own deadline. The gateway pods keep running, so the outage shows up only when something needs the control plane: a pod crash, a node failure, a scaling event. That delay is why teams often find out hours later and blame the wrong change.

  • Why doesn't kubectl delete pod on the kube-apiserver mirror pod pick up the renewed certificates?
    A mirror pod is only a read-only reflection of a static pod the kubelet runs from a file. Deleting it removes the API object, and the kubelet recreates the mirror, but the container keeps running. To restart it, change or move the manifest file, or stop the container through the runtime.
  • After renewal, a colleague's old kubeconfig signed by the cluster CA months ago still works. Why?
    `kubeadm certs renew` re-signs only kubeadm's own files and never changes the CA. A client certificate issued earlier from the same CA stays valid until its own notAfter date, because the API server trusts anything that CA signed. Renewal neither revokes nor extends it.
  • How would you avoid this incident entirely on a cluster you rarely upgrade?
    Schedule `kubeadm certs renew all` well before expiry, followed by a rolling static-pod restart, and alert on residual time from each control-plane node. Better still, upgrade at least yearly: `kubeadm upgrade apply` and `kubeadm upgrade node` renew leaves by default, so the deadline keeps moving.

saying these in an interview costs you the question

  • Rebuild the cluster from scratch because the certificates expired
  • Renewing on one control-plane node fixes all of them
  • Deleting the kube-apiserver mirror pod restarts it with the new files
  • kubeadm certs renew all also rotates the CA and invalidates old kubeconfigs
  • Running workloads stop immediately when control-plane certificates expire
  • Skip checking the node clocks because the error clearly says expired
open as a page

In a kubeadm-built Kubernetes cluster, why do the control-plane certificates expire after one year, and how do you check when they expire?

level: juniorimportance: should knowfreq 52%

basics

~20 s

kubeadm creates a cluster PKI in /etc/kubernetes/pki: CAs valid for ten years and leaf certificates valid for one year, so a stolen leaf is not useful for long. Running kubeadm certs check-expiration on a control-plane node shows each expiry date.

open as a page

How does a Kubernetes kubelet get and rotate its certificates through CertificateSigningRequests, and why do rotateCertificates and serverTLSBootstrap differ at approval?

level: middleimportance: should knowfreq 44%

basics

~20 s

The kubelet submits CertificateSigningRequests to the API server and kube-controller-manager signs them with the cluster CA. Client-certificate CSRs, used with rotateCertificates, are auto-approved. Serving-certificate CSRs, used with serverTLSBootstrap, are not, so someone must approve them.

open as a page

Why does kubeadm certs renew never replace a Kubernetes cluster's CA, and how would you rotate the cluster CA without an outage?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Replacing a Kubernetes cluster CA changes what every component trusts, so kubeadm refuses to do it automatically. Rotating it without an outage means trusting old and new CAs together, re-issuing every leaf from the new CA, then removing the old CA.

open as a page