skip to content

In Kubernetes' version skew policy, how far may kubelets and kubectl drift from kube-apiserver, and why does that force control-plane-first upgrades?

level: middleimportance: must knowfreq 72%

answer

  1. anchored on the API server
  2. newer client, older server is unsafe
  3. kubelet three behind, never ahead
  4. kubectl plus or minus one
  5. controllers wait for every apiserver

basics

~20 s

Kubelets may be up to three minor versions older than kube-apiserver but never newer; kubectl may be one minor older or newer. Since nothing may be newer than the API server, it is upgraded first, then controllers and scheduler, then kubelets.

solid answer

~40 s

The skew policy is anchored on `kube-apiserver`. HA apiserver instances may differ by one minor. `kube-controller-manager`, `kube-scheduler` and `cloud-controller-manager` must not be newer than any apiserver they talk to and may be one minor older. A kubelet must not be newer than the apiserver and may be up to three minors older; `kube-proxy` follows the same bounds. `kubectl` is supported within one minor either way. The direction matters: a newer client may send fields or API versions an older apiserver does not know, so the server must learn them first. That fixes the order: every apiserver, then the controller-manager, scheduler and cloud-controller-manager, then kubelets node by node. In a kubeadm cluster, CoreDNS and kube-proxy move only after every control-plane node is done.

code

bash · 7 lines
bash
kubectl version
# Client Version: v1.35.3
# Server Version: v1.37.0
# WARNING: version difference between client (1.35) and server (1.37) exceeds the supported minor version skew of +/-1

kubectl get nodes
# the VERSION column shows each node's kubelet version

go deeper

for a junior

Remember the headline numbers: kubelet up to three minors older and never newer, kubectl one minor either way, control plane first.

for a middle

Explain the asymmetry: a newer client can send fields or API versions an older server does not know, so the server always upgrades before its clients.

for a senior

Apply it mid-rollout: measure skew from the oldest HA apiserver, hold controllers until every apiserver is done, and never let nodes sit at the edge of the window.

for a principal

Use the three-minor kubelet window deliberately to decouple disruptive node upgrades from control-plane upgrades, without letting it become permanent drift.

## What the version skew policy is A Kubernetes cluster is a set of separately versioned programs talking over the API. The **version skew policy** says which combinations of minor versions the project supports, so that a cluster can be upgraded while it keeps running. Every rule is expressed relative to `kube-apiserver`, because every other component is a client of it. ## The supported windows | Component | Allowed relative to kube-apiserver | |---|---| | Other `kube-apiserver` instances (HA) | newest and oldest within 1 minor | | `kube-controller-manager`, `kube-scheduler`, `cloud-controller-manager` | not newer; up to 1 minor older | | kubelet | not newer; up to 3 minors older | | `kube-proxy` | not newer; up to 3 minors older; up to 3 minors older or newer than the kubelet on its node | | `kubectl` | 1 minor older or newer | Two details are worth saying out loud: - **The kubelet window was two minors before Kubernetes 1.28** and has been three since. kubeadm encodes the current value: its policy constant for kubelet skew is 3, and `kubeadm upgrade plan` reports `There are kubelets in this cluster that are too old` when a target would exceed it. - **HA apiservers change the effective window.** If instances run v1.36 and v1.37 during a rollout, a controller must not be newer than the *oldest* one it might reach, and a kubelet window is measured from the oldest too. ## Why nothing may be newer than the API server The asymmetry is the whole point. An older client talking to a newer server is safe, because the server still serves the API versions and fields the client knows. A newer client talking to an older server is not: - it may send **fields the older server does not know**, which the server drops when it decodes the request, so the client's intent silently disappears; - it may call **API versions or resources** the older server does not serve and get errors; - it may rely on **behaviour** that only the newer server implements. So the server must be upgraded before any of its clients. ## The resulting upgrade order 1. **All `kube-apiserver` instances**, one at a time, staying within one minor of each other. 2. **`kube-controller-manager`, `kube-scheduler`, `cloud-controller-manager`** - only after every apiserver is done, because they must not be newer than any instance they may reach. 3. **Kubelets**, node by node, each node drained first; `kube-proxy` alongside, within its window. 4. **Clients**: `kubectl` on workstations and CI runners, then controllers built on the client libraries, checked against their own compatibility statements. kubeadm follows this order for you on control-plane nodes, and adds one rule of its own: its `addon` phase upgrades the CoreDNS and `kube-proxy` workloads only once every control-plane instance runs the new version. ## A worked example on the portal cluster A clinical-records portal runs on three kubeadm nodes, each a stacked control-plane node that also runs portal Pods. - Control plane at v1.37 with kubelets still at v1.34: **supported** - three minors behind, not newer. - Control plane at v1.36 with kubelets at v1.33, then moving the control plane to v1.37: **unsupported** - the kubelets would be four minors behind; upgrade them first to at least v1.34. - An engineer's laptop with `kubectl` v1.35 against the v1.37 cluster: **outside the window**. `kubectl version` warns that the difference `exceeds the supported minor version skew of +/-1`; most commands still work, but behaviour is no longer guaranteed. - A node whose kubelet was upgraded to v1.37 while the apiserver is still v1.36: **unsupported** - this is exactly the ordering mistake the policy forbids. ## Why the three-minor kubelet window matters operationally Upgrading a kubelet is the disruptive part - the node is drained and its Pods move. The wider window lets operators upgrade the control plane on every release while touching nodes less often, as long as the gap never exceeds three minors. It is a buffer, not a target: nodes left at the edge block the next control-plane hop. ## How to check skew in practice - `kubectl version` on every workstation and CI runner, since `kubectl` is the component most often left behind. - `kubectl get nodes` for the kubelet version of each node, in its `VERSION` column. - The image tags of the control-plane static pods in `kube-system` for each apiserver, controller-manager and scheduler instance. - `kubeadm upgrade plan`, which reports kubelets that a proposed target would leave too far behind before anything is changed.

  • In an HA control plane mid-upgrade, which apiserver version do you measure skew from?
    From the oldest instance a component might reach. Behind a load balancer a controller or kubelet can land on any apiserver, so being within the window of only the newest one is not enough. That is why `kube-controller-manager` and `kube-scheduler` are upgraded only after every apiserver instance runs the new minor.
  • How does kube-proxy's window differ from the kubelet's?
    Relative to `kube-apiserver` it is the same: not newer, up to three minors older. Relative to the kubelet on its own node it may be up to three minors older or newer. That lets kubeadm roll the `kube-proxy` DaemonSet once the control plane is finished, before every kubelet has been upgraded.
  • On a managed Kubernetes service, does the control-plane-first order still apply?
    Yes. The provider performs the control-plane step, but the skew rules are the same: node kubelets must not be newer than the control plane and must stay within three minors of it. Node pools are upgraded after the control plane, by the provider or by you.

saying these in an interview costs you the question

  • Kubelets should be upgraded first because they run the workloads
  • A kubelet may be one minor newer than the API server
  • kubectl works with any cluster version without caveats
  • Controller-manager and apiserver can be upgraded in any order
  • The kubelet window is unlimited as long as nodes report Ready
  • An older client talking to a newer API server is unsupported