Your managed Kubernetes cluster's minor version is nearing the end of the provider's support window. What happens next, and how do you keep a sustainable upgrade cadence?
answer
- three minors a year
- latest three patched, about 14 months
- end of window: provider picks the time
- calendar, staging first, maintenance windows
- PDBs and headroom for node replacement
basics
~20 sPast the window the version stops getting patches, and most providers either upgrade the control plane for you or charge for extended support. Plan roughly three minor upgrades a year, rehearsed on a non-production cluster, before the provider forces one.
solid answer
~40 sUpstream Kubernetes ships about three minor releases a year and patches only the latest three, each for roughly 14 months, so every managed service has a support window per minor version. When yours ends, patches stop, and the provider typically either upgrades the control plane automatically on its own schedule or moves you to a paid extended-support period; either way you have lost control of timing. A sustainable cadence treats upgrades as routine: a calendar that tracks each version's end date, a non-production cluster that moves first, maintenance windows so a forced change lands off-peak, workloads with PodDisruptionBudgets and spare capacity so provider-driven node replacements are safe, and an upgrade runbook that checks for removed API versions. The goal is to always upgrade by choice, well before the deadline, never at it.
code
bash · 5 lines# Control-plane (server) and client versions
kubectl version
# Kubelet version of every node, to spot nodes lagging behind
kubectl get nodes -o custom-columns=NAME:.metadata.name,KUBELET:.status.nodeInfo.kubeletVersiongo deeper
Know that each Kubernetes minor version is only supported for a limited time and that managed clusters must be upgraded regularly.
Explain the upstream cycle (about three minors a year, latest three patched) and what a provider typically does when a version's window ends.
Lay out a cadence: version calendar, staging first, maintenance windows, PDBs with headroom and an API-removal check, so no upgrade is forced.
Account for upgrade work per cluster when sizing the fleet, and decide when extended support is a justified bridge rather than a habit.
## Why every managed version has an expiry date Upstream Kubernetes releases a new **minor version** (1.36, 1.37, ...) roughly three times a year. The project maintains release branches only for the **most recent three minor versions**, and each gets roughly **14 months** of patch releases: about a year of support plus a short period to upgrade. After that, security and bug fixes stop upstream. A managed provider builds its offer on top of that. Each minor version it supports has a published **standard support window**, often aligned with upstream, and some providers add a paid **extended support** period. The provider also applies **patch versions** (1.37.2 to 1.37.3) on its own, usually inside maintenance windows you configure. ## What happens at the end of the window The exact behaviour is provider-specific, but the pattern is consistent: 1. **Notices** start weeks or months ahead in the console, by email or as events. 2. **Patches stop** for that minor version, or continue only for an extra fee. 3. **The provider acts.** Common outcomes are an automatic control-plane upgrade to the next minor version, automatic node upgrades that follow it, or a move into extended support with a higher per-cluster price. 4. **Support narrows.** New features and new node images target newer versions only. The real cost is not the upgrade itself; it is that the provider now chooses the **time**. A forced control-plane upgrade during a sales peak, followed by node replacements that restart every replica of a memory-heavy service, is an incident you scheduled by not scheduling. ## The shape of a sustainable cadence | Practice | What it buys | |---|---| | A version calendar per cluster | Nobody discovers an end date from a notice | | Non-production clusters first, on an earlier channel | Breakage shows up where it is cheap | | Maintenance windows and exclusions | Provider-driven changes land off-peak | | PodDisruptionBudgets and spare capacity | Rolling node replacement does not take a service to zero | | An upgrade runbook with an API-removal check | Manifests using removed API versions are found before the jump | | Automation for the add-ons you own | CNI, ingress, operators move with the cluster | A few rules make the calendar realistic: - **Budget two to three minor upgrades per cluster per year.** With three releases a year and a window of about 14 months, a cluster that upgrades only when forced always sits on the oldest supported version and has no slack. - **Never skip a planned upgrade.** Upgrades move one minor version at a time, so a missed cycle doubles the next one; the ordering and skew rules themselves belong to the upgrade runbook. - **Count clusters.** The work scales per cluster, not per node, which is why fleet size matters when you decide how many clusters to run. - **Own your add-ons.** The provider upgrades its control plane; operators and controllers you installed may not support the new version yet. ## A worked example A team runs a recommendation-model inference server with a 2.6 GiB memory limit per replica on a managed cluster. The current minor version reaches end of standard support in seven weeks. The plan: 1. Week 1: upgrade the staging cluster's control plane, then its nodes; run the inference server's load test. 2. Week 2: confirm the PodDisruptionBudget allows one replica down and that nodes have room for a replacement 2.6 GiB replica during the rolling node update. 3. Week 3: upgrade production inside a configured maintenance window, off-peak. 4. Weeks 4 to 7: buffer for a rollback of the workload or a support case, not for the upgrade itself. ```yaml apiVersion: policy/v1 kind: PodDisruptionBudget metadata: name: recsys-inference spec: maxUnavailable: 1 selector: matchLabels: app: recsys-inference ``` Control-plane upgrades are generally one-way: a provider rarely lets you downgrade a minor version, which is why staging goes first. ## Signals you are falling behind - Production runs the oldest version the provider still supports. - The last upgrade was driven by a provider notice. - Add-on versions are pinned with no owner. - Extended-support charges appear on the bill. ## Why interviewers ask A managed control plane removes the mechanics of upgrading it, but not the obligation. Interviewers look for a candidate who treats the provider's calendar as a hard constraint and builds a routine around it, rather than someone who assumes "managed" means "never think about versions".
- Is paying for extended support a reasonable strategy?As a bridge, yes: it buys months when an upgrade is blocked by a real dependency, such as an operator that does not yet support the new version. As a policy, no: it raises the per-cluster bill, keeps you on an old release with a shrinking ecosystem and only postpones a larger multi-version catch-up. Track every extended-support cluster with an owner and an exit date.
- Why run non-production clusters on an earlier release channel or upgrade schedule?Because control-plane minor upgrades are rarely reversible on a managed service. Moving staging first surfaces removed API versions, add-on incompatibilities and behaviour changes where rollback is not needed. It also gives the team a rehearsed runbook, so production upgrades become routine rather than an event.
- How do maintenance windows and maintenance exclusions change the risk of provider-driven upgrades?A maintenance window limits automatic control-plane and node changes to hours you choose, so they land off-peak with people available. An exclusion blocks automatic changes for a period, such as a sales peak. Neither removes the obligation: exclusions are generally bounded and cannot hold a cluster on a version past its end of support, so they shift timing inside the window rather than extending it.
saying these in an interview costs you the question
- A managed service keeps an old version patched for as long as I want, free.
- Upgrading only when the provider forces it is a fine strategy.
- The provider upgrades my installed operators and controllers too.
- I can always downgrade the control plane if the upgrade breaks something.
- Upgrade effort scales with node count, not with the number of clusters.