skip to content

Before upgrading a 64-node Kubernetes cluster, how do you find every client and manifest still using an API version the target release stops serving, and migrate them safely?

level: seniorimportance: should knowfreq 47%

answer

  1. storage survives, callers break
  2. server evidence plus source evidence
  3. a gauge that only resets
  4. audit names the user agent
  5. get -o yaml lies here

basics

~10 s

Read the removal list for each release you cross, find live callers with kube-apiserver's apiserver_requested_deprecated_apis metric and k8s.io/deprecated audit annotations, scan Git for dormant manifests, then rewrite and verify with fresh audit events.

solid answer

~50 s

Stored objects are safe: the API server keeps each resource at one storage version and converts on read, so removal only breaks clients and manifests that still name the old version. I start with the release notes for every minor release crossed. Then I find live callers on the server: `apiserver_requested_deprecated_apis` has a `removed_release` label, and audit events carry `k8s.io/deprecated` and `k8s.io/removed-release` with the user and user agent. Because that gauge is per instance and resets on restart, I also grep Git for dormant manifests such as monthly jobs. I never use `kubectl get -o yaml` as the inventory, since kubectl chooses the version it reads. I rewrite manifests, using the `kubectl convert` plugin for bulk work, upgrade client libraries in operators, check field changes with `kubectl explain --api-version`, and prove it with timestamped audit events.

code

bash · 3 lines
bash
grep -rn --include='*.yaml' 'apiVersion: autoscaling/v2beta2' deploy/ charts/
kubectl api-versions | grep '^autoscaling/'
kubectl explain hpa.spec --api-version=autoscaling/v2

go deeper

for a junior

Recall that a removed API version breaks manifests and tools that name it, while objects already in the cluster keep working.

for a middle

Explain the storage version, conversion on read, and where the deprecation warning, metric and audit annotations come from.

for a senior

Combine server-side and source-side evidence, name the metric's reset behaviour, and prove the migration with audit events before scheduling the upgrade.

for a principal

Make deprecated-API cleanliness an upgrade gate enforced in CI, and align operator client-library upgrades with the cluster upgrade calendar.

## What actually breaks when a version is removed Each Kubernetes built-in resource is written to etcd at one **storage version**, and kube-apiserver converts to whichever served version a client requests. When a release stops serving a version, stored objects are **not** deleted; they stay readable through the versions that remain. The failure lands on everything that still **names** the removed version: - manifests in Git and CI pipelines that run `kubectl apply`; - packaging tools and GitOps controllers that render or store the old `apiVersion`; - controllers, operators and scripts built on client libraries that request the old version; - dashboards and admin tooling that call the old URL directly. The historical case: `autoscaling/v2beta2` HorizontalPodAutoscaler stopped being served in v1.26. A loyalty-points accrual service whose HPA was sized for a 3,400 requests-per-second burst kept scaling after the upgrade, because the object lived on as `autoscaling/v2`. The next deploy failed with `no matches for kind "HorizontalPodAutoscaler" in version "autoscaling/v2beta2"`. The same pattern returns whenever a beta version is removed. ## Step 1: know what the target release removes Read the release notes and the deprecated API migration guide for every minor release between the current and the target version. Deprecation warnings already name the removal release, because the policy keeps a deprecated beta served for at least 9 months or 3 minor releases. ## Step 2: find live callers from the server side Server-side evidence catches clients you do not know about: 1. **Metric.** kube-apiserver exposes `apiserver_requested_deprecated_apis` with labels `group`, `version`, `resource`, `subresource` and `removed_release`. It is a gauge set to 1 the first time a deprecated version is requested, per API server instance, and it resets when that instance restarts. Query every instance, and treat it as "seen since start", not as a rate. 2. **Audit log.** Each such request carries the audit annotations `k8s.io/deprecated: "true"` and, when known, `k8s.io/removed-release`. The audit event also records the user and user agent, which is how you identify the caller: a CI service account, an old operator, a laptop. 3. **Warnings.** `kubectl` and current client libraries print the server's `Warning` header, so CI logs show lines such as `autoscaling/v2beta2 HorizontalPodAutoscaler is deprecated in v1.23+, unavailable in v1.26+; use autoscaling/v2 HorizontalPodAutoscaler`. On a managed Kubernetes service the control-plane metrics and audit logs are usually available through the provider's logging, sometimes only after you enable them. ## Step 3: find dormant callers from the source side Server-side signals only show clients that ran recently. A monthly batch job or a disaster-recovery manifest may not have called the API since the last restart. Scan the repositories as well: ```bash grep -rn --include='*.yaml' 'apiVersion: autoscaling/v2beta2' deploy/ charts/ kubectl api-versions | grep '^autoscaling/' ``` Do **not** use `kubectl get hpa -o yaml` as the inventory. kubectl asks for a version it chooses and the server converts the stored object to it, so the output says nothing about the version your Git files use. ## Step 4: migrate, then prove it | Caller | Fix | |---|---| | Manifest in Git | Rewrite to the replacement version; check field changes in the migration guide | | Many manifests | The separately installed `kubectl convert` plugin rewrites files between versions | | Controller or operator | Upgrade its client library and release it before the cluster upgrade | | Rendered templates | Update the chart or overlay, then re-render and diff | Rewriting is not always a rename. Some replacement versions changed field names or semantics, so compare the schemas with `kubectl explain <kind> --api-version=<group/version>` and apply to a staging cluster first. Prove the migration with audit events that carry timestamps: no request with `k8s.io/removed-release` for the target release during a full business cycle, including the batch windows. The gauge cannot show that traffic stopped, because it stays at 1 until the API server restarts. ## Storage versions and migration Stored objects need no manual action for a served-version removal. When a release changes a resource's **storage version**, records already in etcd keep their old encoding until something rewrites them. The `StorageVersionMigration` kind in `storagemigration.k8s.io/v1`, driven by a kube-controller-manager controller whose `StorageVersionMigrator` feature gate is GA in v1.37, rewrites every object of a resource at the current storage version. That keeps a later release that can no longer decode the old encoding from failing on old data. ## What a senior answer adds - Gate the upgrade on a clean deprecated-API report, not on a calendar date. - Fail CI when the server's `Warning` header mentions a removal release. - Pin client-library upgrades to the cluster upgrade plan so operators do not lag behind.

  • Why can apiserver_requested_deprecated_apis not prove that a migration is finished?
    It is a gauge that kube-apiserver sets to 1 the first time a deprecated version is requested and keeps until that instance restarts, per instance. After you fix the callers it still reads 1. Use audit events with timestamps, filtered on `k8s.io/removed-release`, over a full business cycle to show the requests have stopped.
  • When does a Kubernetes resource need a storage version migration rather than only a manifest rewrite?
    When a release changes the version a resource is written at. Existing etcd records keep their old encoding until rewritten, and a later release may no longer decode it. A `StorageVersionMigration` object in `storagemigration.k8s.io/v1` makes a kube-controller-manager controller rewrite every object of that resource at the current storage version.

saying these in an interview costs you the question

  • Uses kubectl get -o yaml output as the list of versions in Git
  • Believes a removed served version deletes the stored objects
  • Relies only on the gauge and misses monthly batch jobs
  • Treats every version bump as a pure apiVersion rename
  • Upgrades the cluster before operators move to new client libraries