skip to content

Before upgrading a Kubernetes cluster, how do you find clients still calling an API version the target release removes, and what actually breaks afterwards?

level: seniorimportance: must knowfreq 56%

answer

  1. callers break, stored objects don't
  2. a gauge with removed_release
  3. audit annotation says who
  4. static scan for rare jobs
  5. kubectl get output misleads

basics

~10 s

Watch the apiserver_requested_deprecated_apis metric and audit events annotated k8s.io/removed-release, and scan manifests. After the upgrade, stored objects survive, but every manifest, pipeline or controller still using the removed version fails.

solid answer

~50 s

Stored objects are not the problem: the API server keeps them and serves them at the surviving version. What breaks is every *caller* of the removed version - manifests in Git, deploy pipelines, scheduled jobs, and controllers built on old client libraries - which start failing with `no matches for kind ... in version ...`, possibly weeks later when a rarely run job fires. To find them before the upgrade, use the API server: the `apiserver_requested_deprecated_apis` gauge, labelled with group, version, resource and `removed_release`, says *what* is being called; audit events annotated `k8s.io/deprecated` and `k8s.io/removed-release` say *who*, through user and user agent. Kubectl also prints deprecation warnings. Add a static scan of the manifest repository, fix callers with the `kubectl-convert` plugin or by hand, and watch long enough to cover monthly jobs, because the metric is per apiserver instance and resets on restart.

code

bash · 8 lines
bash
# what is being called (repeat against each apiserver instance)
kubectl get --raw /metrics | grep apiserver_requested_deprecated_apis

# who is calling it (path set by kube-apiserver --audit-log-path)
jq -c 'select(.annotations["k8s.io/removed-release"] != null)
  | {user: .user.username, agent: .userAgent, uri: .requestURI,
     removed: .annotations["k8s.io/removed-release"]}' \
  /var/log/kubernetes/audit/audit.log | sort | uniq -c

go deeper

for a junior

Recall that removed API versions stop being served, so old manifests fail with a no matches for kind error.

for a middle

Explain that stored objects survive and are served at the new version, while every caller of the old version breaks.

for a senior

Show the detection toolkit - deprecated-API gauge, audit annotations, client warnings, static scan - and its blind spots: per-instance, reset on restart, rare jobs.

for a principal

Make removal readiness a standing gate before every minor upgrade, with owners for each caller and operator compatibility checked up front.

## What removal means A Kubernetes API version such as `batch/v1beta1` goes through a deprecation period and is then **removed** in a specific release: from that release on, the API server no longer serves it. `batch/v1beta1` CronJob, for example, was removed in Kubernetes 1.25 in favour of `batch/v1`. The deprecation rules themselves belong to API versioning; what matters for an upgrade is the consequence. ## What survives and what breaks | Thing | After the removing upgrade | |---|---| | Objects already stored in the cluster | kept; served at the surviving version | | `kubectl get` output | unchanged; it shows the preferred version anyway | | Manifests in Git still declaring the old version | `kubectl apply` fails | | Deploy pipelines and scheduled jobs applying them | fail when they next run | | Controllers and operators built on old client libraries | requests to the removed version fail | | Scripts calling REST paths with the old version | receive not-found errors | The failure message from `kubectl apply` names the problem: `no matches for kind "CronJob" in version "batch/v1beta1"`. The trap is timing. The upgrade itself looks clean, because running workloads keep running. The break arrives when a caller next runs - a monthly records-export pipeline, a quarterly disaster-recovery rehearsal - long after the change window closed. ## Finding callers before the upgrade The API server already knows who calls deprecated versions. Use four sources together: - **The `apiserver_requested_deprecated_apis` metric.** A STABLE gauge set to 1 for each deprecated group, version, resource and subresource that has been requested, with a `removed_release` label naming the release that removes it. It answers *what* is called. It is kept per apiserver instance and starts empty after a restart, so read every instance and cover a representative period. - **Audit events.** Requests to deprecated versions carry the annotation `k8s.io/deprecated` set to `"true"`, and `k8s.io/removed-release` when a removal release is set. Each event records `user.username` and `userAgent`, which answers *who*. - **Warnings to clients.** The API server returns a warning such as `batch/v1beta1 CronJob is deprecated in v1.21+, unavailable in v1.25+; use batch/v1 CronJob`, which `kubectl` prints. Useful in CI logs, easy to ignore. - **A static scan** of the manifest repository and rendered templates, because a caller that has not run during the observation window never shows up in the metric. ## A worked plan for the portal The clinical-records portal's manifest repository holds 612 objects, including custom resources for its operators, deployed to a three-node kubeadm cluster. 1. Read the target release's deprecation guide and list every API version it removes. 2. Scrape the metric from each of the three apiservers and filter by `removed_release` equal to the target minor. 3. Query the audit log for `k8s.io/removed-release` over at least one full monthly cycle, and group by user agent. 4. Scan all 612 objects for the removed group-versions; rewrite them with the `kubectl-convert` plugin or by hand, and review the result - a newer version can rename or restructure fields. 5. Check each operator's supported Kubernetes range; a controller that still calls a removed version must be upgraded first. 6. Re-check the metric after the fixes have run, then schedule the upgrade. Chart-level fixes for packaged releases belong to the packaging tool's own upgrade workflow; this plan is about proving that the cluster no longer sees the calls. ## Why kubectl get is not evidence A common false reassurance: `kubectl get cronjob -o yaml` shows `apiVersion: batch/v1`, so everything must be fine. It is not proof. The API server serves the object at whatever version the client requests, and kubectl requests the preferred one - the output says nothing about the version your pipeline sends. ## After the upgrade Rolling the control plane back is not a practical fix for a missed caller; fixing the caller is. Keep the audit query as a standing check before every minor upgrade, since each release can remove something. ## Common mistakes - **Watching for one day.** Rare callers never appear in so short a window. - **Checking only one apiserver.** Each instance keeps its own gauge. - **Blind find-and-replace of apiVersion.** A newer version can move or rename fields, so the converted manifest must be validated, for example with a server-side dry run against the upgraded staging cluster.

  • Why is a week of metric data possibly not enough before the upgrade?
    The gauge only shows versions requested since that apiserver instance started, and only on that instance. Callers that run monthly or quarterly, such as a records-export pipeline or a recovery rehearsal, may not have run yet. Cover a full business cycle and add a static scan of the manifests so rare callers are caught.
  • Does rolling the control plane back fix a pipeline broken by a removed API version?
    Not in practice. Rolling a control plane back a minor is risky and rarely the right call; the pipeline would break again at the next attempt. Fix the caller: update the manifest's apiVersion and any restructured fields, then re-run it.

saying these in an interview costs you the question

  • Objects created with a removed version are deleted by the upgrade
  • kubectl get showing the new apiVersion proves manifests are migrated
  • If the upgrade completes cleanly, no API removal affected us
  • The deprecated-API metric keeps its history across apiserver restarts
  • Only YAML files matter; controllers never call removed versions
  • Changing apiVersion alone is always enough to migrate a manifest