How would you plan a fleet-wide sweep of Helm releases before a removed API version breaks their upgrades?
answer
- Cheap before removal, expensive after
- Audit records, not the chart source
- The never-upgraded release is the risk
- One clean upgrade per release, early
- Pull-request campaign, not a scripted loop
basics
~20 sInventory every release's stored manifest for the doomed group and version, then upgrade each affected release with a corrected chart while the old version still resolves. Doing it before removal turns a fleet of repairs into routine upgrades.
solid answer
~50 sThe strategy hinges on one asymmetry: **before** the API is retired, the fix for each release is a normal upgrade with a corrected chart; **after**, every affected release additionally needs its stored record repaired before it can upgrade at all. So the plan is inventory, then sweep, then freeze. Inventory with `helm get manifest` per release rather than the chart sources: the stored text is what the next upgrade must read, and lagging releases are where the two diverge. Classify: releases whose chart already renders the supported version and merely need one upgrade; releases needing a subchart bump first; and releases nobody owns. Sequence the chart bumps as reviewed changes through whatever pins them, batching by owning team. Treat never-upgraded releases as the highest risk — they look healthy and fail during the first emergency deploy. Gate the cluster work on the sweep reporting clean.
code
bash · 6 linesfor rel in $(helm list -A -o json | jq -r '.[] | .namespace + "/" + .name'); do
ns=${rel%%/*}; name=${rel##*/}
if helm get manifest "$name" -n "$ns" | grep -q 'policy/v1beta1'; then
echo "HIT $rel"
fi
donego deeper
Know that the check is per release and reads the stored manifest, and that upgrading a release early — while the old API still resolves — is what prevents the problem rather than fixing it later.
Explain why the inventory comes from helm get manifest across every release rather than from the charts in Git, and why a release that is behind on chart version is the one that shows up as a hit.
Show the triage into already-fixed, needs-a-chart-change, and unowned, and the operational gate: no cluster change until the audit reruns clean. Mention that rollback targets stay broken after a current-revision sweep.
Own the economics and the calendar — one upgrade per release before retirement versus a per-release record repair afterwards — plus the ownership split with teams that hold the pins, and making the audit a standing job because API retirement is on a schedule.
### Why this is a planning problem, not a debugging problem Take a 41-service platform chart whose subcharts are pinned and deployed by a GitOps controller. When a Kubernetes release retires a group/version those subcharts still emit, the damage is not that anything stops running — the live objects keep serving. The damage is that the **deploy path** for every affected release breaks, silently, until someone tries to use it. That is why this belongs to whoever owns delivery rather than to whoever happens to be on call the day it bites. The single fact the whole plan turns on: before the retirement, each affected release needs one ordinary upgrade with a corrected chart, and the retired version scrolls out of the current revision by itself. After the retirement, that same upgrade cannot run at all until the stored release record is repaired first — per release, by hand or by plugin, with backups and verification. One is a sprint of pull requests; the other is an incident-shaped campaign. ### Inventory from the records, not from Git Audit the stored manifests, not the chart source. What is in Git is what the *next* deploy would render; what is in the release record is what the *next upgrade will have to read*, and those two diverge exactly on the releases that matter — the ones last upgraded fourteen months ago, still on an old subchart, pinned by a controller nobody has touched since. ``` helm list -A -o json | jq -r '.[] | .namespace + "/" + .name' ``` then `helm get manifest` per release and grep for the doomed group. Two refinements are worth the effort. First, walk the revisions you would realistically roll back to, not only the current one, because rollback reads the target's manifest too. Second, record the *chart version* each release is actually running alongside the hit — that column is what turns a list of broken releases into a list of pull requests. ### Classify, then sequence Sort the hits into three piles. **Already fixed upstream.** The chart or subchart in Git already renders the supported version and the release is simply behind. These need one upgrade each and no code change — the cheapest and usually the largest pile. **Needs a chart change.** The subchart still emits the retired version. This is a real change with review, testing and a version bump, and on a 41-service platform chart it may be six or seven subcharts owned by different teams. Batch by owner and give each team the exact list of their releases rather than a policy statement. **Orphans.** Releases nobody claims. These are the ones that produce the 02:00 surprise. Decide explicitly: adopt, migrate, or uninstall — but decide, and record the decision. Sequence the sweep so every release gets one clean upgrade while the old version still resolves, then hold: no cluster change until the audit reruns clean. That gate is the deliverable, and it should be automated so it can be rerun rather than being a spreadsheet someone updated once. ### The parts people underestimate **The controller is in the loop.** A release pinned by a GitOps controller cannot be upgraded by hand in any lasting way — the pin in Git decides. The sweep is therefore a pull-request campaign with review latency, not a scripted loop, and its throughput is bounded by how fast owning teams merge. Plan calendar time accordingly. **Rollback targets are part of the blast radius.** Sweeping the current revision leaves older revisions unusable as rollback targets. Usually acceptable — a forward fix is the answer during an incident anyway — but say it out loud rather than discovering it mid-incident. **Bounded history helps you.** Helm keeps a limited number of revisions (`--history-max` defaults to 10), so releases that are upgraded regularly shed the offending revisions on their own. That is an argument for making the sweep part of a normal release cadence rather than a one-off exercise. **Make it recurring.** Beta APIs are retired on a schedule, so the audit is a standing job, not a project. Wire the grep into the delivery pipeline so a rendered manifest naming a retired group fails its own change, and keep the fleet audit on a timer. The organisations that get hurt by API removals are not the ones that lack a fix; they are the ones that only look after the upgrade.
- Why audit stored manifests rather than the chart sources in Git?Git tells you what the next deploy would render; the release record tells you what the next upgrade has to read. They differ precisely on releases that are behind — old chart version, never upgraded, pinned and forgotten. Those are the ones that break, so the record is the authoritative inventory and the chart source is only the fix list.
- How do you handle releases whose chart is pinned by a GitOps controller?The pin in Git decides, so a manual upgrade does not stick. Treat the sweep as a pull-request campaign: bump each pin to a chart version rendering the supported API, route it to the owning team, and measure progress by merged pins rather than by commands run. Its throughput is review latency, which is what you plan calendar time around.
- What is left unfixed after the sweep, even when the audit reports clean?Older revisions. Sweeping the current revision of every release leaves earlier ones still carrying the retired version, so they remain unusable rollback targets until history rolls past them. That is normally acceptable — a forward fix is the incident answer regardless — but it should be stated, not discovered while someone is trying to roll back.
saying these in an interview costs you the question
- Audits chart sources instead of stored release manifests
- Plans the sweep after the API is already gone
- Ignores releases nobody has upgraded in a year
- Assumes hand-run upgrades stick under a controller
- Treats it as a one-off project, not a recurring audit