skip to content

How do you repair a Helm release whose stored manifest names a removed apiVersion, without deleting the running workload?

level: seniorimportance: should knowfreq 44%

answer

  1. The workload is healthy; the record is not
  2. Rewrite the record, or re-adopt the objects
  3. Ownership metadata makes adoption possible
  4. Repair the past, bump the chart for the future
  5. Pause a controller before repairing by hand

basics

~20 s

Either rewrite the retired apiVersion inside the stored release records, or delete the release history and reinstall under the same name and namespace so Helm adopts the live objects. Pair either repair with a chart bump.

solid answer

~50 s

There are two repairs that do not touch the workload. **Rewrite the record**: edit the retired `apiVersion` inside the stored manifest of the revisions Helm will read — at minimum the current one — either with a plugin built for this (the Helm project's `mapkubeapis`, checking it supports your Helm major version) or by hand on the `sh.helm.release.v1.<name>.v<rev>` Secret. History and rollback targets survive. **Re-adopt**: delete that release's stored records and run `helm upgrade --install` with the same release name and namespace; Helm adopts the existing objects because they already carry `app.kubernetes.io/managed-by: Helm` and matching `meta.helm.sh/release-name` and `release-namespace` annotations, and `--take-ownership` (Helm 3.17+) covers objects whose metadata does not match. You lose history and pruning of orphans. Uninstall-and-reinstall is the last resort — it deletes the workload. Bump the chart either way, or the next render writes the retired version back.

code

yaml · 8 lines
yaml
metadata:
  name: fraud-scoring
  namespace: payments
  labels:
    app.kubernetes.io/managed-by: Helm
  annotations:
    meta.helm.sh/release-name: fraud-scoring
    meta.helm.sh/release-namespace: payments

go deeper

for a junior

Know that the running objects are healthy and that the repair targets Helm's record, not the workload. Recognise that helm uninstall would delete the live resources and is not the first answer here.

for a middle

Explain the two non-destructive routes — rewriting the retired apiVersion in the stored records, or deleting the records and reinstalling so Helm adopts the objects — and name the ownership label and annotations that make adoption possible.

for a senior

Weigh the routes on what each costs: history and rollback targets versus procedural risk, plus the orphaned-resource trap when Helm has no previous manifest to prune against. Show that you back up records before editing and verify with helm get manifest afterwards.

for a principal

Own the sequencing across a fleet and the delivery system around it: pausing a controller that would re-run the failure, landing the chart bump in Git so the repair is not undone, and deciding which repair becomes the team's standard runbook.

### The constraint The live objects are fine and must stay up — a fraud-scoring endpoint cannot take an outage to fix a bookkeeping problem. What is broken is Helm's own record: the manifest stored with the current revision names an `apiVersion` the cluster no longer serves, so Helm cannot build objects from it, and every `helm upgrade` and `helm rollback` aborts. Any repair therefore has to change what Helm holds, not what the cluster runs. ### Option 1 — rewrite the stored record in place The most surgical fix replaces the retired group/version string inside the stored manifest of the revisions Helm will actually read. The Helm project publishes a plugin for exactly this job, `mapkubeapis`, which walks a release's records and rewrites known retired APIs in place; check that the build you install supports your Helm major version, since Helm 4 rebuilt the plugin format (`plugin.yaml` now declares `apiVersion: v1`, a `type` and a `runtime`, and `helm plugin install --verify` defaults to true). Doing it by hand means fetching the `sh.helm.release.v1.<name>.v<rev>` Secret, decoding it, editing the manifest inside, and writing it back. It is entirely doable and entirely unforgiving — a botched record is worse than the error you started with, so snapshot the Secret first. Which revisions? At minimum the current one, because that is what the next upgrade reads. If you want rollback to remain usable you have to repair every revision you might roll back to, since rollback reads the target's manifest as well. **Keeps:** the workload, the full history, rollback targets, and the release's identity. **Costs:** the fiddliest procedure of the three, and it must be repeated per release. ### Option 2 — drop the history and re-adopt Helm will adopt a pre-existing object into a release when that object already carries the ownership metadata Helm writes: the label `app.kubernetes.io/managed-by: Helm` plus the annotations `meta.helm.sh/release-name` and `meta.helm.sh/release-namespace` matching the release you are installing. Objects installed by Helm in the first place already have all three. So: 1. delete the release's stored records (`sh.helm.release.v1.<name>.v*`); 2. run `helm upgrade --install <same name> -n <same namespace>` with a chart that renders supported API versions. Helm finds no history, treats it as a fresh install, discovers the objects already exist, checks their ownership metadata, and takes them over. Nothing restarts. Since Helm 3.17, `--take-ownership` extends this to objects whose ownership metadata does *not* match — useful when some resources were created outside the original release. **Keeps:** the workload and a clean, correct record going forward. **Costs:** the entire revision history and every rollback target; and because Helm now believes nothing existed before, anything the old release owned that the new chart no longer renders is silently orphaned rather than deleted. Inventory those before you delete the records. ### Option 3 — uninstall and reinstall Correct only when the workload can be recreated. `helm uninstall` deletes the objects; you take an outage and lose anything not backed by durable storage. `helm.sh/resource-policy: keep` on the pieces you cannot lose makes uninstall skip them — but then you are re-adopting those objects anyway on the way back in, which is option 2 with extra steps and a partial outage. ### The wrinkle when a controller owns the release If the chart version is pinned by a GitOps controller, the controller is going to re-run the failing upgrade on its interval and re-fail it, and it may fight your repair or reinstate the old pin. Suspend reconciliation for that release before you start, do the repair, then land the chart bump in Git so the controller's next run is the corrected one. Repairing the record while Git still pins the old chart just means the retired `apiVersion` is rendered straight back into the next revision. ### The pairing rule Whichever repair you pick, it fixes the past. The chart bump fixes the future. Ship both, verify with `helm get manifest <release> -n <ns>` that the current record no longer contains the retired group, and only then hand the release back to normal delivery.

  • What do you lose by deleting the release records and reinstalling, beyond the history?
    Pruning. Helm decides what to delete by comparing the new render against the previous stored manifest; with no previous manifest there is nothing to compare, so any resource the old release owned that the new chart no longer renders is left running and unmanaged. Inventory the old manifest before you delete the records, and clean up the orphans by hand.
  • If you repair only the current revision's record, what is still broken?
    Rollback. A rollback builds objects from the target revision's stored manifest as well as the current one, so any earlier revision you did not repair remains an unusable rollback target. Either repair every revision you would realistically roll back to, or accept that the release's recovery path is a forward fix until the old revisions age out of history.
  • Why is uninstall with helm.sh/resource-policy: keep not a clean shortcut here?
    It stops Helm deleting the annotated objects, but the release record still goes away, so on reinstall you are adopting those objects anyway — which is the re-adopt path with an extra uninstall and a partial outage for everything you did not annotate. If adoption is where you end up, go there directly.

saying these in an interview costs you the question

  • Reaches for helm uninstall on a live production release
  • Thinks helm rollback repairs the stored record
  • Believes a chart bump alone finishes the repair
  • Deletes release Secrets without a backup
  • Forgets a GitOps controller will re-run the bad upgrade
  • Assumes adoption works without the ownership metadata

context