skip to content

Helm 3's bug-fix support has ended. How would you sequence a move to Helm 4 across your CI images and hundreds of live releases?

level: principalimportance: should knowfreq 30%

answer

  1. Two halves: the client, and the write path
  2. The deadline is not the interesting part
  3. Nothing to convert; everything to audit
  4. Pin the version before you bump it
  5. Give the coexistence period an end date

basics

~20 s

Treat it as a client rollout, not a data migration: releases and charts are adopted in place. Pin the CLI version per image, sweep automation for changed flags and plugin wiring, roll out by blast radius, and schedule the switch to server-side apply as a separate, per-release change.

solid answer

~50 s

The programme has two independent halves and conflating them is the common mistake. The first is **replacing the client**: releases are adopted in place — same `sh.helm.release.v1.<name>.v<rev>` records, same `apiVersion: v2` charts, same commands — so the work is auditing every `helm` invocation in pipelines for renamed, removed and now value-taking flags, plus plugin installs and post-renderer wiring. Because the record format is unchanged, both clients read the same history, which makes CLI rollback a genuine escape hatch. The second half is **changing the write path**: on upgrade, `--server-side` defaults to `auto` and inherits, so old releases stay client-side indefinitely unless you move them deliberately. I would pin an exact CLI version per image rather than tracking latest, roll out in waves by blast radius, and keep a register of which releases have been flipped so the estate does not end up with two invisible write paths forever.

go deeper

for a junior

Know that the move is mostly about the tool and the pipelines around it, not about the charts, and that pinning an exact CLI version in an image is what makes a bump reversible.

for a middle

Be ready to describe the concrete audit: find every place a Helm binary runs, check each invocation against the new flag behaviour, and verify that the same chart renders identically under both clients before anything ships.

for a senior

Show the staging judgement — non-production first, stateless before stateful, one real upgrade and one real rollback per wave — and explain why the apply-path change is scheduled separately from the client bump.

for a principal

Own the decisions nobody else can make: whether the fleet converges on one apply model, how long two clients may coexist, what the migration is allowed to clean up, and how you avoid a forced move under incident conditions after support ends.

## Frame it correctly first The deadline is real — Helm 3's bug-fix support ends 9 September 2026, with 3.21.x as its final line — but the deadline is the least interesting part. What makes this a leadership question is that the migration looks small and has one genuinely risky component hidden inside it. **The small part:** Helm 4 adopts Helm 3 releases in place. There is no conversion tool because nothing needs converting: revisions are still Secrets named `sh.helm.release.v1.<name>.v<rev>`, `HELM_DRIVER` still selects the storage backend, `apiVersion: v2` is still the chart format, and the top-level command set is identical. Nobody has to touch a chart. **The risky part:** the default write path. New installs by a Helm 4 client apply server-side. Upgrades of an existing release default to `--server-side=auto`, which inherits the previous revision's method — so every release created under Helm 3 keeps using the client-side three-way merge. That is safe, and it is also a fork in the estate: after the rollout you have releases on two different apply models, and only your records tell you which is which. ## Sequence the work **Phase 0 — inventory.** Two lists. Every place a `helm` binary runs (CI images, base images, bastions, laptops, any controller or job that shells out to it), and every release, with the client version that last wrote it. `helm history` gives you the second; the first is a grep exercise across pipeline definitions and Dockerfiles. Expect the surprise to be in the first list: shared base images that half a dozen teams inherit without knowing. **Phase 1 — pin, then bump.** If images track a floating `latest` Helm, fix that before anything else. A migration you cannot roll back is not a migration. Pin an exact version per image so the bump is a reviewable commit and a revert is one line. **Phase 2 — sweep automation.** Every invocation gets read against the Helm 4 CLI: flags renamed, one list flag removed, two flags that became value-taking rather than boolean, plugin installation now verifying by default, and the post-renderer interface changed shape. Parse-time failures are the good case because they are loud; the ones to hunt for are invocations whose meaning shifted rather than breaking. A fast, cheap check is to run the pipeline's render and preview steps under the new client in a scratch namespace and diff the output against the old client's — charts should render identically, and any difference is a finding. **Phase 3 — waves by blast radius.** Non-production first, then internal-only production, then customer-facing. Within a wave, prefer stateless workloads before stateful ones: an invoice-rendering worker is a cheaper first subject than a Redis chart carrying a StatefulSet and a PersistentVolumeClaim, where an unexpected object replacement has data attached to it. Each wave is: bump the pinned version, run one upgrade of a real release, verify, then let the wave run for a period long enough that scheduled jobs and rollbacks have exercised the new client. **Phase 4 — the apply-path decision, separately.** This is the phase teams skip. Decide, per release or per class of release, whether it moves to server-side apply and when. It is not free: the first server-side upgrade of a long-lived object introduces field-manager ownership where there was none, so fields another controller or a past `kubectl edit` set can conflict. Do it as its own change, on its own schedule, with a preview and a diff — not as a side effect of a CLI bump. ## What to decide as a lead rather than delegate - **Whether the fleet converges on one apply model, or lives with two.** Two is legitimate but must be recorded, because six months later nobody remembers which releases inherited what, and the drift behaviour after a manual edit differs between them. - **How long the two clients coexist.** Both read the same records, so coexistence is technically easy and therefore tends to last forever. Give it an end date, or you will run an unsupported binary in some corner of the estate into next year. - **Whether the migration pays for cleanup.** A sweep of every `helm` invocation is a rare licence to delete plugins nobody uses, unpin charts nobody owns, and remove the pipeline that still installs a plugin for a command last run in 2023. ## How you know it worked Define done as: no image ships a Helm 3 binary; every pipeline's `helm` invocations parse and behave identically under the new client; every release has a recorded apply model; and one real rollback has been exercised under Helm 4 on a release that Helm 3 installed. The last one is the test people skip, and it is the one that proves the adoption story rather than assuming it. ## The failure modes to name A big-bang bump of a shared base image, so every team's pipeline changes on the same afternoon. A migration that quietly flips production to a new write path without anyone deciding to. And the opposite failure: a two-year "we'll do it later" tail on an unsupported client, where the eventual forced move happens under an incident rather than a plan.

  • If the migration goes wrong mid-rollout, what is your actual rollback?
    Repin the image to the previous Helm 3 version. Because the release-record format is unchanged, the older client reads the same history and can upgrade and roll back the same releases, so reverting the binary is a real option rather than a hope. The exception is any release you already flipped to server-side apply: those objects now carry field-manager ownership, so treat them as a separate, smaller rollback question rather than assuming the whole estate reverts cleanly.
  • How do you keep the estate from permanently carrying two different apply models?
    Record the model per release the moment it is decided, in whatever registry already tracks ownership of a service, and give the mixed state an expiry date the same way you would give a feature flag one. Then migrate in the same waves you used for the CLI, stateless before stateful. Without a record, the model becomes hidden state discovered only when someone hand-edits an object and the drift behaviour surprises them.
  • A team says they will just skip Helm 4 until the next major. What is your response?
    That the choice is not "upgrade or stand still" but "upgrade on a plan or upgrade during an incident". Once bug-fix support ends, a security fix or a Kubernetes-side change forces the move on someone else's timetable, and the flag and plugin sweep still has to happen — just faster and without a rehearsal. The cost of the move does not decrease with delay; only the amount of control you have over when it lands does.

Swapping the CLI is like replacing the crew's tools overnight; changing the apply model is like changing how the crew is allowed to divide up the work, and only the second one needs a rehearsal.

saying these in an interview costs you the question

  • Plans a data migration for releases that need none
  • Treats the CLI bump and the apply-path switch as one change
  • Bumps a shared base image for every team at once
  • Leaves images tracking a floating latest Helm version
  • Assumes charts must be re-authored for the new major
  • Sets no end date for running both clients

context