skip to content

What happens to a Helm upgrade when its pre-upgrade migration Job fails?

level: middleimportance: must knowfreq 62%

answer

  1. The release stops where the hook stopped
  2. Which manifests had been applied at that moment
  3. Helm tracks objects, not rows
  4. Restoring manifests is not restoring data
  5. The timeout ends the wait, not the Pod

basics

~20 s

Helm aborts the upgrade. None of the chart's own manifests are applied, the old pods keep serving, and the attempt is recorded as a failed release. The schema changes the Job already made stay: Helm never undoes a hook's side effects.

solid answer

~40 s

A hook failure is a release failure. For a `pre-upgrade` hook the chart's manifests have not been applied yet, so the workload is untouched and the previous version keeps serving — the release simply ends marked failed. Two things bite. First, Helm's wait on the hook is bounded by `--timeout`; when it expires Helm calls the release failed but Kubernetes keeps the migration Pod running, so writes may still be in flight. Second, Helm's model is Kubernetes objects, not data: whatever the migration already committed stays committed, and `--rollback-on-failure` (Helm 4's name for what was `--atomic`) restores the previous revision's *manifests* only. Contrast a `post-upgrade` hook failure, where the new manifests are already live when the release is marked failed.

code

bash · 4 lines
bash
helm upgrade myrelease ./platform-chart --timeout 20m
# Error: UPGRADE FAILED: pre-upgrade hooks failed
kubectl logs job/myrelease-migrate-27
helm history myrelease

go deeper

for a junior

Remember the headline: a failed hook fails the whole upgrade, and with a pre-upgrade hook the application is never touched, so the old version keeps serving. Fixing the cause and re-running the upgrade is the recovery.

for a middle

Explain the ordering that produces that outcome — release record, pre-upgrade hooks, chart manifests, post-upgrade hooks — and why Helm restores manifests but never data. Know that --rollback-on-failure is Helm 4's name for --atomic.

for a senior

Show that you have lived with the timeout: Helm stops waiting, Kubernetes does not stop the Pod, and a hasty retry runs two migrations at once. Be able to say what evidence survives on the cluster and for how long.

for a principal

Own the policy question: given that a failed migration leaves a half-changed schema Helm cannot repair, decide where migrations are allowed to run inside the release at all, and what the standard timeout and retry rules are across the estate.

### The sequence Helm actually runs When you run `helm upgrade`, Helm renders the chart, writes a release record for the attempt into the namespace (a Secret named `sh.helm.release.v1.<name>.v<rev>`, in a pending state), then applies the `pre-upgrade` hooks, waits for them, applies the chart's own manifests, and finally applies the `post-upgrade` hooks. A hook that ends in failure stops that sequence where it stands and the release attempt is marked failed. Where it stops decides everything about the blast radius. ### A failed pre-upgrade hook At the moment a `pre-upgrade` migration Job fails, Helm has not applied a single one of the chart's own manifests. The Deployment in the cluster is still the previous revision's Deployment; the pods serving traffic are still running the old image. From the workload's point of view nothing happened at all. Helm exits non-zero, and the attempt shows in the release history as a failed revision. This is the property that makes the pattern worth using: a migration that cannot complete stops the code that depends on it from ever starting. The failure mode is a stalled deploy, not a crash loop of new pods talking to an old schema. ### What Helm does not do Helm's model of the world is Kubernetes objects it applied and can re-apply. It has no model of the rows a hook's Pod changed. If the migration added two of five columns and then failed, those two columns are still there when Helm returns. There is no automatic compensation, and no flag adds one. The flag people reach for here is `--rollback-on-failure`. In Helm 4 that is the real name; `--atomic` survives as a deprecated alias for it, and setting it also defaults `--wait` to the `watcher` strategy. What it does is restore the previous revision's manifests when a release fails. For a pre-upgrade hook failure that is close to a no-op, because nothing new was applied — and in every case it restores manifests, never data. Treating `--rollback-on-failure` as a database safety net is the single most common misconception on this topic. ### The timeout is not a kill Helm waits for the hook Job to reach a terminal state, and that wait is bounded by `--timeout` (five minutes by default). When the timeout expires Helm stops waiting and calls the release failed. Kubernetes was never told anything: the Job keeps running and its Pod keeps writing. So a "failed" release can sit alongside a migration that goes on to succeed a few minutes later, and a retry launched immediately can put a second migration Pod next to the first. Any migration longer than the timeout needs the timeout raised — and needs a lock, because the assumption of a single runner no longer holds. ### The failed Job object stays Helm does not clean up a hook that failed. The Job object and its Pod are left in the namespace, which is deliberate: the Pod's logs are how you find out what went wrong. What eventually removes it is the hook's delete policy, which by default deletes the previous hook object immediately before creating the new one on the next attempt. That is worth knowing precisely because it means the evidence is available until the next upgrade, and no longer. ### Contrast: a post-upgrade hook failure The same rule — hook failure equals release failure — has a very different meaning at the other end. A `post-upgrade` hook runs after the chart's manifests have been applied, so when it fails the new Deployment is already rolling out. Helm marks the release failed while the cluster is running the new version. A smoke test placed there is therefore a detector, not a gate; only a `pre-` hook is a gate. ### Recovering There is no special command. You fix the cause and run `helm upgrade` again, which re-renders and re-runs the hook. That is why the migration has to be safe to run a second time: the retry is the normal path, not an exception. If the failed attempt was interrupted rather than completed — the process killed mid-flight — Helm may still hold the release in a pending state and reject the next operation until that is cleared, which is a different problem from a hook that simply reported failure. ### What to check first The Pod belonging to the failed hook Job is the primary evidence, and the release history tells you whether the attempt got as far as applying manifests at all. Those two facts — did the workload change, and what did the migration Pod say — are the whole triage.

  • Does --rollback-on-failure help here?
    Not with the database. In Helm 4 it restores the previous revision's manifests when a release fails (`--atomic` is a deprecated alias for it). After a pre-upgrade hook failure there is barely anything to restore, since no chart manifest was applied — and in no case does it undo rows the migration committed.
  • How is this different when the hook is post-upgrade instead?
    A post-upgrade hook runs after the chart's manifests are applied, so the new version is already rolling out when the hook fails and the release is marked failed. A smoke test there detects a bad release; it does not prevent one. Only a pre- hook gates the rollout.
  • The upgrade failed on timeout but the migration Pod is still Running. What now?
    Nothing killed it — Helm only stopped waiting. Let it finish or delete it deliberately, but do not fire a retry blindly: a second hook Job would run a second migration concurrently with the first. Raise `--timeout` above the migration's real runtime once you know it.

saying these in an interview costs you the question

  • Says Helm rolls the database back when the hook fails
  • Thinks the new Deployment is applied anyway and just marked failed
  • Believes the timeout kills the running migration Pod
  • Treats a post-upgrade smoke test as a rollout gate
  • Expects a dedicated command to recover the failed release

context