skip to content

Why would `helm upgrade` keep failing in its pre-upgrade hook phase because the migration Job already exists, and how do you fix it?

level: seniorimportance: should knowfreq 47%

answer

  1. Every retry fails the same way
  2. Something is sitting on the hook's name
  3. A Job's pod template cannot be edited
  4. Read the rendered hook's annotations
  5. One policy value replaced the default

basics

~20 s

A hook Job from an earlier run is still holding the name, and the hook's helm.sh/hook-delete-policy does not include before-hook-creation, so nothing clears it first. Delete the leftover to unblock, then add before-hook-creation to the chart's policy list.

solid answer

~40 s

The migration hook has a deterministic name so it can be re-run, and an earlier attempt left an object on that name. Because Helm cannot edit a Job's pod template in place, the create fails, the hook phase fails, and every retry fails identically. The cause is almost always the annotation: a policy of `hook-succeeded` alone replaces the default `before-hook-creation`, so a run that failed — or was interrupted before Helm reached its deletion step — leaves the Job behind forever. Read the Job's Pod logs, delete it, re-run the upgrade to unblock, then fix the chart with `helm.sh/hook-delete-policy: before-hook-creation,hook-succeeded` so every run starts from a clean name. Renaming the hook per revision also dodges the collision but leaves objects accumulating with nothing to reap them.

code

bash · 6 lines
bash
helm history order-checkout -n team-checkout-7
kubectl get jobs -n team-checkout-7
helm get hooks order-checkout -n team-checkout-7
kubectl logs -n team-checkout-7 job/order-checkout-migrate
kubectl delete job -n team-checkout-7 order-checkout-migrate
helm upgrade order-checkout ./order-checkout -n team-checkout-7

go deeper

for a junior

Know that a hook Job with a fixed name can collide with one left by an earlier run, and that the annotation on the hook decides whether that leftover is cleared. Deleting it manually is a valid emergency step, not a fix.

for a middle

Explain why the retry never self-heals: the object holds the name, a Job's pod template cannot be edited, and the missing before-hook-creation means nothing clears it. Name the commands you would run to confirm it.

for a senior

Show a diagnosis order — release history, live objects, rendered hook annotations — and that you read the failed Pod's logs before deleting the evidence. Be ready to explain why only some namespaces are affected.

for a principal

Own the systemic angle: a chart installed by many teams must be re-runnable from any state, so hook idempotence and cleanup are release-engineering standards you set, and manual per-namespace intervention is a signal the chart contract is incomplete.

### The shape of the failure A pre-upgrade hook that runs schema migrations is the classic case. The hook's Job has a stable name so it can be found and re-run; the chart's helper truncates that name at 63 characters, so it is byte-identical on every upgrade of a given release. The upgrade reaches the pre-upgrade phase, Helm tries to create the Job, and a Job with that name is already sitting in the namespace from a previous attempt. The hook phase fails, the upgrade aborts, and it fails the same way on every retry — nothing about the situation changes on its own. Helm cannot patch its way out either: a Job's pod template cannot be edited after creation, so "just update the existing one" is not available to it. ### Diagnosing it Work from the release outward. 1. `helm history <release> -n <ns>` shows the failed revision and confirms the release never got past the hook. 2. `kubectl get jobs -n <ns>` shows the leftover, and its completion count tells you which outcome left it there — a `Complete` Job means the hook succeeded once and nothing removed it, a Job with failures means the run that aborted the previous upgrade is still on disk. 3. `helm get hooks <release> -n <ns>` prints the hook manifests as rendered for that release. Read the `helm.sh/hook-delete-policy` annotation on the migration Job. This is where the answer almost always is. Two annotation states produce this failure. Either the policy list is `hook-succeeded` only — so a successful run cleans up, an unsuccessful one does not, and `before-hook-creation` is no longer in force because writing any value replaces the default — or the policy is `hook-succeeded,hook-failed`, which cleans up both outcomes but not a run that never completed: a Job whose Pod was evicted, a namespace whose quota blocked the Pod, or an upgrade the operator interrupted before Helm reached the deletion step. In both cases the pre-creation delete is missing, and that is the only policy that guarantees a clean slate at the start of a run. ### Why only some namespaces A multi-tenant chart installed once per team namespace makes this look like an environment problem. Three of fourteen namespaces fail; the other eleven upgrade fine. The eleven are simply the ones whose last migration run completed cleanly and deleted its Job. The three are where something went wrong once — weeks ago, possibly under a different chart version — and the wreckage has been waiting on the hook's name ever since. Nothing is special about those clusters; they are the ones that exercised the missing policy. ### Fixing it The unblock is manual and immediate: delete the leftover Job in each affected namespace and re-run the upgrade. Do read its Pod logs first — that Job is the record of why the earlier run failed, and once it is gone the reason is gone. The actual fix is in the chart: put `before-hook-creation` back into the policy list. ```yaml annotations: "helm.sh/hook": pre-upgrade "helm.sh/hook-weight": "0" "helm.sh/hook-delete-policy": before-hook-creation,hook-succeeded ``` Now every run starts by clearing whatever holds the name, whatever left it there, and the failure cannot recur. Keeping `hook-succeeded` stops completed Jobs accumulating; leaving `hook-failed` off means the next failure is still readable in the namespace, which is what you wanted in step 2 above. ### The alternative, and why it is usually worse The other way to dodge a name collision is to stop reusing the name — build the hook Job's name from `.Release.Revision` or another per-run value so each upgrade creates a distinct object. It works, and it preserves every run's logs. It also means nothing ever deletes anything: after a year of weekly releases each namespace carries a pile of dead Jobs and Pods, consuming object quota and making `kubectl get jobs` useless, and the leftovers survive uninstall because Helm does not track hook resources. If you take that route you need a deliberate reaping story — Kubernetes' own time-to-live field on a Job is the one Helm's documentation points at — rather than an implicit one. ### What not to do `--no-hooks` will get the upgrade through, and it will get it through by skipping the migration, which is how a schema-dependent release ships against an unmigrated database. Deleting the Job by hand as a runbook step is the same failure repeated on a schedule: it is the correct emergency action and an admission that the chart is unfinished. And be sceptical of a fix that only makes the hook Job's name more unique without saying who deletes it — uniqueness converts a hard failure into a slow one. ### Prevention for a chart many teams install Treat the hook's delete policy as part of the chart's contract, not as an operator's concern: consumers install into namespaces you do not see and will not clean up. Give every run-to-completion hook `before-hook-creation` so the chart is re-runnable from any state, decide explicitly whether failures are kept for evidence or reaped, and name hook objects from the release so one tenant's cleanup can never target another's object.

  • The chart already sets hook-succeeded,hook-failed. How can a leftover still block the upgrade?
    Both of those fire only after Helm has judged a completed hook. A run that never reaches that point — the Pod evicted, the namespace quota refusing it, an operator interrupting the upgrade, the client losing the cluster — leaves the object with no policy having fired. Only `before-hook-creation` covers an arbitrary starting state, which is why it belongs in the list even when completion-time policies are already there.
  • Why not just give the hook Job a unique name per revision?
    It removes the collision and preserves every run's logs, but nothing then deletes anything: each upgrade adds a Job and its Pods to the namespace, they count against object quota, and they survive uninstall because Helm does not track hook resources. If you choose per-run names you owe the chart an explicit reaping story rather than an implicit one.
  • An engineer unblocked the release with --no-hooks. What did that actually do?
    It shipped the new application version without running the migration. The upgrade succeeds and looks healthy, and the workload now runs against a schema it does not expect — a worse failure than the one that was blocking, and one that surfaces as application errors rather than a clear hook failure. It also destroys the evidence path, since the next run may finally clean up the leftover Job.

saying these in an interview costs you the question

  • Reaching for --no-hooks to push the upgrade through
  • Deleting the Job by hand as the permanent fix
  • Believing hook-succeeded also clears failed runs
  • Assuming Helm patches the existing Job in place
  • Blaming the cluster before reading the hook's annotations
  • Adding a per-run name suffix with no cleanup plan

context