skip to content

A Helm upgrade fails with "pre-upgrade hooks failed" — what changed in the cluster, and how do you find why?

level: seniorimportance: should knowfreq 45%

answer

  1. Hooks run before anything else
  2. So what did not get applied?
  3. The hook is still a real object
  4. Unless an annotation deleted it
  5. helm get hooks shows what rendered

basics

~20 s

Nothing from the chart's ordinary manifests was applied — hooks run first — so the previous revision is still serving, though a failed revision is recorded. Find the cause in the hook's own Job and Pod, and use helm get hooks to see which hook ran.

solid answer

~50 s

Helm applies `helm.sh/hook: pre-upgrade` resources in `helm.sh/hook-weight` order and waits for each before touching the chart's ordinary manifests, so a hook failure means none of them were applied: the old objects are still running, and Helm records the new revision as `failed`. What the hook itself did to the outside world, though — a migration, a data backfill — is half-done and Helm cannot undo it. To find why, run `helm get hooks <release>` to see the rendered hook manifests, their events, weights and delete policies, then read the corresponding Job and its Pod in the release namespace. The trap is `helm.sh/hook-delete-policy: hook-failed`: Helm deletes the failed hook resource immediately and its logs go with it, so you must either strip that policy for the retry or have the hook ship its output somewhere durable.

code

yaml · 18 lines
yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: digest-builder-migrate
  annotations:
    "helm.sh/hook": pre-upgrade
    "helm.sh/hook-weight": "-5"
    # hook-failed would delete this Job, and its logs, on failure
    "helm.sh/hook-delete-policy": before-hook-creation,hook-succeeded
spec:
  backoffLimit: 0
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: migrate
          image: registry.internal/digest-builder:2.14.3
          args: ["migrate", "--up"]

go deeper

for a junior

Know that Helm runs pre-upgrade hooks before it applies the chart's other resources, and that the hook is a normal Kubernetes object you can go and look at in the release's namespace.

for a middle

Be ready to explain the ordering and the weights, and to name the command that prints a release's rendered hooks rather than guessing which hook ran from the chart source.

for a senior

Show the state deduction first — what is and is not running — then the evidence path, including the delete-policy trap and the fact that a hook's real-world side effects survive the failure.

for a principal

Own the chart-authoring consequence: hooks a fleet retries must be idempotent and must not delete their own failure evidence, and irreversible work in a pre-upgrade hook is a design decision, not a detail.

`Error: UPGRADE FAILED: pre-upgrade hooks failed: ...` is the most informative failure Helm produces, because the ordering guarantee behind it tells you the cluster state without your having to look. ## What Helm had done by then On `helm upgrade`, Helm renders the chart, records a new revision, then applies the manifests annotated `helm.sh/hook: pre-upgrade`. Hook resources are applied in ascending `helm.sh/hook-weight` order, and Helm waits for each one to reach a successful terminal state — for a Job, completion — before moving to the next weight. Only after every pre-upgrade hook has succeeded does Helm apply the chart's ordinary resources. So when the message says the pre-upgrade hooks failed, the deduction is firm: **none of the chart's ordinary manifests were applied**. The Deployments, Services and ConfigMaps of the previously deployed revision are untouched and still serving traffic. The release, however, is not clean: a new revision exists in the history with status `failed`, and `helm status <release>` will report `failed` with the same text in its description. ## What Helm cannot tell you The hook is a real Kubernetes object doing real work. If it ran a schema migration against a database, that migration is now half-applied, and stopping the upgrade did not roll it back. Helm's model of a failure is "the release is failed"; it has no model of your data. This is why an interviewer probes here: a candidate who says "the cluster is untouched" has understood the manifests and missed the hook's side effects. ## Finding the cause Start from the chart's declaration rather than from the cluster, because you first need to know *which* hook of possibly several failed and what it was supposed to do. `helm get hooks <release>` prints the rendered hook manifests exactly as Helm applied them — names, `helm.sh/hook` events, `helm.sh/hook-weight` values and `helm.sh/hook-delete-policy` annotations. This is the only view that shows what the templates actually produced, which matters when the hook does not live in your chart at all. Consider an email-digest builder whose service chart consumes a shared library chart, one of twelve services doing so: the migration Job is defined by a named template in the library chart, so the file you would open in your own repository does not contain the manifest that ran. `helm get hooks` does. With the hook named, the failure itself is in the hook's Job and the Pod it created, in the release namespace. Reading a failed Pod — events, exit codes, container logs — is ordinary Kubernetes work; the Helm-side skill is knowing that this is where to go and that the object is still there to be read. ## The evidence trap Whether it is still there depends on an annotation. | `helm.sh/hook-delete-policy` | When the failed Job is deleted | What is left to read | | --- | --- | --- | | Not set — Helm applies `before-hook-creation` | The previous run's hook resource is deleted just before the next run creates its replacement | The failed Job survives until you retry | | `hook-succeeded,hook-failed` | The moment it fails | A failure message and no detail: the Pod, and therefore its logs, go with it | The responses are all deliberate: - retry from a chart copy with the failed-deletion policy removed so the evidence survives; - have the hook write its output to a durable place, such as log shipping or a status object, rather than only to stdout; - or reproduce the hook's work as a standalone Job you control. A chart you ship to other teams should not delete its failure evidence by default, and that is a chart-authoring consequence of an operations problem. ## Retrying Fix and re-run is usually the move, but two things bite. First, a pre-upgrade hook runs again on every attempt, so a hook that is not idempotent — an `ALTER TABLE` without a guard, an append-only backfill — will misbehave on the second run; hooks in charts that get retried must be written to be re-runnable. Second, if you have already performed the hook's work by hand and just want the manifests applied, `--no-hooks` skips hooks on that upgrade — but it skips *all* of them, including any post-upgrade hook the chart relies on, so it is a deliberate one-off, not a default. The answer an interviewer is listening for has three parts: the ordering deduction about what did and did not change, the two-step path from `helm get hooks` to the hook's own Pod, and the awareness that the hook's side effects and its delete policy are the parts Helm will not clean up for you.

  • The failed hook Job is gone before you can read it. What happened, and what do you change?
    The chart annotated the hook with `helm.sh/hook-delete-policy` including `hook-failed`, so Helm removed the Job and its Pod as soon as it failed. Retry from a chart copy without that policy, or better, change the chart: default to `before-hook-creation` so the previous run's resource is cleared only when the next one starts, and have the hook emit its diagnostics somewhere that outlives the Pod.
  • Is it safe to just run the upgrade again after a pre-upgrade hook failed?
    Only if the hook is idempotent. Helm re-runs every pre-upgrade hook on each attempt, so a migration that is not guarded, or a backfill that appends, will run a second time. Charts whose hooks perform irreversible work should make them re-runnable; where they are not, complete the work manually and use `--no-hooks` once, accepting that it skips the chart's other hooks too.

saying these in an interview costs you the question

  • Says the failed upgrade left everything untouched
  • Assumes Helm rolls back the hook's side effects
  • Looks for the hook manifest only in the local chart
  • Reruns a non-idempotent migration hook without checking
  • Thinks --no-hooks skips only the failed hook
  • Never notices the hook Job was deleted on failure

context