A Helm pre-upgrade hook Job fails because a ConfigMap the same chart renders does not exist yet - why?
answer
- Two phases, not one apply
- The weight never leaves its own phase
- The old revision's copy hides the bug
- Make the input a hook too, or move events
basics
~20 sPre-upgrade hooks are a phase that finishes before Helm applies any of the chart's ordinary manifests, and hook-weight orders hooks only against other hooks. Make the ConfigMap a hook too, or move the work to post-upgrade.
solid answer
~50 sThe hook phase and the manifest phase are separate steps, and weights only sort *within* one of them. On `helm upgrade`, Helm runs every `pre-upgrade` hook to completion and only then applies the chart's ordinary resources. A ConfigMap that is just a normal template is therefore not in the cluster yet, and no weight you give the hook Job changes that: the annotation cannot interleave a hook into the ordinary manifest set. On a first install the ConfigMap is absent entirely; on an upgrade the previous revision's copy exists, which is why the bug hides until the hook needs the *new* content. The three fixes: annotate a copy of the ConfigMap as a hook with a lower weight so it is created in the same phase, render the data straight into the hook Job's spec, or move the work to a `post-upgrade` hook.
code
yaml · 29 linesapiVersion: v1
kind: ConfigMap
metadata:
name: fanout-reshard-input
annotations:
"helm.sh/hook": pre-upgrade
"helm.sh/hook-weight": "-10"
data:
shards: "48"
---
apiVersion: batch/v1
kind: Job
metadata:
name: fanout-reshard
annotations:
"helm.sh/hook": pre-upgrade
"helm.sh/hook-weight": "0"
spec:
activeDeadlineSeconds: 96
template:
spec:
restartPolicy: Never
containers:
- name: reshard
image: registry.example.internal/fanout-tools:1.4.2
args: ["reshard"]
envFrom:
- configMapRef:
name: fanout-reshard-inputgo deeper
Take away the shape: hooks on a pre-* event all run before any of the chart's normal resources are applied, and the weight annotation only sorts hooks against other hooks. That alone explains the missing ConfigMap.
Explain the phase order of an upgrade and why the weight cannot cross it, then give the concrete fix - annotate the ConfigMap as a hook with a lower weight, or render the value straight into the Job.
Diagnose why it passed for months: the previous revision's ConfigMap masked the ordering bug until the content changed. Weigh the three fixes, and know that a post-* hook is not a guarantee the new workload is ready.
Decide where lifecycle work belongs at all. Sequencing encoded in chart hooks is invisible to reviewers and hard to test; argue for pushing it into the application's startup path or a dedicated job, and set the rule for when a chart may own it.
### Two phases, and the weight only sorts inside one A Helm operation is a sequence of phases, not one big apply. On an upgrade, Helm renders the chart, records a new pending revision, runs the `pre-upgrade` hooks to completion, applies the chart's ordinary manifests, and then runs the `post-upgrade` hooks. `helm.sh/hook-weight` orders the hooks *inside* a single one of those hook phases. It is compared only against the weights of the other hooks firing on the same event. It is not a global position in the release, and there is no weight - however large or however negative - that places a hook between two of the chart's ordinary resources. That is the whole shape of the problem: the ordering the author wants crosses a boundary the annotation cannot cross. ### Why it looks intermittent A chat fan-out service whose chart is packaged by CI and pushed to an OCI registry might carry a `pre-upgrade` Job that rebuilds routing shards and reads the shard count from a ConfigMap the chart also renders. The chart installs fine on a fresh namespace only if the Job tolerates the ConfigMap's absence; more often the first install fails outright. On an upgrade it usually *succeeds*, because the ConfigMap from the previous revision is still in the cluster - the hook reads yesterday's shard count and nobody notices. The failure surfaces the day the upgrade is the one that changes the shard count: the hook runs before the new ConfigMap is applied, reads the old value, and either errors or, far worse, quietly does the wrong work. If the hook blocks waiting for content that will only exist after it finishes, it deadlocks and dies at the command's hook timeout - a 96-second timeout on that Job just means you find out 96 seconds sooner. ### Fix one: bring the data into the hook phase Anything a hook needs can itself be a hook. Annotate a copy of the ConfigMap as a `pre-upgrade` hook with a lower weight than the Job: ```yaml apiVersion: v1 kind: ConfigMap metadata: name: fanout-reshard-input annotations: "helm.sh/hook": pre-upgrade "helm.sh/hook-weight": "-10" data: shards: "48" ``` Now both objects are in the same pool, Helm sorts `-10` before the Job's `0`, and the ConfigMap exists when the Job starts. Two consequences to state before an interviewer states them for you: this object is now a hook, so it lives outside the release's ordinary manifest set and its lifecycle is governed by the hook delete-policy annotation rather than by ordinary upgrade behaviour; and you now have two ConfigMaps rendering the same content, so generate both from one template helper rather than letting them drift. ### Fix two: render the value into the hook itself Often the hook does not need a ConfigMap at all - it needs a value. Passing it through `args` or `env` rendered from `.Values` removes the ordering problem instead of solving it, and it is the cheapest option when the payload is small. ### Fix three: move to a post-* event If the work genuinely belongs after the chart's resources exist, a `post-upgrade` hook is the correct event, and then the ConfigMap has been applied before the hook runs. The trap is readiness rather than existence: Helm always waits for each *hook* to finish, but whether it waits for the chart's own workloads before running the post phase depends on the wait strategy you asked for. In Helm 4, `--wait` is strategy-valued (`watcher`, `legacy`, `hookOnly`) and omitting it means `hookOnly` - hooks are waited for, workloads are not - so a `post-upgrade` hook can start while the new Pods are still rolling. Ask for a waiting strategy if the hook needs the new version actually serving. Helm 3 behaves comparably by default: no `--wait` means no wait on workloads, and its `--wait` is a plain boolean. ### What not to do Do not put a large negative weight on the ConfigMap in the hope of dragging an ordinary manifest earlier - weights on non-hook manifests are ignored. Do not add a sleep to the hook: the ConfigMap is not late, it is in a phase that has not started. And do not reach for a separate pre-apply step in the pipeline that applies the ConfigMap out of band, because you have then split one release into two sources of truth, and a rollback restores only half of it.
- Why does this chart upgrade cleanly for months and then fail on one particular release?Because the ConfigMap from the previous revision is still in the cluster, the hook reads stale content and appears to work. The failure surfaces on the upgrade that actually changes that content: the hook runs before the new version is applied, so it acts on the old value. A fresh install into an empty namespace exposes the bug immediately.
- If you move the work to a post-upgrade hook, is the new version guaranteed to be serving traffic when it runs?No. The chart's manifests have been applied, but whether Helm waits for the workloads to become ready before the post phase depends on the wait strategy. Helm 4 defaults to waiting for hooks only, so a post-upgrade hook can start against Pods that are still rolling out; ask for a waiting strategy if the hook needs the new version live.
- A colleague suggests giving the plain ConfigMap a very negative helm.sh/hook-weight. What happens?Nothing useful. A weight on a manifest that is not marked as a hook is ignored - it stays an ordinary resource applied in the manifest phase, after the pre-upgrade hooks have already finished. The annotation only has meaning on a hook, which is exactly why the ConfigMap has to become one for that approach to work.
saying these in an interview costs you the question
- Thinks a big enough weight orders a hook after normal resources
- Puts a hook weight on a non-hook manifest
- Adds a sleep to the hook to wait for the ConfigMap
- Blames the API server or a slow cluster
- Assumes post-* hooks run only once Pods are ready
- Applies the ConfigMap out of band from the pipeline