skip to content

A ConfigMap annotated helm.sh/hook: post-install is stale in all 12 namespaces after 23 upgrades — why does Helm never update it?

level: seniorimportance: should knowfreq 42%

answer

  1. The upgrade reported success anyway
  2. Which command fires that event
  3. What an upgrade actually compares
  4. Two reads: manifest versus hooks
  5. Hooks are actions, not managed state

basics

~20 s

Two reasons stack. The post-install event fires only on the original install, never on an upgrade, and hook resources are excluded from the release manifest, so even the diff Helm computes on upgrade cannot see the object. Nothing will ever reconcile it.

solid answer

~40 s

`post-install` fires once, on the install that created the release; 23 upgrades fire `post-upgrade` and nothing else, so the hook is never re-applied. Underneath that, the deeper problem is that hook manifests are removed from the release manifest — the document Helm diffs old against new on every upgrade — so the object is not tracked state at all. `helm get manifest` will not list it, `helm rollback` will not restore it, and deleting the hook from the chart deletes nothing in the cluster. `helm get hooks` shows the *current* rendered text, which is why the change looks correct on paper while all 12 namespaces still serve the original content. The fix is not another event: a resource that must track the chart is an ordinary template, not a hook.

code

bash · 5 lines
bash
for ns in $(kubectl get ns -o name | cut -d/ -f2 | grep '^ingest-'); do
  echo "== $ns"
  helm get manifest telemetry-gateway -n "$ns" | grep -c 'gateway-sampling-rules' || true
  helm get hooks    telemetry-gateway -n "$ns" | grep -c 'gateway-sampling-rules' || true
done

go deeper

for a junior

Take away the smaller half of the answer: post-install fires on the install and never again, so a manifest annotated with it is applied once in the life of a release no matter how many upgrades follow.

for a middle

Explain what an upgrade actually compares — the stored release manifest against the newly rendered one — and why a manifest removed from that document can never be diffed, patched or pruned however the chart changes.

for a senior

Diagnose it live: two reads (helm get manifest versus helm get hooks) plus the live object, then argue for removing the annotation rather than adding events, and plan the ownership-metadata problem the conversion creates across every namespace.

for a principal

Set the boundary for your chart authors: hooks own side effects with a success signal, the release manifest owns everything that must converge, prune and roll back. Say how you would find the charts already violating it before the next silent no-op upgrade.

### The incident A telemetry ingest gateway chart ships a ConfigMap of sampling rules and, when it was written, someone annotated it `helm.sh/hook: post-install` — probably to control when it appeared relative to the gateway Pods. The chart has since moved through 23 upgrades and is deployed into a 12-namespace fleet. An engineer edits the sampling rules in `templates/`, ships chart 4.11.2, watches `helm upgrade` report success in all 12 namespaces, and finds every gateway still applying the rules from the original install months ago. ### Reason one: the event never fires again `post-install` is fired by `helm install` and by nothing else. An upgrade fires `pre-upgrade` and `post-upgrade`. So across 23 upgrades this manifest was rendered 23 times and applied zero times. That alone explains the staleness, and it is what most candidates spot. ### Reason two: it is not release state The more important reason is structural. Helm renders the chart, then removes every `helm.sh/hook`-annotated manifest from the **release manifest** — the single document stored on the release and the only thing an upgrade compares. Hook manifests are stored separately on the release. Consequently, for anything a hook creates: - `helm get manifest <release>` does not list it (`helm get hooks <release>` does); - an upgrade computes no diff for it, so it is never patched, whatever the chart now says; - removing it from the chart removes nothing from the cluster — there is no prune for an object that was never in the manifest; - `helm rollback` replays a previous revision's stored manifest, which does not contain it, so a rollback cannot restore it either. Adding `post-upgrade` to the annotation papers over the symptom — the object would be re-applied each upgrade — while leaving the object outside the release forever: still absent from `helm get manifest`, still unpruned when the chart drops it, still unrestorable on rollback. ### The diagnostic that makes it obvious Run both reads against one namespace: ```bash helm get manifest telemetry-gateway -n ingest-eu | grep -c sampling-rules # 0 helm get hooks telemetry-gateway -n ingest-eu | grep -c sampling-rules # 1 kubectl get configmap gateway-sampling-rules -n ingest-eu -o yaml # old content ``` The hook text stored with the current revision is the *new* text, which is exactly why the change looks shipped: the release record was updated, the cluster was not. ### The fix Drop the annotation and let the ConfigMap be an ordinary template. From the next upgrade it becomes release state: diffed, patched, pruned when removed, restored on rollback, and — if the Pods consume it as a checksum-annotated volume — able to trigger a restart on change. One wrinkle: the live object was created by a hook and does not carry the ownership metadata Helm writes on resources it manages (`app.kubernetes.io/managed-by: Helm` plus `meta.helm.sh/release-name` and `meta.helm.sh/release-namespace`). Helm refuses to adopt a conflicting existing object rather than silently taking it over, so per namespace you either delete the object first, add the metadata by hand, or take ownership explicitly on the upgrade. In a 12-namespace fleet that is a short scripted loop, not a redesign. ### The design rule worth stating Hooks are for **actions**: migrate a schema, warm a cache, take a lease, register with a discovery service, deregister on the way out. Their value is the side effect, and Helm's contract is only that they run at a named point and must succeed. Anything whose value is the object continuing to exist and match the chart is **state**, and state belongs in the release manifest. A chart that manages state through hooks looks fine for exactly as long as nobody upgrades, removes or rolls anything back — and then it fails in the least visible way available, by reporting success.

  • You convert the hook ConfigMap into an ordinary template. What do you have to do about the object already in the cluster?
    It was created by a hook, so it lacks the ownership metadata Helm writes on managed resources — the `app.kubernetes.io/managed-by: Helm` label and the `meta.helm.sh/release-name` and `meta.helm.sh/release-namespace` annotations. Helm will not silently adopt it; per namespace you delete the object first, add that metadata by hand, or tell the upgrade to take ownership. Plan the loop across all 12 namespaces before shipping the chart change.
  • Would annotating it post-install,post-upgrade have been an acceptable fix?
    It fixes only the visible symptom. The object would be re-applied on every upgrade, but it stays outside the release manifest: still missing from `helm get manifest`, still not pruned when the chart drops it, still not restored by `helm rollback`, and still invisible to any diff-based review of what an upgrade will change. If the object must track the chart, remove the annotation instead of adding events to it.
  • Does helm rollback undo anything a hook did?
    No. A rollback re-applies the target revision's stored release manifest, which by definition contains no hook resources, and it fires `pre-rollback`/`post-rollback` rather than replaying the hooks that ran on the way up. Any side effect a hook had — an object created, a schema migrated, a row written — survives the rollback untouched unless a rollback hook is written specifically to reverse it.

saying these in an interview costs you the question

  • Expects hook resources to appear in helm get manifest
  • Thinks an upgrade re-applies post-install hooks
  • Believes helm rollback reverts hook-created objects
  • Assumes deleting the hook from the chart deletes the object
  • Reaches for a force-replace upgrade to reconcile it
  • Blames cluster drift rather than the chart's design

context