skip to content

Hooks & Tests

Ordinary manifests annotated to run at a point in the release lifecycle, plus the in-cluster smoke test that proves the release works. The database-migration Job lives here, and a hook failing mid-upgrade is how a release gets stuck.

part ofHelmoverview, primer and where to startread it →
on this pageshow

explore

questions

19

What does `helm test` do when you run it against an installed release?

level: juniorimportance: must knowfreq 70%

answer

  1. It needs something already running
  2. Works on a release, not a chart
  3. Marked resources, created on demand
  4. Pod exit status decides the verdict
  5. Non-zero exit is the CI gate

basics

~20 s

helm test looks up an installed release, creates the manifests in it that carry the helm.sh/hook: test annotation - usually Pods - waits for them to finish in the release namespace, and reports pass or fail through its exit code.

solid answer

~50 s

`helm test RELEASE` operates on a release that is already installed: Helm loads the release, picks out the hook manifests annotated `helm.sh/hook: test`, creates them in the release's namespace, and waits for each Pod to terminate. A Pod whose containers all exit 0 reaches `Succeeded` and counts as a pass; anything else is a failure, and the command exits non-zero so a pipeline can gate on it. Because the Pod runs *inside* the cluster it can do things no offline check can - resolve the release's Service name, connect to it, use a Secret the chart rendered. `--logs` dumps the test Pods' output so you can see why one failed. It is a live smoke test, not a chart check: it never runs during `helm install` unless you invoke it, and it does not create a new release revision.

code

yaml · 12 lines
yaml
apiVersion: v1
kind: Pod
metadata:
  name: ledger-api-smoke
  annotations:
    "helm.sh/hook": test
spec:
  restartPolicy: Never
  containers:
    - name: smoke
      image: curlimages/curl:8.6.0
      args: ["--fail", "--silent", "http://ledger-api:8080/healthz"]

go deeper

for a junior

Be ready to say the command needs a release that is already installed, that it runs the resources marked as test hooks in the release's namespace, and that the Pod's exit code decides pass or fail.

for a middle

Explain the mechanics: Helm creates the test resources, waits for each Pod to terminate, and treats Succeeded as a pass. Be clear that no revision is created and that the run is separate from install and upgrade.

for a senior

Show that you know what the command is worth operationally - a non-zero exit is the gate, --logs is the triage, and a chart with no tests passes silently. Say how you keep that silent pass from becoming a green pipeline that proves nothing.

for a principal

Own the policy question: which releases must ship a test, what a test is allowed to assert against production data, and whether the smoke check belongs in Helm at all versus a probe or a synthetic check owned by the platform.

## What the command actually is `helm test RELEASE_NAME` is Helm's only command that runs *your* code against a release that is already live. Everything else Helm offers in the checking family works on chart source: it renders templates, validates values, or diffs YAML. `helm test` does none of that. It takes the name of an installed release, finds the resources in it that are marked as test hooks, creates them in the cluster, and tells you whether they succeeded. ## What has to exist first Two things. A release of that name must exist in the namespace you point at (`-n`, or whatever `HELM_NAMESPACE` resolves to) - if there is no release, the command errors out before it touches the cluster. And the chart that release was installed from must contain at least one manifest annotated `helm.sh/hook: test`. If the chart has none, `helm test` succeeds trivially and has proved nothing at all, which is a real trap: a green test on a chart with no tests looks identical in CI to a green test that actually connected to something. The scaffolding `helm create` emits includes a connection test under `templates/tests/` - a small Pod that fetches the release's own Service. That file is a starting point, not a smoke test worth trusting; most teams replace its body with a request that exercises a real endpoint. ## The lifecycle of a run Helm creates the test resources, then waits. The unit of judgement is the Pod: every container in it must terminate, and terminate with exit status 0, for the Pod to reach the `Succeeded` phase. That is why a test Pod must be written as a **command that ends** - a request, an assertion, an exit - rather than as a server. A Pod that keeps running is not passing slowly; it is not passing at all, and it will burn the whole wait budget before Helm gives up. When every test Pod has finished, Helm reports each one as passed or failed and sets its own exit status accordingly. Non-zero on failure is what makes the command usable as a pipeline step. `--logs` prints the test Pods' container output after the run completes, which is normally the fastest way to see what a failing assertion actually said. `--filter name=<pod-name>` narrows a multi-test chart to a single test while you iterate. ## What it does *not* touch `helm test` does not create a new revision. The release's revision number, its stored manifest and its stored values are all unchanged by a test run, which is why running it repeatedly is cheap and why a passing test is not itself recorded as part of the deployment history. It is also not part of `helm install` or `helm upgrade`: the install finishes, the release is already serving, and the test runs afterwards as a separate command. If you want the test to gate a rollout, your pipeline has to run it and act on the exit code itself. It is also not a substitute for the pre-merge checks a chart needs. Rendering the templates, validating a values file against a schema and asserting on the produced YAML all happen before anything reaches a cluster and catch a different class of defect - a typo in a template, an unset required value. `helm test` catches the class those cannot see: the chart rendered fine, the objects were accepted by the API server, and the thing still does not work because a Service selector matches nothing, a Secret key is spelled differently from the one the app reads, or a dependency is unreachable from that namespace. ## Why the distinction matters in an interview Candidates routinely describe `helm test` as "validating the chart". It validates nothing about the chart. It runs a Pod in a cluster against a running release and reports whether that Pod exited cleanly. Everything the test proves comes from what you wrote inside it, and everything it cannot prove comes from the fact that it is one process, run once, from inside one namespace.

  • Does `helm test` create a new release revision?
    No. It creates the test resources in the cluster and reports a verdict, but the release's revision number, stored manifest and stored values are untouched. That is why re-running it is cheap, and why a passing test leaves no trace in `helm history` - if you want a record that the smoke test passed, your pipeline has to keep it.
  • What exactly makes a chart test count as passed?
    The test Pod must reach the `Succeeded` phase: every container in it terminates, and terminates with exit code 0. A non-zero exit makes the Pod `Failed` and `helm test` exits non-zero. A container that never terminates never passes - it simply consumes the wait budget until Helm gives up.
  • A chart has no test-annotated manifests. What does `helm test` report?
    It has nothing to run and reports success. That is indistinguishable in CI from a test that genuinely connected to the app, so a pipeline that gates on `helm test` should also assert that the chart actually ships at least one test - otherwise deleting the test file silently turns the gate green forever.

It is the smoke test an electrician runs after wiring a building: flip the switch and see whether the light comes on. It says nothing about whether the wiring diagram was drawn correctly - only that this one circuit works right now.

saying these in an interview costs you the question

  • Says helm test validates the chart's templates or YAML
  • Thinks helm test runs automatically as part of helm install
  • Believes it needs the chart directory checked out locally
  • Confuses helm test with a lint or schema check
  • Assumes a long-running test container eventually counts as passing
  • Thinks a test run bumps the release revision

context

open as a page

What does the helm.sh/hook annotation do to a manifest in a Helm chart's templates/ directory?

level: juniorimportance: must knowfreq 70%

basics

~20 s

It turns that manifest into a lifecycle hook. Helm keeps it out of the release manifest and applies it on its own at the named point, such as pre-install or post-upgrade, waiting for it to succeed before the operation continues.

open as a page

How does a Helm chart run a database schema migration before the new pods start?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Ship the migration as a Kubernetes Job template inside the chart, annotated with helm.sh/hook set to pre-install,pre-upgrade. Helm applies that Job and waits for it to complete before applying the chart's own manifests, so the schema changes before any new pod starts.

open as a page

What do the three values of Helm's helm.sh/hook-delete-policy do, and which applies when it is omitted?

level: middleimportance: must knowfreq 61%

basics

~20 s

before-hook-creation deletes an object of the same kind and name just before the hook is created; hook-succeeded deletes it once the hook succeeds; hook-failed deletes it when the hook fails. With the annotation absent, Helm applies before-hook-creation.

open as a page

Helm defines nine helm.sh/hook events — which of them fire during helm upgrade, and which cannot?

level: middleimportance: must knowfreq 58%

basics

~20 s

Only pre-upgrade and post-upgrade fire on helm upgrade. The install, delete, rollback and test events belong to their own commands, so a hook that must run on both the first deploy and every later one has to list two events.

open as a page

What happens to a Helm upgrade when its pre-upgrade migration Job fails?

level: middleimportance: must knowfreq 62%

basics

~20 s

Helm aborts the upgrade. None of the chart's own manifests are applied, the old pods keep serving, and the attempt is recorded as a failed release. The schema changes the Job already made stay: Helm never undoes a hook's side effects.

open as a page

Two Helm pre-install hook Jobs carry no helm.sh/hook-weight - in what order do they run?

level: middleimportance: must knowfreq 55%

basics

~20 s

Both default to weight 0, so they tie. Helm still runs them one at a time in a deterministic order that falls back to the resource name, not the template file order. Give them explicit weights instead.

open as a page

After `helm uninstall`, a Job created by a chart's pre-install hook is still in the namespace. Why?

level: juniorimportance: should knowfreq 44%

basics

~20 s

Helm does not track hook resources in the release manifest, so uninstall removes only the chart's ordinary resources. A hook object is removed solely by its own helm.sh/hook-delete-policy; with no matching policy it outlives the release.

open as a page

In a Helm chart, what does the helm.sh/hook-weight annotation control?

level: juniorimportance: should knowfreq 46%

basics

~20 s

helm.sh/hook-weight is a quoted integer that sorts the hooks firing on the same event, lowest first. Negatives are allowed and a missing annotation means 0. Helm runs each hook to completion before creating the next.

open as a page

Why would `helm test` hang for five minutes and then fail with a timeout?

level: middleimportance: should knowfreq 50%

basics

~20 s

Because helm test waits for each test Pod to terminate and its --timeout defaults to 5m0s. A container that never exits, or a Pod that never starts, burns that budget and reports a timeout instead of a failure.

open as a page

What does Helm's --no-hooks flag do, and when is reaching for it dangerous?

level: middleimportance: should knowfreq 38%

basics

~20 s

It makes one Helm operation skip every hook that operation would have fired, including subchart hooks. It is all-or-nothing and skipped hooks are never queued for later, so any invariant a hook maintained silently does not hold.

open as a page

Why would `helm upgrade` keep failing in its pre-upgrade hook phase because the migration Job already exists, and how do you fix it?

level: seniorimportance: should knowfreq 47%

basics

~20 s

A hook Job from an earlier run is still holding the name, and the hook's helm.sh/hook-delete-policy does not include before-hook-creation, so nothing clears it first. Delete the leftover to unblock, then add before-hook-creation to the chart's policy list.

open as a page

Your `helm test` passed after upgrading a payments ledger release - what has it proved?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Only that one Pod, run once from inside the release namespace, exited 0. That covers in-cluster reachability of whatever it touched at that instant - not the correctness of the new version, the rendered configuration, or anything outside the cluster.

open as a page

A ConfigMap annotated helm.sh/hook: post-install is stale in all 12 namespaces after 23 upgrades — why does Helm never update it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Two reasons stack. The post-install event fires only on the original install, never on an upgrade, and hook resources are excluded from the release manifest, so even the diff Helm computes on upgrade cannot see the object. Nothing will ever reconcile it.

open as a page

Your Helm pre-upgrade migration Job failed midway and every retry fails on a change that already exists — how do you make it re-runnable?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Make the migration decide from recorded state rather than assumption: a ledger table of applied versions, one transaction per step where the engine allows it, and a lock so only one runner migrates at a time. A retry then resumes instead of repeating.

open as a page

A Helm pre-upgrade hook Job fails because a ConfigMap the same chart renders does not exist yet - why?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Pre-upgrade hooks are a phase that finishes before Helm applies any of the chart's ordinary manifests, and hook-weight orders hooks only against other hooks. Make the ConfigMap a hook too, or move the work to post-upgrade.

open as a page

A Helm rollback restores the previous revision's manifests but not the database it migrated — how do you keep releases reversible?

level: principalimportance: should knowfreq 44%

basics

~20 s

Only ship migrations the previous release's code can still run against: additive changes now, destructive ones in a later release once the old version is gone. Helm restores manifests in seconds; whether that restoration actually works is decided entirely by the schema.

open as a page

Which copy of a chart test does `helm test` run - your working tree's or the release's?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

The release's. Helm loads the installed release from its storage record, which already holds the rendered hook manifests, and runs those - so edits to a test file on disk have no effect until you upgrade the release.

open as a page

Why would you give a Helm chart hook a negative helm.sh/hook-weight?

level: middleimportance: nice to knowfreq 24%

basics

~10 s

Because 0 is the default, every hook that omits the annotation sits at 0 - including hooks from subcharts you do not control. Only a negative weight is guaranteed to run ahead of them.

open as a page