skip to content

Operating & Triage

Running releases somebody else installed: reading what is on the cluster now, telling a fresh render from the stored manifest, and turning a failure message into the next command.

part ofHelmoverview, primer and where to startread it →
on this pageshow

explore

questions

24

How does the Helm CLI decide which cluster and namespace a command targets?

level: juniorimportance: must knowfreq 74%

answer

  1. Helm remembers nothing between commands
  2. The same file another CLI reads
  3. One flag beats one environment variable
  4. Then the context's own namespace field
  5. Last resort is literally default

basics

~20 s

Helm keeps no target state of its own. It acts on the kubeconfig's current context unless --kube-context overrides it, and targets the namespace from -n, else HELM_NAMESPACE, else the namespace recorded on that context, else default.

solid answer

~50 s

Every `helm` invocation resolves its target from scratch — there is no daemon and no remembered selection. For the cluster, Helm loads a kubeconfig the same way the standard Kubernetes client libraries do: `--kubeconfig` or the `KUBECONFIG` environment variable picks the file, the file's current context picks cluster and credentials, and `--kube-context` overrides that context for one command. Helm never writes back to the file, so there is no Helm command that switches contexts. For the namespace the order is `-n`/`--namespace`, then `HELM_NAMESPACE`, then the namespace field on the selected context, then `default`. This is not cosmetic: a release record lives as a Secret in one namespace, so from the wrong namespace `helm status` reports the release missing and `helm upgrade --install` happily creates a second, independent release of the same name. Helm creates a missing namespace only when `--create-namespace` is passed.

code

bash · 8 lines
bash
# Nothing inherited from the shell: cluster and namespace are both explicit
helm upgrade --install billing-cron billing-charts/billing-cron \
  --kube-context prod-eu \
  --namespace billing --create-namespace \
  -f prod-values.yaml

# Confirm what Helm resolved before acting
helm env | grep HELM_NAMESPACE

go deeper

for a junior

Be ready to state the namespace order out loud: -n, then HELM_NAMESPACE, then the context's namespace, then default. Know that Helm reads your kubeconfig and never edits it.

for a middle

Explain why the namespace is part of a release's identity: the release record is a Secret in that namespace, which is why the same command in two namespaces produces two independent releases.

for a senior

Show the diagnosis habit — confirm context and namespace before touching a release you did not install, and be able to describe the duplicate-release and ownership-metadata failures this causes in production.

for a principal

Own the guardrail: make target selection explicit in every pipeline, decide whether --create-namespace is allowed at all, and avoid shared kubeconfigs whose current context any tool can move under you.

### Helm is a client, and it remembers nothing There is no Helm daemon and no in-cluster Helm component. Every `helm` invocation works out from scratch which API server to talk to and which namespace to treat as the release's home, using only the flags on that command line, the environment, and your kubeconfig. A large share of "it worked on my machine" reports about Helm end at one of those two resolutions. ### Choosing the cluster Helm builds its Kubernetes client the way other clients built on the standard Kubernetes client libraries do. The kubeconfig file is chosen by `--kubeconfig` if you pass it, otherwise by the `KUBECONFIG` environment variable, otherwise by the default path in your home directory. Inside that file one context is marked current, and a context names a cluster, a user and optionally a namespace. `--kube-context <name>` selects a different context for that single command; `HELM_KUBECONTEXT` is the environment binding of the same flag. The important asymmetry with `kubectl` is that **Helm only reads**. It has no command that changes your current context, because that state belongs to the kubeconfig file, not to Helm. If a script needs a particular cluster, the script passes `--kube-context` (or `--kubeconfig`) on every command rather than assuming whatever context the file happens to point at. ### Choosing the namespace Four sources, in this order: 1. `-n` / `--namespace` on the command line. 2. the `HELM_NAMESPACE` environment variable. 3. the `namespace` field of the selected kubeconfig context. 4. the literal string `default`. The flag beats the environment variable, and both beat the context. The last fallback is a genuine trap: an operator with no namespace on their context and no `-n` habit spends their day acting on `default` while believing they are "in" the namespace they were looking at a minute ago in another tool. ### Why the namespace is part of the release's identity A Helm release is not cluster-scoped. Its record is stored as a Secret named `sh.helm.release.v1.<name>.v<rev>` **in the release's namespace**, so a release is identified by namespace *plus* name. Four consequences follow directly: - `helm status` and `helm get` from the wrong namespace report that the release does not exist, even while the workload is plainly running. - `helm upgrade --install` from the wrong namespace does not fail. Finding no record there, it *installs* — you now own two independent releases with the same name in two namespaces, each with its own revision history. - If those two renders collide on cluster-scoped or already-existing objects, Helm refuses with an ownership-metadata error complaining that the object's `meta.helm.sh/release-namespace` annotation does not match the namespace it is being asked to manage from. - `.Release.Namespace` inside templates is exactly this resolved value, so anything a chart renders from it — a ConfigMap reference, a service DNS name, an annotation — changes silently with the flag. ### A worked example A subscription-billing cron ships as a chart with a `values.schema.json` contract and an 11-value production override file. It belongs in the `billing` namespace. An engineer whose current context defaults to `platform-ops` runs `helm upgrade --install billing-cron billing-charts/billing-cron -f prod-values.yaml` with no `-n`. Helm finds no `billing-cron` release in `platform-ops`, installs a fresh revision 1 there, and the cron begins running twice — once from each namespace. `helm list` in `billing` still shows the old revision, unchanged, which is exactly why the engineer's first instinct ("the upgrade did not apply") is wrong. ### Getting it right - In automation, pass `--kube-context` and `-n` explicitly on every command. Inheriting either from the ambient environment is the single most common cause of a change landing in the wrong place. - Interactively, `helm env` prints the resolved `HELM_NAMESPACE`, so you can confirm what Helm thinks before acting. - Remember that Helm does not create namespaces implicitly: `install` (or `upgrade --install`) into a namespace that does not exist fails unless you add `--create-namespace`. - Export `HELM_NAMESPACE` when you want a whole shell session pinned, but keep in mind it is invisible to the next person reading your terminal history, whereas `-n` is not.

  • Does `helm install -n billing` create the billing namespace if it does not exist?
    No. Helm fails rather than creating it. You have to pass `--create-namespace`, which is available on `install` and on `upgrade --install`. Without it the API rejects the objects because their namespace is absent, and you get a failed release rather than a partially created one. Many teams deliberately omit the flag in production so that a typo'd namespace is an error instead of a new, empty namespace nobody owns.
  • Your kubeconfig holds four clusters. How do you keep a deploy script from ever hitting the wrong one?
    Pass `--kube-context` on every Helm command, or give the script its own `--kubeconfig`/`KUBECONFIG` pointing at a file containing only the intended cluster. Never rely on the current context: it is shared mutable state that another tool or another human can change between two lines of the same script. Helm cannot switch it for you, and it has no confirmation prompt.

Helm is a courier who reads the address off the envelope you hand over on each trip. It keeps no address book of its own, so if you leave the address blank it delivers to the default door.

saying these in an interview costs you the question

  • Says Helm remembers the namespace you last used
  • Thinks Helm has its own context-switching command
  • Assumes Helm creates a missing namespace automatically
  • Believes HELM_NAMESPACE overrides an explicit -n flag
  • Treats a release as cluster-wide, so namespace is cosmetic
  • Expects upgrade --install to fail in the wrong namespace

context

open as a page

When you replace the Helm 3 CLI with Helm 4, what happens to releases already installed in the cluster?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Nothing has to be done to them. Helm 4 adopts Helm 3 releases in place: the release record is still a Secret in the same format and v2 charts still install, so there is no migration tool and nothing to convert.

open as a page

What does `helm list` show by default, and how do its namespace and status flags change that?

level: juniorimportance: must knowfreq 76%

basics

~20 s

helm list shows releases in one namespace - the current kube context's, or the one -n names; -A lists every namespace. Helm 4 lists all statuses by default, so flags like --failed narrow the list rather than widening it.

open as a page

Why does helm upgrade still fail on a removed apiVersion after the chart was fixed to emit the supported one?

level: middleimportance: must knowfreq 70%

basics

~20 s

Because Helm must decode the previous release's stored manifest on every upgrade, and that text is frozen at install time. Fixing the chart changes what you render, not the retired apiVersion already recorded with the live release.

open as a page

Why does `helm status` still report a release as deployed after someone hand-edits its live objects?

level: middleimportance: must knowfreq 62%

basics

~20 s

Helm stores the manifest it applied in the release record and never watches the cluster afterwards. helm status reads that record, so it reports the outcome of the last operation, not whether the live objects still match it.

open as a page

A helm upgrade ends with "Error: UPGRADE FAILED" — how do you classify what actually failed?

level: middleimportance: must knowfreq 66%

basics

~20 s

Read the clause after UPGRADE FAILED. Helm names the class there: a failed hook, an API-server rejection such as an admission webhook denial or a field-manager conflict, or an expired wait reported as a context deadline. Each leaves the cluster differently.

open as a page

A helm upgrade fails with 'no matches for kind "Ingress" in version "networking.k8s.io/v1beta1"' — what is the cluster telling you?

level: juniorimportance: should knowfreq 58%

basics

~10 s

The cluster no longer serves that apiVersion, so Helm cannot map the kind to any resource the API server offers. Some manifest Helm is reading still names networking.k8s.io/v1beta1, which the upgraded cluster dropped.

open as a page

What does `helm env` print, and how do you use it when Helm behaves differently on two machines?

level: middleimportance: should knowfreq 52%

basics

~20 s

helm env prints the client's fully resolved settings as shell-style assignments: the cache, config and data home directories, the repository cache and config paths, the plugin directory, the registry credentials file, and the resolved namespace. It contacts no cluster.

open as a page

What does `helm diff upgrade` compare, and how is that different from `helm upgrade --dry-run=server`?

level: middleimportance: should knowfreq 48%

basics

~20 s

The helm-diff plugin's upgrade subcommand renders the new chart and prints a unified diff against the manifest stored in the release record. A server-side dry-run renders and sends the objects to the API server for validation and admission, printing the manifest rather than a diff.

open as a page

Why does Helm 4's --server-side=auto leave a release installed by Helm 3 applying client-side?

level: middleimportance: should knowfreq 44%

basics

~20 s

On helm upgrade and helm rollback, Helm 4's --server-side is a string defaulting to auto, which inherits the apply method recorded for the previous revision. A revision written by Helm 3 was applied client-side, so the release keeps the old three-way merge until you ask for server-side explicitly.

open as a page

What is the difference between `helm get values` and `helm get values --all` for a release?

level: middleimportance: should knowfreq 58%

basics

~20 s

helm get values prints only the overrides supplied at the last install or upgrade, as stored with that revision. --all prints the computed tree: those overrides merged over the chart's and subcharts' defaults, which is what the templates resolved against.

open as a page

How do you repair a Helm release whose stored manifest names a removed apiVersion, without deleting the running workload?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Either rewrite the retired apiVersion inside the stored release records, or delete the release history and reinstall under the same name and namespace so Helm adopts the live objects. Pair either repair with a chart bump.

open as a page

A CI job's `helm upgrade` resolves a chart version older than the one published an hour ago — why, and how do you make it deterministic?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Helm resolves a repo/chart reference against the repository index cached on that machine, not against the server. Nothing refreshes that cache on a timer, so a runner keeps whatever it downloaded until helm repo update runs.

open as a page

Piping `helm get manifest` into `kubectl diff` shows a large diff on a release nobody touched — why?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Most of that diff is expected: fields other controllers own, server-assigned bookkeeping metadata, and objects the stored manifest never covered. Real drift is what a human or an unexpected actor wrote, so triage the diff by asking who owns each changed field.

open as a page

You inherit a Helm release nobody documented - which commands establish what is deployed, and in what order?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Find it with helm list -A, read helm status for the current state and the last operation's description, then helm get metadata, values, values --all, manifest and hooks for the chart version, overrides and applied YAML - before changing anything.

open as a page

A Helm upgrade fails with "pre-upgrade hooks failed" — what changed in the cluster, and how do you find why?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Nothing from the chart's ordinary manifests was applied — hooks run first — so the previous revision is still serving, though a failed revision is recorded. Find the cause in the hook's own Job and Pod, and use helm get hooks to see which hook ran.

open as a page

A helm upgrade fails with "context deadline exceeded" after 7 minutes — is the rollout stuck or just slow?

level: seniorimportance: should knowfreq 54%

basics

~20 s

That message means Helm applied everything and then gave up waiting; Kubernetes rejected nothing. Decide by watching the workload converge: progressing replicas mean the timeout was too tight, a parked rollout means the workload will never be ready.

open as a page

Helm 3's bug-fix support has ended. How would you sequence a move to Helm 4 across your CI images and hundreds of live releases?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat it as a client rollout, not a data migration: releases and charts are adopted in place. Pin the CLI version per image, sweep automation for changed flags and plugin wiring, roll out by blast radius, and schedule the switch to server-side apply as a separate, per-release change.

open as a page

What does `helm version` report, and what does it tell you about the cluster?

level: juniorimportance: nice to knowfreq 26%

basics

~20 s

helm version reports build information about the binary on your PATH only: its version, git commit, tree state and Go version. It says nothing about the cluster, because Helm has had no in-cluster component since Helm 3.

open as a page

Why does `helm get manifest` omit a release's hook resources, and which command shows them?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Helm splits a render into ordinary manifests and hook manifests - anything carrying a helm.sh/hook annotation. Only the ordinary ones become the release manifest that helm get manifest prints; the hooks are stored separately and printed by helm get hooks.

open as a page

After moving CI to Helm 4, why does helm plugin install now fail for a plugin that installed fine under Helm 3?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Helm 4 verifies plugin provenance on install by default: helm plugin install --verify is true unless you say otherwise. An unsigned plugin that Helm 3 installed silently now fails, and the opt-out is --verify=false — there is no --allow-insecure-plugins flag.

open as a page

How would you plan a fleet-wide sweep of Helm releases before a removed API version breaks their upgrades?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Inventory every release's stored manifest for the doomed group and version, then upgrade each affected release with a corrected chart while the old version still resolves. Doing it before removal turns a fleet of repairs into routine upgrades.

open as a page

Before a production `helm upgrade`, what must a change preview prove, and where does it stop helping?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

A preview should prove three separate things: the delta against what Helm last applied, that the cluster would accept the objects, and that nothing drifted since the last upgrade. It cannot prove the rollout will be healthy, and it never prevents the next out-of-band edit.

open as a page

Would you set --rollback-on-failure on every helm upgrade your CI runs, across a dozen service charts?

level: principalimportance: nice to knowfreq 29%

basics

~20 s

Not uniformly. Automatic rollback restores service quickly but erases the pods, events and hook output you triage from, and it cannot undo a hook's side effects. Decide per chart, and capture diagnostics before anything reverts.

open as a page