skip to content

Why do deploy jobs run helm upgrade --install rather than helm install?

level: middleimportance: must knowfreq 74%

answer

  1. Neither plain command is re-runnable
  2. Release identity is name plus namespace
  3. One run, one revision, even unchanged
  4. History is a rolling window of ten
  5. Exit zero is not rolled out

basics

~20 s

helm install fails when the release name already exists and helm upgrade fails when it does not, so neither is safe to re-run. helm upgrade --install picks whichever applies, letting one deploy job serve a first deploy and every later one.

solid answer

~40 s

A deploy job cannot know whether the release already exists: the first run in a new namespace is an install, the next forty are upgrades. `helm upgrade --install NAME CHART` (shorthand `-i`) looks for a release with that name in the target namespace, upgrades it if found and installs it if not, so the same job works on a brand-new environment and on a long-lived one. Two consequences matter in a pipeline. Every successful run writes a new revision as a `sh.helm.release.v1.<name>.v<rev>` Secret even when nothing changed, and `--history-max` (default 10) prunes the oldest, so busy pipelines keep only a rolling window. And a zero exit code means the manifests were applied, not that the pods are healthy - Helm does not wait for workloads unless you ask it to.

code

bash · 5 lines
bash
helm upgrade --install ledger charts-repo/payments-ledger \
  --version 2.14.3 \
  -n payments --create-namespace \
  -f envs/prod.yaml \
  --history-max 20

go deeper

for a junior

Remember why the flag exists: plain install fails when the release is already there and plain upgrade fails when it is not, so one job needs both behaviours. Being able to write the command with a namespace is the expectation.

for a middle

Explain the mechanics beneath it: the lookup is by name and namespace, each run writes a new revision Secret whether or not content changed, and history is bounded by --history-max.

for a senior

Show the operational judgment: set the job timeout above the slowest realistic upgrade, serialise deploys per release, and be explicit about whether the job's success means applied or means healthy.

for a principal

Own the contract the deploy step publishes to the rest of the organisation - what a green deploy asserts, how long the rollback window is, and who is responsible for observing rollout health when Helm returns early.

`helm install` refuses to run when a release with that name already exists in the namespace, and `helm upgrade` refuses to run when it does not. A deploy job cannot know which situation it is in. The very first deploy of a service is an install; the next forty deploys are upgrades; and standing up a fresh environment, or an engineer running `helm uninstall` to clean something up, resets the count. Hard-coding either command gives you a job that works until it does not. `helm upgrade --install NAME CHART`, shorthand `-i`, collapses the two paths. Helm looks for a release with that name in the target namespace: if it finds one, it upgrades; if it does not, it installs. That is the single line of a deploy job that is safe to re-run, and it is why almost every pipeline in the wild uses it. ## Scope, and the flags that come with it The lookup is by release name **and** namespace. `ledger` in `payments` and `ledger` in `payments-staging` are two unrelated releases, so a deploy job must always pass `-n`, never rely on the ambient namespace. `--create-namespace` covers the bootstrap case where the namespace itself does not exist yet, which is the usual reason a first deploy into a new environment fails. Note that the two branches are not identical: on the install branch there is no previous release, so anything that reconciles new input against what a previous revision stored simply has nothing to reconcile. ## Every run makes a revision Each successful upgrade writes a new release revision, stored as a Secret named `sh.helm.release.v1.<name>.v<rev>` in the release namespace. Helm does not detect a no-op and skip: a pipeline that redeploys on every merge to `main` will cheerfully produce revision 43 whose stored manifest is byte-identical to revision 42. `--history-max`, default 10, prunes the oldest superseded revisions as new ones arrive, so a busy deploy job keeps a rolling window - the ten most recent revisions are inspectable and can be rolled back to, and anything older is gone. Teams that deploy dozens of times a week and expect to roll back to *last month* are usually surprised by this. ## A zero exit code is not a healthy workload `helm upgrade --install` returning 0 means the manifests were applied and any hooks Helm blocked on completed. It does not by itself mean the new pods started. Helm does not wait for workload readiness unless asked, so a deploy job whose only success criterion is the exit code will report a green deploy while the payments-ledger pods crash-loop on a bad config value. If the job's contract is *the workload is healthy*, ask Helm to wait; if it is *the desired state has been submitted*, say so explicitly, and let something else observe the rollout. ## The two failure modes pipelines actually hit The first is a failed first install. If install fails and nothing rolls it back, the release exists at revision 1 with status `failed`, and there is no deployed revision for a later `upgrade --install` to upgrade from, so subsequent runs error out rather than repairing themselves. The fix is to uninstall the failed release and install again - which is why some pipelines add an automatic rollback on failure precisely so a failed first attempt does not leave a corpse. The second is a killed job. Suppose the payments-ledger deploy takes about seven minutes and the job's step timeout is six. The runner kills the Helm process mid-upgrade; nothing writes a terminal status, so the release is left in a pending state, and the next pipeline run is refused with the message that another operation - install, upgrade or rollback - is in progress. That lock is pessimistic and does not expire on its own; a human clears it. The lesson for pipeline design is that the job's timeout must exceed the slowest realistic upgrade, and that two runs must never deploy the same release concurrently. ## What interviewers listen for A weak answer stops at *it installs if missing, upgrades if present*. A strong one adds that the release identity is name plus namespace, that every run consumes a revision slot bounded by `--history-max`, and that exit 0 is a statement about applying manifests rather than about running pods.

  • Does a zero exit code from helm upgrade --install mean the new version is serving traffic?
    No. It means the manifests were applied and the hooks Helm blocked on finished. Helm does not watch your Deployments for readiness unless you ask it to wait, so pods can still be pulling an image or crash-looping when the command returns. Decide what the job promises: either ask Helm to wait and give it a timeout longer than a real rollout, or state that the job only submits desired state and let a separate check confirm health.
  • A pipeline redeploys on every merge. What does --history-max do to that release's history?
    Each successful run writes a new revision Secret, even if nothing changed. `--history-max`, default 10, prunes the oldest superseded revisions as new ones are written, so the release keeps a rolling window rather than growing without bound. That is good for etcd and bad for the assumption that you can roll back to any past state: after ten deploys, the revision you wanted is gone. Raise the value deliberately if your rollback window needs to be longer.
  • Two pipeline runs for the same release start at once. What happens?
    The second one is refused: Helm reports that another operation - install, upgrade or rollback - is already in progress for that release. The lock is pessimistic and is not released by a timeout, so if the first process is killed rather than finishing, the release stays in a pending state until somebody clears it. Pipelines avoid this by serialising deploys per release rather than by retrying into the lock.

saying these in an interview costs you the question

  • Thinking upgrade --install is idempotent and skips no-op runs
  • Assuming release names are unique cluster-wide
  • Treating exit code zero as proof pods are running
  • Expecting a killed Helm process to unlock itself
  • Believing a failed first install self-heals on the next run
  • Never passing -n and relying on the ambient namespace

context