skip to content

Would you set --rollback-on-failure on every helm upgrade your CI runs, across a dozen service charts?

level: principalimportance: nice to knowfreq 29%

answer

  1. Availability and evidence pull opposite ways
  2. A revert restores manifests only
  3. Enabling it changes waiting too
  4. Budgets belong to charts, not pipelines
  5. Collect before you revert

basics

~20 s

Not uniformly. Automatic rollback restores service quickly but erases the pods, events and hook output you triage from, and it cannot undo a hook's side effects. Decide per chart, and capture diagnostics before anything reverts.

solid answer

~40 s

`--rollback-on-failure` buys availability and costs evidence. When it fires, the failed revision's pods, their events and any hook output are gone before an engineer looks, so recurring failures become unexplainable. It also cannot undo what a pre-upgrade hook did outside Kubernetes, so a chart with an irreversible migration can be reverted to old code against new data — worse than staying failed. And enabling it changes waiting behaviour: in Helm 4 it defaults `--wait` to `watcher`, so charts that previously never waited now can fail and revert on a slow rollout. My default is per-chart: on for stateless services with measured timeout budgets, off for charts with migrating hooks, and in both cases a pipeline step that collects `helm status`, `helm get hooks` and pod diagnostics before any revert.

code

bash · 10 lines
bash
set -euo pipefail

if ! helm upgrade digest-builder ./charts/digest-builder -n mail \
      --wait=watcher --timeout 12m; then
  helm status digest-builder -n mail            > artifacts/status.txt
  helm get hooks digest-builder -n mail         > artifacts/hooks.yaml
  kubectl get events -n mail --sort-by=.lastTimestamp > artifacts/events.txt
  helm rollback digest-builder -n mail --wait=watcher --timeout 12m
  exit 1
fi

go deeper

for a junior

Know that Helm can revert a release automatically when an upgrade fails, and that this is a choice someone made rather than default behaviour.

for a middle

Be ready to say what a revert actually restores — the previous revision's manifests — and what it leaves alone, including anything a hook did outside Kubernetes.

for a senior

Show the operational cost: an automatic revert removes the pods and events you would have read, so recurring failures become unexplainable unless the pipeline collects diagnostics first.

for a principal

Own the policy across charts: which services trade evidence for availability, how per-chart timeout budgets are derived and reviewed, and how you keep a dozen pipelines from drifting into a dozen different answers.

This is a policy question wearing a flag's clothes, and the interesting part is that the obvious answer — "always roll back, it protects production" — trades away the thing that stops the same outage recurring. ## What the flag does `--rollback-on-failure` tells Helm that when an upgrade fails it should revert the release to the previous revision as part of the same operation. In Helm 4 this is the flag's real name; `--atomic` survives as a deprecated alias, so a fleet's scripts written against Helm 3 keep working but should be renamed. Setting it also defaults `--wait` to the `watcher` strategy — a consequence worth stating out loud, because it changes behaviour for every chart that previously ran without waiting: those upgrades now observe readiness, can now fail on a slow rollout, and will now revert. Turning the flag on fleet-wide is not one change, it is two. ## The case for it For a stateless service, a failed upgrade that leaves half-applied objects or a crash-looping new version is a live incident, and reverting in the same operation shortens it to the length of one rollout. It also keeps the release record honest: rather than accumulating failed revisions that nobody cleans up, the release returns to a known-good, deployed state that the next upgrade can build on. ## The case against it Three arguments, in rising order of severity. 1. **Evidence.** Triage of a failed upgrade lives in objects that the revert destroys: the pods that would not become ready, their events, the Job a failed hook left behind. If your pipeline reverts before anyone looks, every failure report reads "the deploy failed and then it was fine", and a flaky, intermittent failure never gets diagnosed. The release history preserves the failed revision's manifest and description, which is genuinely useful, but a manifest is not a reason. 2. **Side effects.** Reverting the release reverts the Kubernetes manifests and nothing else. If a pre-upgrade hook ran a schema migration and a later step failed, the automatic rollback puts the *old* application code back in front of *new* schema. That combination can be worse than a failed upgrade that simply never replaced the running pods, and it is a per-chart property: a chart with an irreversible pre-upgrade hook wants human judgement, not an automatic revert. 3. **The previous revision must actually be deployable.** Rollback re-applies the stored manifest of the prior revision, which references images and configuration that must still exist. A fleet that garbage-collects old image tags aggressively has a rollback path that quietly stops working, and the discovery moment is an incident. ## A defensible position Decide per chart, and say so with reasons rather than reaching for uniformity. | Chart | Policy | Why | | --- | --- | --- | | Stateless services whose rollout time you have measured | Enable it | Set `--timeout` per chart from observed rollout duration with headroom, never from a shared default | | Charts with migrating or otherwise irreversible hooks | Leave it off, fail loudly, and page | The correct next action is a decision, not a revert | | A shared library chart consumed by twelve service charts | Consistency of *mechanism* rather than of setting | Express the timeout and failure policy as values the library chart's deployment tooling reads, so each service declares its own budget in one reviewed place instead of twelve pipeline files drifting apart | A budget below the honest rollout time plus automatic rollback produces a service that can never deploy — for an email-digest builder whose replicas drain a queue before reporting ready, a seven-minute rollout under a five-minute default budget is exactly that trap. ## Capture before you revert The strongest version of the answer refuses the binary. Because Helm performs the revert inside the same operation, a pipeline cannot interleave collection with it. So for services that need fast recovery *and* diagnosable failures, do not use the flag: run the upgrade without it, and on failure have the pipeline collect `helm status <release>`, `helm get hooks <release>`, the workload's pods, logs and events into the build's artefacts, and only then run the revert itself as an explicit step. That is a few more lines of pipeline in exchange for every failure being explainable, and it makes the recovery an auditable action rather than an invisible one. ## What to measure Whatever you choose, the fleet-level signals are the same: how often upgrades fail, how many of those failures were deadline expiries rather than genuine rejections, and how many were reverted without anyone reading a diagnosis. A high count in the last bucket means your automation is converting outages into mysteries, and that is the argument for changing the policy — not a preference about flags.

  • Your platform enables automatic rollback fleet-wide and two charts immediately start failing. Why?
    Because the flag also turns waiting on. In Helm 4 setting `--rollback-on-failure` defaults `--wait` to `watcher`, so charts that previously returned as soon as the manifests were accepted now observe readiness. Any chart whose rollout is slower than the timeout budget, or whose workload was quietly never becoming ready, now fails and reverts. The failures were always there; the flag made them visible.
  • How would you decide the timeout budget for each chart rather than sharing one number?
    Derive it from observed rollout durations at production replica counts, including image pull on a cold node and any readiness gate the workload imposes, then add headroom for scale-up. Record it with the chart, not in a pipeline file, so it is reviewed alongside the workload changes that alter it. Revisit it when replica counts or readiness behaviour change, since a budget set once ages badly.
  • If you leave automatic rollback off, what keeps a failed release from lingering?
    An explicit recovery step and an owner. The pipeline should fail loudly, publish the collected diagnostics, and page the service team rather than leaving a failed release for whoever notices. For services that must self-heal, run the revert as its own pipeline step after collection, so the action is logged, attributable and reviewable instead of happening invisibly inside the upgrade.

saying these in an interview costs you the question

  • Enables automatic rollback everywhere without exceptions
  • Thinks a rollback undoes a hook's schema migration
  • Ignores that the flag also enables waiting
  • Shares one timeout budget across every chart
  • Assumes the previous revision's images still exist
  • Never collects diagnostics before reverting

context