What does helm upgrade --rollback-on-failure do, and how does --atomic relate to it?
answer
- It decides what the cluster looks like after a failure
- Same flag, two names, one deprecated
- The exit code is still non-zero
- History gains revisions, it never loses them
- Restores the previous revision's stored manifest
basics
~10 s--rollback-on-failure makes helm upgrade restore the previous release revision automatically when the upgrade fails, so no half-applied change is left running. In Helm 4 --atomic is a deprecated alias of that same flag.
solid answer
~50 s`helm upgrade --rollback-on-failure` turns a failed upgrade into an automatic `helm rollback`: if the upgrade errors, a hook fails, or the new objects never become ready within the timeout, Helm restores the manifest of the previous revision instead of leaving the release half-changed. The command still exits non-zero, so CI sees the failure — the difference is what the cluster looks like afterwards. `--atomic` is the Helm 3 name for this behaviour and survives in Helm 4 only as a deprecated alias, so new automation should be written with `--rollback-on-failure`. Two consequences matter in practice. Setting it also makes Helm actually wait for the release to come up rather than returning as soon as the objects are accepted, and the restore is applied as a *new* revision on top of the failed one — Helm never rewrites history, it appends to it.
code
bash · 9 lines# Helm 4 spelling: restore the previous revision if this upgrade fails
helm upgrade chat-fanout ./chat-fanout \
-n team-atlas \
-f team-atlas-values.yaml \
--rollback-on-failure \
--timeout 4m17s
# Same behaviour, deprecated Helm 3 spelling still accepted
helm upgrade chat-fanout ./chat-fanout -n team-atlas --atomicgo deeper
Be ready to say in one sentence what the flag does and give its current name. Knowing that --atomic is now just a deprecated alias of --rollback-on-failure is the detail interviewers listen for, because most scripts you will read still use the old name.
Explain the mechanics: which failures trigger it, that the restore re-applies the previous revision's stored manifest, and that the result is an extra revision rather than a rewound one. Be able to say why the flag implies waiting.
Show that you know what the rollback cannot reverse — hook side effects, migrated data, CRDs from crds/ — and that the command still exits non-zero. Interviewers want to hear you treat it as damage limitation, not as a safety net that makes upgrades risk-free.
Own the policy question: which environments and which chart shapes get automatic rollback, and how the team keeps forward-fix and auto-rollback from fighting each other. Be ready to argue where the flag is actively wrong.
## The problem the flag solves A plain `helm upgrade` is not a transaction. Helm renders the chart, sends the resulting objects to the API server, records a new revision, and returns. If the API server rejects one object halfway through the set, or a pre-upgrade hook exits non-zero, or the new Deployment rolls out pods that crash-loop, you are left with a release whose stored revision says one thing and whose live workloads are a mixture of old and new. Somebody now has to decide, by hand and usually at a bad moment, whether to go forward or back. `--rollback-on-failure` makes Helm make that decision for you, in the same command, in the "go back" direction. ## What it actually does When the flag is set and the upgrade fails, Helm performs a rollback to the revision that was deployed before this upgrade started. The rollback re-applies the *stored manifest* of that previous revision — the exact YAML Helm kept in the release record — so the cluster returns to the object shapes and the values that revision was installed with. Three things are worth being precise about: 1. **The command still fails.** `--rollback-on-failure` is not a way to make a broken deploy look green. Helm reports the original error and exits non-zero; CI must still treat the run as a failed deploy. The flag changes the *state left behind*, not the verdict. 2. **The rollback is a new revision, not an erasure.** If the release was at revision 12 and the upgrade to 13 fails, you end with revision 13 recorded as failed and revision 14 holding the restored content of 12. Helm's history is append-only, which is what lets you see afterwards that an automatic rollback happened at all. 3. **It only covers what Helm applies.** Rolling the manifest back does not undo side effects. A schema migration that a pre-upgrade hook already ran, rows a Job already deleted, data written to a volume, or a CustomResourceDefinition installed from the chart's `crds/` directory (which Helm installs once and never upgrades or removes) all stay exactly as the failed attempt left them. ## The naming, and why it changed In Helm 3 this behaviour was spelled `--atomic`, and the name promised more than the mechanism delivers — "atomic" suggests a database transaction with no observable intermediate state, whereas Helm genuinely applies the new objects, observes them failing, and then applies the old ones again. Clients watching the cluster see all of that. Helm 4 renamed the flag to `--rollback-on-failure`, which describes the actual behaviour, and kept `--atomic` as a deprecated alias so existing pipelines keep working. Write the new name in anything you author today; recognise the old one, because most scripts and blog posts in the wild still say `--atomic`. ## The waiting side effect The flag is close to useless without waiting, so Helm couples the two: setting `--rollback-on-failure` defaults `--wait` to the `watcher` strategy. Without waiting, "failure" can only mean "the API server rejected something" or "a hook failed" — the most common real failure, a new version that deploys cleanly and then crash-loops, would never be seen because Helm would have exited successfully long before. Helm 4 does not wait for workloads unless asked (an omitted `--wait` means `hookOnly`), so this defaulting is what makes the flag catch readiness failures at all. ## Where it does not apply The flag's name says *rollback*, but there is nothing to roll back to on a first install. If you pass it to `helm install` (or to `helm upgrade --install` for a release that does not exist yet) and the install fails, Helm removes the release instead — a very different outcome, and the one that surprises people. ## A concrete shape A multi-tenant chart installed once per team namespace is a good example of where the flag earns its place: each tenant upgrade is independent, the blast radius is one namespace, and an operator upgrading forty namespaces in a loop does not want to hand-inspect the six that failed. With `--rollback-on-failure`, the failures self-restore to the last known-good revision and the operator reads the exit codes. ```bash helm upgrade chat-fanout ./chat-fanout \ -n team-atlas \ -f team-atlas-values.yaml \ --rollback-on-failure ``` If the new chat fan-out pods never reach readiness, this command restores the previous revision and exits non-zero. If they come up healthy, it behaves like any other upgrade.
- Does --rollback-on-failure make the helm upgrade command succeed when the deploy was bad?No. Helm reports the underlying error and exits non-zero; the flag only changes the state left in the cluster. A pipeline that treats the run as green because the cluster looks healthy again has misread it — the new version was rejected, and something still has to be fixed before the next attempt.
- After an automatic rollback, what does the release's revision number look like?It goes up, not down. The failed attempt is recorded as its own revision with a failed status, and the restored content lands as the next revision after it. Helm's release history is append-only, so 'we are back on the old manifest' and 'we are back on the old revision number' are different statements.
- Which parts of a failed upgrade does the automatic rollback not undo?Anything that is not a Helm-applied manifest. Work already done by hooks or Jobs — a database migration, deleted rows, files written to a volume — persists, and CustomResourceDefinitions installed from the chart's crds/ directory are never upgraded or removed by Helm at all. The rollback restores object definitions, not effects.
It is less like a database transaction and more like a decorator who paints the room, sees you hate it, and repaints it the old colour before leaving — the wrong colour was genuinely on the wall for a while, and the paint on the carpet stays.
saying these in an interview costs you the question
- Says the failed upgrade is never applied at all
- Claims the command exits zero after a successful rollback
- Thinks the rollback deletes the failed revision from history
- Believes --atomic and --rollback-on-failure are different behaviours
- Assumes it undoes database migrations run by hooks
- Expects the same flag on install to restore a previous version