A helm install --rollback-on-failure fails. What does Helm leave behind, and why is that a problem?
answer
- There is nothing behind a first install
- Tidy and diagnosable pull in opposite directions
- The pods you needed are deleted with the release
- upgrade --install picks its behaviour by what exists
- Some objects survive the cleanup anyway
basics
~20 sOn a first install there is no previous revision to restore, so Helm uninstalls the release instead: the objects and the release record are removed. The cleanup is deliberate, but it also deletes the failed pods you needed to diagnose.
solid answer
~50 sThe flag behaves asymmetrically. On `helm upgrade` it restores the previous revision; on `helm install` — and on `helm upgrade --install` when the release does not exist yet — there is no previous revision, so Helm removes the release entirely. Afterwards `helm list` shows nothing, `helm history` has nothing, and the failed pods, their events and their logs are gone with the objects that owned them. That is the whole point when you are installing a multi-tenant chart into forty team namespaces and want no debris, and it is exactly wrong when you are bringing a new service up for the first time and need to know *why* it failed. The usual working rule: keep the flag for upgrades of things already running, and for a first install either drop it and clean up by hand after inspecting, or reproduce the failure with a render-and-validate run before installing at all.
code
bash · 9 lineshelm install chat-fanout ./chat-fanout \
-n team-atlas --create-namespace \
-f team-atlas-values.yaml \
--rollback-on-failure --timeout 6m43s
# After the failure, there is nothing left to look at:
helm list -n team-atlas # no release
helm history chat-fanout -n team-atlas # no history to read
kubectl get pods -n team-atlas # the failed pods are already deletedgo deeper
Know the headline: with this flag, a first install that fails leaves no release at all, so there is nothing in helm list afterwards. If you were expecting to look at the broken pods, install without it.
Explain why the behaviour differs between install and upgrade — no previous revision means nothing to restore — and that helm upgrade --install picks its path from whether the release already exists rather than from the subcommand name.
Demonstrate the operational tradeoff: cleanliness versus evidence, and a concrete plan for keeping diagnostics when the flag is on. Mentioning what survives the uninstall — CRDs from crds/, kept resources, external side effects — is what separates a lived answer from a read one.
Own the policy split across environments and commands, and the retry discipline that goes with it. Be ready to say how you stop teams from looping a failing install, and how first-run failures get triaged when the tooling is designed to erase them.
## The asymmetry One flag, two outcomes, decided by whether a previous revision exists: - **Upgrade of an existing release** — restore the previous revision's stored manifest. The service keeps running on the old version. - **Install of a new release** (including `helm upgrade --install` when nothing is installed yet) — uninstall. There is no old version to go back to, so "back" means "gone". Both are defensible readings of "leave no half-applied change behind". They just feel very different at 3 a.m. ## What "gone" actually means Say you are rolling a multi-tenant chart out to a new team namespace, driving it from an 11-value override file, and the chat fan-out Deployment never becomes ready because one of those 11 values names a Secret key that does not exist: ```bash helm install chat-fanout ./chat-fanout \ -n team-atlas --create-namespace \ -f team-atlas-values.yaml \ --rollback-on-failure --timeout 6m43s ``` Helm waits, the pods never pass readiness, the timeout expires, and Helm uninstalls. Now: - `helm list -n team-atlas` shows nothing; there is no release. - `helm history chat-fanout -n team-atlas` has nothing to show — the release record is gone, so you cannot even see the manifest that was attempted. - The Deployment, its ReplicaSet and its pods have been deleted, so `kubectl logs` and `kubectl describe pod` have nothing to read. - Namespace events survive briefly, but they expire, and they are the thinnest possible evidence. You are left with the terminal output of the failed command and the values file, and you get to guess. The cost is not the missing release — you can install it again — it is the missing *evidence*, and evidence about a startup failure is only available while the failed pods exist. ## Why the flag still earns its place on install The cleanup is genuinely valuable at scale. A failed install without the flag leaves a release in a failed state whose objects are still there: half a chart's worth of Services, ConfigMaps, maybe a PVC, plus a release record that a subsequent `helm install` with the same name will refuse to write over. For an operator installing the same chart once per team namespace across many namespaces, that debris multiplies and someone has to clean it namespace by namespace. The flag turns a partial install into no install, which is the state the next attempt wants to start from. So the tension is real: *tidy* and *diagnosable* are opposites here. ## How practitioners resolve it A few habits, roughly in order of how often they are used: 1. **Split the policy by command.** Automatic cleanup on upgrades of running services; first installs run without it, and the pipeline cleans up explicitly after capturing diagnostics. This is the common shape because the risk profiles genuinely differ — a failed upgrade endangers live traffic, a failed first install endangers nothing. 2. **Fail earlier, where nothing is applied.** A server-side dry run catches schema violations, bad API versions and rejected objects before any of this matters. It cannot catch a container that starts and dies, but it removes a large class of first-install failures from the picture entirely. 3. **Capture before you lose it.** If the flag must stay on, have the pipeline collect pod logs and describe output *during* the wait, not after the command returns — by then the objects are already deleted. 4. **Retry deliberately, not blindly.** Because the release is fully removed, re-running the install is clean and safe, which tempts people into a retry loop. A retry loop over a deterministic failure just burns the timeout repeatedly and destroys the evidence each time round. ## The things that survive the uninstall anyway "Uninstall" is not "the namespace is as it was". CustomResourceDefinitions installed from the chart's `crds/` directory are never deleted by Helm. Objects annotated `helm.sh/resource-policy: keep` are left in place. Anything a hook created outside the release, or wrote to an external system, is untouched. So the outcome of a failed install with the flag set is usually "the release is gone and a few things are not", and a second attempt can meet a leftover object it did not expect to find. ## The interview answer in one line On upgrade it protects the running service; on install it protects the namespace's tidiness at the cost of your ability to debug — and knowing which of the two you are invoking is a property of whether the release already exists, not of which subcommand you typed.
- Does helm upgrade --install with the flag behave like an install or an upgrade?Whichever it actually performed. If the release already exists it upgrades, and a failure restores the previous revision; if nothing is installed the same command performs an install, and a failure removes the release. The subcommand you typed does not decide it — the presence of an existing release does.
- What would you do differently to keep the diagnostics from a failed first install?Run the first install without the flag, wait for readiness, and on failure capture describe output and container logs before uninstalling explicitly. If the flag has to stay on, collect diagnostics during the wait from a parallel step, because once the command returns the objects are already deleted.
- After a failed install cleans itself up, is the namespace back to exactly its previous state?Usually not quite. CustomResourceDefinitions installed from the chart's crds/ directory are never removed by Helm, objects annotated helm.sh/resource-policy: keep are left in place, and anything a hook did to an external system is untouched. A retry can therefore meet leftovers it did not create.
saying these in an interview costs you the question
- Expects a failed install to roll back to an earlier chart version
- Says the release stays visible in helm history afterwards
- Thinks the failed pods remain available for inspection
- Assumes the namespace is restored exactly as before
- Believes upgrade --install always takes the upgrade path
- Retries the same install repeatedly instead of diagnosing