Why does Helm's --rollback-on-failure flag also change whether the upgrade waits?
answer
- Helm can only react to what it stays to see
- Two flags, one of them useless alone
- Crash-looping pods arrive after the apply returns
- The flag changes a default, not a hard lock
- A timeout is treated as the failure
basics
~20 sHelm can only roll back failures it has observed. Setting --rollback-on-failure therefore defaults the wait strategy to watcher, so Helm stays until the release is ready; without waiting it would exit successfully before crash-looping pods ever appeared.
solid answer
~50 sRollback-on-failure is only as good as Helm's definition of failure, and that definition depends entirely on how long Helm sticks around. If Helm returns as soon as the API server has accepted the objects, the only detectable failures are apply errors and failed hooks — a new image that is accepted, scheduled and then crash-loops looks like a clean success. So Helm couples the two: passing `--rollback-on-failure` defaults `--wait` to the `watcher` strategy, which follows the release's resources until they are ready or the timeout expires, and a timeout counts as the failure that triggers the rollback. This matters more in Helm 4 than it did in Helm 3, because an omitted `--wait` in Helm 4 means `hookOnly` — Helm does not wait for workloads unless asked. You can still pass `--wait` explicitly alongside the flag to choose a different strategy, but disabling waiting reduces the rollback to catching only submission-time errors.
code
bash · 10 lines# No waiting for workloads: only apply errors and failed hooks can trigger anything
helm upgrade chat-fanout ./chat-fanout -n team-atlas -f team-atlas-values.yaml
# --rollback-on-failure defaults the wait strategy to watcher, so a rollout that
# never becomes ready inside the timeout counts as the failure that rolls it back
helm upgrade chat-fanout ./chat-fanout \
-n team-atlas \
-f team-atlas-values.yaml \
--rollback-on-failure \
--timeout 7m23sgo deeper
Remember the practical effect: adding this flag makes the upgrade command take as long as the rollout instead of returning immediately. If your pipeline step suddenly got much slower after someone added it, this is why.
Explain the causal chain — no waiting means no observed failure means no rollback — and name what Helm 4 does when --wait is omitted. Being able to separate synchronous apply failures from asynchronous convergence failures is the core of the answer.
Bring the timeout into it. Show that you size the deadline from real rollout durations and that you know a too-short timeout rolls back healthy releases, which is its own kind of outage. Be able to diagnose 'it did not roll back' from first principles.
Frame it as a budget: every deploy now blocks a pipeline slot for the length of a rollout. Be ready to discuss where that cost is acceptable, where you would rather have an external health gate, and how you keep timeouts from being copy-pasted across workloads with very different startup profiles.
## Two knobs that only make sense together Helm has two separate concerns at the end of an upgrade: *how long do I stay and watch?* and *what do I do if what I see is bad?* `--wait` answers the first, `--rollback-on-failure` answers the second, and the second is meaningless without the first. ## What counts as a failure, and when Think of an upgrade as passing through phases: 1. **Render** — the chart is templated and the objects are produced. Errors here (bad template, schema violation) fail before anything reaches the cluster, and there is nothing to roll back because nothing was applied. 2. **Hooks and apply** — pre-upgrade hooks run, then the objects are written to the API server. Failures here are *synchronous*: a hook exits non-zero, or the API server rejects an object. Helm sees these no matter what. 3. **Convergence** — the controllers do their work. A Deployment creates a new ReplicaSet, pods are scheduled, images are pulled, containers start, probes pass or do not. Failures here are *asynchronous* and arrive seconds to minutes after the apply. Most of the failures you actually care about live in phase 3: an image tag that does not exist, a config value that makes the process exit on boot, a readiness probe that never passes, a missing Secret key. None of those are visible at the moment Helm finishes applying. If Helm does not wait, it exits successfully at the end of phase 2. The upgrade is recorded as deployed, the command returns zero, and `--rollback-on-failure` never fires — because from Helm's point of view nothing failed. The pipeline goes green while the chat fan-out service crash-loops. ## The coupling Rather than let people configure that footgun, Helm makes the flags interact: setting `--rollback-on-failure` defaults `--wait` to `watcher`. The `watcher` strategy is the event-driven readiness waiter — it follows the release's resources and reports when they have converged, rather than exiting the moment the write succeeded. Now phase 3 has an outcome Helm can see: either the resources become ready inside the timeout, or the timeout expires and Helm treats that as the failure that triggers the restore of the previous revision. The consequence people trip over is a *change in shape*, not just in duration. Without the flag, `helm upgrade` is a fire-and-forget submission that returns in seconds. With it, the command's runtime becomes the rollout's runtime, and its timeout becomes the deadline for the whole release to converge. A pipeline step that used to take four seconds can take minutes, and a timeout that is shorter than a genuinely slow-but-healthy rollout will roll back a deployment that was going to succeed. ## Why this is sharper in Helm 4 In Helm 3, `--wait` was a boolean and `--atomic` implied it. In Helm 4, `--wait` takes a *strategy* value, and the default when the flag is omitted is `hookOnly` — Helm waits for hooks and not for workloads. That makes the unwaited case the norm rather than an opt-out, and makes the defaulting behaviour of `--rollback-on-failure` the thing that gives it teeth. A bare `--wait` means `watcher`; the legacy polling waiter from Helm 3 is still available as `legacy` for anyone who needs the old readiness semantics. You can override the strategy explicitly — the flag defaults it, it does not lock it. That is occasionally useful (for example forcing the legacy waiter for a resource whose readiness the newer waiter judges differently) and almost always a mistake if what you choose is a strategy that does not watch workloads, because you then keep the cost of the flag and lose its benefit. ## Reading it in an incident When someone says "we had `--rollback-on-failure` on and it did not roll back", there are three ordinary explanations, and this coupling is behind two of them. Either the failure was outside what Helm watches (the pods were ready and the service was still wrong — readiness is not correctness), or the waiting was overridden or the timeout was so long that the human intervened first, or the run failed at render time and there was nothing applied to restore. Knowing which of the three you are in is the whole diagnosis. ```bash # Returns as soon as the objects are accepted; a crash-looping rollout is invisible here helm upgrade chat-fanout ./chat-fanout -n team-atlas -f team-atlas-values.yaml # Waits for readiness with the watcher strategy, and restores the previous revision if it never arrives helm upgrade chat-fanout ./chat-fanout -n team-atlas -f team-atlas-values.yaml --rollback-on-failure ```
- Which failures would still trigger a rollback even if Helm did not wait at all?Only the synchronous ones: a hook that exits non-zero and an object the API server rejects, such as an immutable field change or a validation error. Everything that happens after the write — image pull failures, crash loops, probes that never pass — is invisible to a command that has already returned.
- What is the risk of pairing --rollback-on-failure with a short timeout?You roll back healthy deployments. The timeout becomes the deadline for the entire release to converge, so a slow image pull, a large rollout or a workload with a long startup probe can exceed it while nothing is actually wrong, and Helm dutifully restores the previous revision. Size the timeout from the worst observed healthy rollout, not the typical one.
- Can you keep --rollback-on-failure but choose a different wait strategy?Yes — the flag sets a default, so an explicit --wait alongside it wins. That is legitimate when you need the legacy waiter's readiness semantics for a particular resource, but choosing a strategy that does not watch workloads leaves you paying for the flag while catching only submission-time errors.
saying these in an interview costs you the question
- Thinks rollback works without Helm waiting
- Says --wait is still a boolean in Helm 4
- Believes an omitted --wait waits for pods
- Assumes a timeout is not treated as a failure
- Claims the flag makes the wait strategy unchangeable
- Confuses pods being ready with the release being correct