skip to content

Would you make --rollback-on-failure mandatory for every production helm upgrade in CI?

level: principalimportance: should knowfreq 36%

answer

  1. It restores manifests, never effects
  2. A property of the chart, not the environment
  3. The timeout stops being a patience setting
  4. Frequent automatic rollbacks are a smell
  5. Reverting after full exposure is a weak guarantee

basics

~20 s

No blanket rule. Automatic rollback fits stateless workloads whose previous revision is still safe to run, and is dangerous for releases with irreversible hook side effects, very slow rollouts, or schema changes the old version cannot read. Decide per chart shape.

solid answer

~50 s

Mandating it sounds like pure safety, but it buys one property — the cluster returns to the last known-good manifest — at four costs: every deploy blocks a pipeline slot for the length of a rollout, the timeout becomes a deadline that can roll back healthy-but-slow releases, evidence disappears on first installs, and the restored revision may be incompatible with side effects the failed attempt already committed. That last one is the real argument. A chart whose pre-upgrade hook migrates a schema forward is not safe to roll back to code that predates the migration; the automatic restore turns a failed deploy into a genuinely broken one. My default is to enable it for stateless services whose previous revision is unconditionally safe, disable it where a hook mutates shared state, size the timeout from the worst healthy rollout rather than the typical one, and treat forward-fix as the policy wherever automatic reversal cannot be proven correct.

code

bash · 10 lines
bash
# Revertible tier: stateless service, previous revision always safe to run
helm upgrade chat-fanout ./chat-fanout -n team-atlas \
  -f team-atlas-values.yaml \
  --rollback-on-failure --timeout 8m11s

# Forward-only tier: a pre-upgrade hook migrates shared state, so reverting the
# manifest would put old code in front of a migrated schema. Wait, do not revert.
helm upgrade chat-archive ./chat-archive -n team-atlas \
  -f team-atlas-values.yaml \
  --wait --timeout 8m11s

go deeper

for a junior

You will not usually set this policy, but know that the flag is on or off per deploy for a reason. If a chart runs a database migration before the new version starts, reverting the manifest afterwards is not automatically safe.

for a middle

Be able to name the costs, not just the benefit: the pipeline now blocks for the rollout, the timeout becomes the failure definition, and the restore does not undo work a hook already did. That is enough to explain why it is not simply always on.

for a senior

Show judgement about which charts get it. Sizing timeouts from real rollout data, splitting first installs from upgrades, and treating a high rollback rate as a quality signal are the things that show you have operated this rather than read about it.

for a principal

Own the policy shape and its exemption path: how charts are classified, who signs off on moving one, and how you stop 'always on' from being sold as a replacement for progressive delivery. Be ready to defend deliberately turning it off for a critical service.

## What the mandate is actually proposing "Always pass `--rollback-on-failure`" is a proposal to make one recovery strategy — revert — the automatic response to every deploy failure, chosen by a flag rather than by a person. Whether that is right depends on a property of each chart that the flag itself cannot inspect: *is the previous revision still a correct thing to be running after this attempt partially happened?* ## The case for making it the default - **Bounded exposure.** Without it, a failed upgrade leaves the release in a mixed state for however long it takes a human to notice and decide. With it, the window is the timeout. - **Uniform end state.** Every failed deploy ends in the same place: the last known-good manifest. That is enormously easier to write runbooks against than "it depends what failed". - **It scales with fan-out.** A multi-tenant chart installed once per team namespace, upgraded across dozens of namespaces in one run, cannot have a human triage each failure. Self-restoring failures plus exit codes is the only shape that works. - **It forces waiting.** Because the flag defaults the wait strategy to watcher, adopting it also ends the pattern of pipelines reporting success before the rollout has converged. Some teams want the flag mostly for that. ## The case against making it universal **Irreversibility.** This is the one that actually decides it. Rolling back restores *manifests*, not *effects*. If a pre-upgrade hook has already migrated a database schema, the automatic restore puts the old application code back in front of a new schema — a state nobody tested, and often worse than the failed upgrade it replaced. The same applies to a Job that consumed a queue, an external system that was called, or data already written in the new format. For those charts, forward-fix is the only defensible policy, and the flag encodes the wrong one by default. **Timeout as verdict.** With the flag on, the timeout stops being a patience setting and becomes the definition of failure. Size it from typical rollouts and you will roll back healthy releases whose images were slow to pull or whose startup probes are long; size it very generously and the pipeline slot is held for that long on every failure. Neither is free, and a single organisation-wide number will be wrong for both the fastest and the slowest workload. **Evidence destruction.** On a release that does not exist yet the flag uninstalls rather than restores, taking the failed pods and their logs with it. Teams bringing up new services hit this repeatedly, conclude the tooling is hostile, and start disabling the flag everywhere — including where it was doing good. **Failure during recovery.** The restore is itself an apply that can fail, and a failed automatic rollback leaves a release in a worse and less-expected state than the failure that triggered it. Rare, but it is the scenario where the mandate has made things harder rather than easier. ## How I would actually write the policy 1. **Classify charts, not environments.** The question is "is reverting safe for this thing?", which is a property of the chart's hooks and its data model. Two tiers is usually enough: *revertible* (stateless services, config-only releases, anything whose previous revision is unconditionally safe) and *forward-only* (anything with a state-mutating hook or a schema contract). Default the first tier on, the second off, and require a reason to move a chart between tiers. 2. **Make the timeout a per-chart value with a floor.** Derive it from observed healthy rollout durations with headroom, not from a shared default. A chart whose pods take minutes to warm caches needs its own number. 3. **Separate first-install from upgrade.** New releases install without the flag, with an explicit diagnostic capture and cleanup step. Established releases upgrade with it. 4. **Keep the human path.** Automatic rollback is a first response, not an incident process. The pipeline must still fail loudly, page, and leave the failed revision visible in history for the follow-up. 5. **Watch the metric that matters.** Count automatic rollbacks per chart. A chart that rolls back regularly is not being protected by the flag, it is being masked by it — the flag is absorbing a quality problem that should be fixed upstream in tests or in the rollout strategy. ## The honest summary Automatic rollback is a good default for the majority of charts and a genuinely bad one for the minority that change shared state. A mandate with no exemption path optimises for the common case and quietly creates the worst incidents; a mandate with a documented exemption path, per-chart timeouts, and a separate rule for first installs is close to right. What I would not accept is "always on" being treated as a substitute for progressive delivery — the flag reverts a bad release *after* it has been fully applied, which is a strictly weaker guarantee than never sending it all the traffic in the first place.

  • Which single chart property most strongly argues against automatic rollback?
    A hook or Job that mutates shared state irreversibly — most commonly a schema migration. Restoring the previous manifest puts code that predates the migration back in service against data that has already moved, which is an untested combination and frequently worse than the failed upgrade. Those charts need forward-fix, not automatic reversal.
  • How would you choose the timeout that goes with the flag?
    Per chart, from the slowest healthy rollout you have actually observed, plus headroom for a cold image pull. A single organisation-wide number is simultaneously too short for the slow workloads — rolling back releases that were fine — and needlessly long for the fast ones, since the pipeline holds the slot for the whole window on failure.
  • If a chart triggers automatic rollbacks most weeks, what do you conclude?
    That the flag is masking a quality problem rather than solving one. Something upstream is wrong: flaky readiness, undersized resources, insufficient pre-deploy verification, or a rollout that is genuinely too slow for its deadline. I would treat the rollback rate as a per-chart signal and fix the cause rather than widening the timeout.
  • Is automatic rollback a substitute for a progressive rollout?
    No. The flag reverts a release that has already been applied in full, so every user was exposed before the revert began. Progressive exposure limits who sees the bad version at all. Automatic rollback is a cheap floor under the failure case; it does not give you the property that a staged rollout gives you.

saying these in an interview costs you the question

  • Treats automatic rollback as risk-free in every case
  • Ignores hooks that already migrated shared state
  • Uses one organisation-wide timeout for every chart
  • Calls it equivalent to a progressive rollout
  • Assumes the automatic rollback itself cannot fail
  • Counts frequent rollbacks as the flag working well

context