skip to content

Chart Granularity

How much one chart should ship: an umbrella release covering a whole system versus one release per service, and what a shared revision history costs when a single subchart fails.

part ofHelmoverview, primer and where to startread it →
on this pageshow

questions

4

Why does one failing subchart in an umbrella Helm release affect every other service in it?

level: middleimportance: must knowfreq 62%

answer

  1. Helm records status per what, exactly?
  2. Nothing is undone unless you asked
  3. Automatic rollback reverts the innocent too
  4. One failed status, fourteen blocked teams
  5. --history-max 10 shared across everyone

basics

~20 s

An umbrella install is one release, so an upgrade is one operation with one outcome: work already applied stays live, the release is marked failed, and an automatic rollback reverts every service, not just the broken one.

solid answer

~50 s

Helm tracks state per release, not per subchart. An upgrade of an umbrella renders and applies the whole thing; if one subchart's resource fails to apply, or fails to become ready when you asked Helm to wait, the release status becomes `failed` while everything applied before the failure stays live in the cluster — Helm does not undo work on its own. If you ran with `--rollback-on-failure` (Helm 4's name for the old `--atomic`), Helm reverts the *entire* release to the previous revision, including the thirteen services that deployed perfectly. Either way the next team to ship must upgrade the same release out of a failed or reverted state, and `helm rollback` cannot exclude their service. That coupling — one status, one revision history, one rollback unit — is the real cost of the umbrella.

code

bash · 7 lines
bash
# Detect failure and revert automatically - for the WHOLE release
helm upgrade platform ./platform -n platform \
  --rollback-on-failure --timeout 9m30s

# What the release records afterwards
helm history platform -n platform
helm status platform -n platform

go deeper

for a junior

Remember that an umbrella install is one release with one status. If an upgrade fails, what was already applied stays in the cluster and the release is marked failed — Helm does not tidy up by itself.

for a middle

Explain the mechanics: one release record per revision, a status that covers the whole rendered manifest, --rollback-on-failure reverting everything, and Helm 4's hookOnly default meaning readiness is not checked unless you ask.

for a senior

Show the operational consequence. Say who is blocked while the release sits failed, why a shared --history-max of 10 can prune the revision you wanted, and how you would decide between fixing forward and automatic rollback during an incident.

for a principal

Frame it as a design constraint you chose. Helm's units of failure, rollback and concurrency are all the release, so granularity is the only real lever — argue when merging blast radii is acceptable and when it is an outage waiting to be scheduled.

### The unit of state is the release Helm records a release, not a component. Each `helm upgrade` on a release named `platform` writes a new record — the Secret `sh.helm.release.v1.platform.v<rev>` in the release namespace — containing the values used, the full rendered manifest, and a status. There is no per-subchart status field anywhere in that record, because as far as Helm is concerned the umbrella *is* the chart. So any question of the form "what happened to the other services" resolves to "whatever happened to the release". ### What actually happens when a subchart fails Take a `platform` release at revision 26 containing a document-indexing job alongside thirteen other services, upgraded to chart version 4.11.3. Three things can go wrong, and they behave differently. **The manifest is rejected.** If a subchart renders invalid YAML or a resource the API server refuses, the upgrade fails during apply. Resources applied before that point are already in the cluster and stay there. The release lands in `failed`, and the cluster is now a partial mixture of revisions 26 and 27. **A workload never becomes ready.** This one surprises people. In Helm 4, `--wait` is strategy-valued — `watcher`, `legacy` or `hookOnly` — and **omitting it entirely defaults to `hookOnly`**, so Helm waits for hooks and then returns. A document-indexing job whose pods crash-loop after a bad config change does not fail the upgrade at all: `helm status` says `deployed` while the workload is broken. You only get readiness-based failure when you ask for it with `--wait` (which means `watcher`) or `--wait=legacy`, and then a single unready workload from any subchart fails the whole upgrade, however healthy the other thirteen are. **A hook fails.** A failed pre-upgrade hook aborts before the main manifest is applied, so the whole upgrade — every service — simply does not happen. ### Automatic rollback makes the coupling total `--rollback-on-failure` is the Helm 4 flag (`--atomic` remains as a deprecated alias); setting it also defaults the wait strategy to `watcher`, so Helm now genuinely watches workloads. On failure it rolls the release back to the previous revision. That is the behaviour people want, and it is exactly where the umbrella bites: the rollback restores revision 26's stored manifest for *everything*. A catalog change that shipped fine in the same upgrade is reverted with the broken indexer, and the team that shipped it finds out from a graph, not from Helm. And note that rollback is a *forward* operation: it creates revision 28 whose content is revision 26's manifest. Nothing is erased; the history just grows. ### Everybody now shares the recovery After a failure, the release carries `failed` status and the next deploy of any service is an upgrade of the same release on top of it. If an operation was interrupted rather than completing, the release can be left in a pending state and Helm's pessimistic lock refuses new operations with "another operation (install/upgrade/rollback) is in progress" — one team's stuck upgrade blocks fourteen services' deploys. There is no command that scopes an upgrade or a rollback to one subchart; the flags select values and behaviour, never a component. A quieter version of the same coupling is history retention. `--history-max` defaults to 10, so a release keeps ten revisions. In an umbrella that everyone deploys through, ten revisions can be a single busy afternoon — meaning the revision you actually wanted to roll back to has already been pruned away, and it was pruned by other teams' deploys. ### What you can and cannot do about it Within an umbrella you can reduce the odds and the ambiguity: run upgrades with an explicit wait strategy so failures are detected rather than discovered later, decide deliberately whether automatic rollback is better than a partial state for your system, raise `--history-max` so the useful revision survives, and review `helm template` output in CI so rendering failures never reach the cluster. What you cannot do is make the release granular. If the answer a team needs is "roll back my service and only my service", the fix is not a flag — it is a different chart boundary, which means splitting the umbrella into per-service releases. The crisp interview formulation: Helm's unit of state, of failure, of rollback and of concurrency is the release, and an umbrella chart deliberately makes that unit as large as your system.

  • The upgrade reported success but the document-indexing job's pods are crash-looping. How is that possible?
    Because Helm 4 defaults to the `hookOnly` wait strategy, an upgrade that applies cleanly returns success without watching workloads become ready. The API server accepted the Deployment; the pods failed afterwards. Pass `--wait` (the `watcher` strategy) or `--wait=legacy` when you want readiness to decide the outcome, and pair it with a timeout you are willing to hold the release lock for.
  • Is `--rollback-on-failure` the right default for an umbrella release?
    It is a real tradeoff. It guarantees a known state instead of a half-applied one, but it reverts services that deployed correctly and it holds the release lock for the whole watch-then-revert cycle. On a large umbrella, teams often prefer leaving the failure visible and fixing forward. On a small, single-owner release it is usually worth switching on.
  • Does `helm rollback` delete the failed revision from the history?
    No. Rollback is a forward operation: it creates a new revision whose stored manifest is the one from the revision you named, so the failed revision remains in `helm history` as a record. That is useful for the post-mortem, and it is also why an umbrella with many contributors burns through the default `--history-max` of 10 quickly.

saying these in an interview costs you the question

  • Believes a failed upgrade automatically undoes what it applied
  • Thinks Helm waits for workloads to be ready by default
  • Expects helm rollback to revert only the broken subchart
  • Says each subchart keeps its own release status
  • Claims rollback deletes the failed revision from history
  • Treats an umbrella upgrade as an atomic transaction

context

open as a page

What is an umbrella Helm chart, and how does it differ from one release per service?

level: juniorimportance: should knowfreq 55%

basics

~20 s

An umbrella chart is a parent chart whose content is mostly its dependencies on other charts. Installing it produces ONE Helm release with one revision history, so everything in it upgrades and rolls back together. Per-service releases give each its own history.

open as a page

An umbrella Helm chart bundles 14 services plus a third-party ingress-controller chart, and its 7-minute upgrade blocks every team. Which signals say split it?

level: seniorimportance: should knowfreq 44%

basics

~10 s

Split when deploy cadences diverge, when one team's failure blocks others, when the shared history no longer answers "what changed for my service", and when a bundled third-party chart upgrades on someone else's schedule.

open as a page

How would you set Helm chart granularity policy across a platform of many services and teams?

level: principalimportance: should knowfreq 36%

basics

~20 s

Default to one Helm release per independently deployable unit: the release is the unit of rollback, failure and concurrency. Allow umbrellas only where a set is genuinely installed as one product, and build the composition layer the split needs.

open as a page