A helm upgrade ends with "Error: UPGRADE FAILED" — how do you classify what actually failed?
answer
- The banner is not the message
- Four classes, four cluster states
- Did anything reach the API server?
- Deadline exceeded means applied, not rejected
- helm status keeps the description
basics
~20 sRead the clause after UPGRADE FAILED. Helm names the class there: a failed hook, an API-server rejection such as an admission webhook denial or a field-manager conflict, or an expired wait reported as a context deadline. Each leaves the cluster differently.
solid answer
~50 sHelm prefixes every upgrade error with `Error: UPGRADE FAILED:`, and the useful part is the clause after it. `pre-upgrade hooks failed: ...` means Helm never reached the chart's ordinary manifests — hooks run first, so nothing else was applied. An `admission webhook ... denied the request` message, or a field-manager conflict under server-side apply, is the API server refusing one object mid-apply, so the release is left partially applied. `context deadline exceeded` means the opposite: every object *was* applied and Helm's wait ran out. There is also a silent fourth case — in Helm 4 an upgrade without `--wait` waits only for hooks, so a workload that never becomes ready fails nothing and the release reports `deployed`. Confirm afterwards with `helm status <release>`, whose description preserves the message, and `helm get hooks <release>` to see which hooks the chart declared.
code
bash · 8 lineshelm upgrade digest-builder ./charts/digest-builder -n mail --timeout 7m --wait
# Error: UPGRADE FAILED: pre-upgrade hooks failed: ...
# the same message, recoverable later, in the release description
helm status digest-builder -n mail
# which hooks the chart declared, in weight order
helm get hooks digest-builder -n mailgo deeper
Be ready to say that the words after Error: UPGRADE FAILED: are the diagnosis, and to name the command that shows a release's status. Do not offer re-running the upgrade as the first step.
An interviewer expects you to walk the classes: hook failure, API-server rejection, expired wait, and the case where nothing failed because Helm never waited — and to say what is running in the cluster in each.
Show that you reach for the release record rather than the scrollback: the description on the failed revision, the rendered hooks, and the elapsed time as evidence of which class you are in.
Own the question of what your platform captures automatically when an upgrade fails, so that engineers are classifying from an artefact rather than from a pipeline log that has already rotated.
`helm upgrade` fails in a small number of distinguishable ways, and Helm tells you which one in the clause that follows its `Error: UPGRADE FAILED:` prefix. Reading that clause *is* most of the triage, because each class leaves the cluster in a different state and demands a different next move. Re-running the command before reading it is the classic weak answer. | Class | What is running now | | --- | --- | | A hook failed (`pre-upgrade hooks failed: ...`) | The previous revision, untouched | | The API server refused an object | Some objects new, some old | | The wait expired (`context deadline exceeded`) | The new manifest, still converging | | Nothing failed, no error at all | The new manifest, workload never ready | ## A hook failed The message reads `pre-upgrade hooks failed: ...`. Helm applies manifests annotated `helm.sh/hook: pre-upgrade` before it applies the chart's ordinary resources, in `helm.sh/hook-weight` order, waiting for each to complete. When one fails, Helm stops there: none of the chart's ordinary manifests were applied, and the objects from the previously deployed revision are still running untouched. A new revision is recorded with status `failed`, but the workload did not change. The corollary matters more than the message: whatever the hook itself did to the outside world — a schema migration, a cache flush, a call to another service — is half-done, and Helm has no concept of undoing it. ## The API server refused an object This arrives in two shapes. The first is an admission denial — `admission webhook "..." denied the request: ...` — where a validating webhook, typically an admission policy engine a platform team installed, rejected a rendered object. The object was never created. The second shape exists because Helm 4 makes server-side apply the default write path: another field manager (a human's `kubectl apply`, an autoscaler, another controller) already owns a field your chart is now setting, and the apply comes back as a conflict naming that manager and the field path. The resolution there is `--force-conflicts`, which is a *different* flag from `--force` — in Helm 4 `--force` is a deprecated alias of `--force-replace`, which recreates resources, and the two are mutually exclusive. Conflating them is a red flag. Either way, Helm was part-way through applying when it stopped, so some resources may already carry the new spec while others do not. ## The wait expired The message is `context deadline exceeded` (Helm 3's polling waiter phrased the same situation as `timed out waiting for the condition`). Nothing was rejected and nothing failed to apply. Helm wrote every object successfully, waited for the workload to report ready, and `--timeout` ran out first. The cluster holds the *new* manifest; the release is marked `failed` regardless; Kubernetes may still be converging while you read the error. This is the class most often misdiagnosed as "Helm broke the deploy". ## Nothing failed and the application is still broken In Helm 4 `--wait` is strategy-valued — `watcher` (event-driven), `legacy` (the Helm 3 poller) or `hookOnly` — and *omitting* it means `hookOnly`: Helm waits for hooks and not for your Deployment. An upgrade whose pods never become ready therefore reports success, and `helm status` shows `deployed`. When the scenario is "the pipeline went green and the service is down", this is the answer, and no error message exists to read. ## Confirming it after the fact Two commands carry you from the terminal to the record. - `helm status <release>` prints the release's status and its description, and the description of a failed revision preserves the same failure text — which is how you recover the cause a week later when the CI job's log has rotated away. - `helm get hooks <release>` prints the rendered hook manifests for the release: which hooks the chart declared, their events, their weights and their delete policies. That tells you what *should* have run and in what order, which is exactly the question a `pre-upgrade hooks failed` message raises and does not answer. ## A worked example An email-digest builder is deployed from a service chart that consumes a shared library chart, one of twelve services doing so. A `helm upgrade` with `--timeout 7m --wait` comes back seven minutes later with `Error: UPGRADE FAILED: context deadline exceeded`. The elapsed time equalling the timeout almost exactly is the tell: this is class 3, not class 2. Nothing was rejected, the new manifest is live, and the question is now whether the digest builder's rollout is progressing slowly or parked — not whether Helm did anything wrong. Had the same run failed in under a second with `pre-upgrade hooks failed`, the correct conclusion would be the opposite: the digest builder is still running its old spec, and the failure is entirely inside the library chart's migration hook. The discipline to demonstrate is that you name the class before you name a fix, and that you can say, for each class, what is running in the cluster right now.
- The CI log has rotated away. Where do you still find why last week's upgrade failed?`helm status <release>` prints the release description, and Helm writes the failure text into the description of the revision it recorded. Since every revision is stored as a Secret named `sh.helm.release.v1.<name>.v<rev>` in the release namespace, that record outlives the terminal and the pipeline log. In Helm 4 the description always prints — the old `--show-desc` flag was removed.
- Why can a helm upgrade report success while the service is down?Because Helm only waits for what you asked it to wait for. In Helm 4, omitting `--wait` selects the `hookOnly` strategy, so Helm waits for hooks and returns as soon as the API server accepts the manifests. A Deployment whose new pods crash-loop or never pass readiness produces no Helm error at all, and the release status stays `deployed`.
- An upgrade stops with a conflict naming another field manager. Which flag clears it, and which one must you not reach for?`--force-conflicts` takes ownership of the disputed fields under server-side apply. Do not reach for `--force`: in Helm 4 that is a deprecated alias of `--force-replace`, which deletes and recreates resources rather than resolving field ownership, and it is mutually exclusive with `--force-conflicts`. The conflict itself usually means someone edited the object by hand or another controller owns that field legitimately.
Like reading a build log: the banner at the bottom only says it failed; the first line above it says which stage stopped and how much of the pipeline ever ran.
saying these in an interview costs you the question
- Assumes every failed upgrade rolled itself back
- Reads context deadline exceeded as a rejected resource
- Assumes a failed upgrade left the cluster untouched
- Re-runs the upgrade before reading the message
- Says --force resolves a server-side apply conflict
- Believes a deployed status proves the pods are ready