How do you recover a Helm release stuck in pending-upgrade after the CI job was killed mid-upgrade?
answer
- Least invasive rung that works
- First prove nothing is still running
- Return to the last terminal revision
- Deleting records is surgery, not routine
- The record and the cluster now disagree
basics
~20 sConfirm nothing is still running, then take the least invasive rung that works: roll back to the last revision that reached deployed. Only if that is impossible do you delete the stuck revision's release Secret by hand.
solid answer
~50 sWork a ladder. **First, prove it is abandoned** — check when the pending revision was written and that no job is still running against that release; clearing a record under a live client gives you two writers. **Second, `helm rollback <release> <last-deployed-revision> -n <ns>`.** Rollback appends a new revision built from the stored manifest of the one you name, and when it finishes that revision is `deployed`, so the pending state is no longer the top of the history and the next upgrade proceeds. **Third, if there is no deployed revision to return to** — a release stuck in `pending-install` — `helm uninstall` it and install again. **Last, delete the stuck revision's release Secret** so Helm sees the previous revision as newest. Then reconcile: run the real upgrade again and diff, because whatever the killed run pushed is live and unrecorded.
code
bash · 11 lines# 1. Which revision is stuck, and which one was last good?
helm history billing-cron -n billing
# 2. Least invasive: append a revision built from the last deployed one
helm rollback billing-cron 30 -n billing
# 3. Last resort: drop only the stuck revision's record
kubectl delete secret sh.helm.release.v1.billing-cron.v31 -n billing
# 4. Reconcile: the cancelled run may have applied part of the release
helm upgrade billing-cron ./billing-cron -n billinggo deeper
Know that the fix is a deliberate recovery, not a retry: look at the history, roll back to the last revision marked deployed, and ask for help before deleting anything in the cluster by hand.
Explain why rollback clears the block — it appends a revision that ends in a terminal status — and why pending-install has no rollback target and needs an uninstall instead.
Show the ladder and the judgement between its rungs: rule out a live client, prefer Helm's own write path, treat record deletion as surgery on one revision only, and finish by reconciling the drift the abandoned run left behind.
Own the systemic answer: pipeline budgets that let Helm write its verdict, per-release serialisation, and a policy for who may touch release records, so recovery is not improvised with cluster-admin credentials during an incident.
Take a concrete case. A chart for a subscription-billing cron, which bundles an optional database subchart, is upgraded by a nightly pipeline. The pipeline was cancelled seventeen minutes into the run. `helm history billing-cron` now shows revision 30 `deployed` and revision 31 `pending-upgrade`, and every subsequent `helm upgrade` refuses with `another operation (install/upgrade/rollback) is in progress`. The release Secret for revision 31 is 1.1 MiB, because the bundled database subchart renders a lot of manifest. ## Rung 0 — prove that nothing is running The pending status is Helm's belief that a write is in flight. Before you falsify that belief, confirm it is false: look at the pending revision's age, and at whether any pipeline run or engineer still has a `helm upgrade` open against this release in this namespace. If the answer is yes, wait. Clearing a record while a client is still working produces two writers with different ideas of the current revision, and the loser can overwrite the winner's result. This is the rung people skip, and it is the one that turns a five-minute fix into an incident. ## Rung 1 — roll back to the last deployed revision ``` helm history billing-cron -n billing helm rollback billing-cron 30 -n billing ``` Rollback does not rewind history — it appends. It creates revision 32, populated from the manifest stored with revision 30, and drives it through the usual pending-then-terminal cycle. When it lands as `deployed`, the top of the history is a terminal status again and the guard no longer fires. This is the least invasive rung because Helm does the work through its normal write path: the release record stays coherent, the revision numbering stays intact, and the cluster is pushed back toward a state Helm actually has on file. Name the revision explicitly rather than relying on the default. On a stuck release the "previous" revision is not always the one you want, and being explicit is also self-documenting in an incident channel. ## Rung 2 — no deployed revision to return to If the stuck status is `pending-install`, revision 1 is the only revision and there is nothing behind it. Rollback has no target. The recovery is `helm uninstall billing-cron -n billing` followed by a fresh install: uninstall removes whatever the abandoned install managed to create, and you start clean. This is why `pending-install` and `pending-upgrade` are genuinely different problems despite looking alike in a listing. ## Rung 3 — remove the stuck revision record If rollback is not available or itself refuses, the last lever is to delete the release record for the stuck revision, so that Helm's read of "newest revision" lands on the previous, terminal one: ``` kubectl get secret -n billing -l owner=helm,name=billing-cron kubectl delete secret sh.helm.release.v1.billing-cron.v31 -n billing ``` Two cautions. First, this is surgery on Helm's own bookkeeping — the record is not a text file you can casually hand-edit, and flipping a status in place means decoding and re-encoding the stored payload, which is far more error-prone than deleting one revision and letting the previous one stand. Second, **delete only the stuck revision**. If you delete every revision Secret for the release, Helm forgets the release entirely: `helm list` no longer shows it, and a subsequent install collides with the objects still running, which already carry another release's ownership metadata and will be rejected until you explicitly adopt them. ## Rung 4 — reconcile, because the record and the cluster disagree Every rung above fixes the *record*. None of them tells you what the cancelled run had already pushed. Revision 31 may have applied part of the release — the database subchart's objects but not the cron's, or an updated ConfigMap with the old workload still referencing it — and Helm has no note of any of it. So the last step is always to run the real upgrade again and confirm the outcome: a rendered diff before applying, then the workload's own health after. Rolling back to revision 30 puts revision 30's manifest back, which reverts most of that drift, but resources the newer chart added and the older one never mentioned are not in revision 30's manifest and so are not removed by the rollback either. ## Prevention These are manufactured, not natural. Bound the pipeline's own timeout so it is longer than the Helm timeout you pass, so Helm always gets to write its verdict rather than being shot first. Serialise deploys per release so two runs cannot race. And prefer letting Helm fail loudly — a release in `failed` is an ordinary state you can upgrade or roll back straight out of, whereas a release in `pending-upgrade` needs a human with cluster credentials.
- Why delete only the stuck revision's release Secret rather than all of the release's Secrets?Deleting one revision just moves Helm's idea of "newest" back to the previous, terminal revision, and the release survives with its history. Deleting all of them erases the release: Helm no longer lists it, and a fresh install runs into the objects still running, which carry ownership metadata pointing at a release Helm no longer knows about, so the install is rejected until you deliberately adopt them.
- After a successful rollback, why can the cluster still differ from what revision 30 described?Because the rollback replays revision 30's stored manifest, and anything the abandoned run created that revision 30 never mentioned is simply absent from that manifest — so it is not reverted and not deleted. Objects the newer chart added, or side effects such as a job that already ran, survive. That is why the last step is a real upgrade plus a diff, not a declaration of victory after the rollback returns.
- How would you keep your pipeline from manufacturing stuck releases in the first place?Make the pipeline's own cancellation and wall-clock budget longer than the timeout you pass to Helm, so Helm always reaches its final write and lands in failed rather than pending. Serialise deploys per release so two runs cannot race, and avoid killing deploy jobs by reflex. A failed release is self-service to fix; a pending one needs someone with cluster credentials.
saying these in an interview costs you the question
- Deleting release Secrets as the first move
- Clearing the record without checking for a live run
- Hand-editing the stored release payload in place
- Deleting every revision Secret for the release
- Assuming rollback also reverts what the killed run created
- Trying to force the upgrade through with more flags