A `helm rollback` reports success on a 41-service platform chart, yet the incident continues. What does a rollback not restore?
answer
- It restores objects, not consequences
- Ask what happened outside the API server
- Some resources are deliberately never reverted
- Reverting can itself be the damage
- Applied is not the same as running
basics
~20 sA rollback re-applies a stored manifest, so it restores only what that manifest describes. Data a workload already wrote, volume contents, CRDs installed from crds/ and anything outside the cluster stay as they are — and rendered credentials are reverted, which can be worse.
solid answer
~50 s`helm rollback` restores **desired object state**, nothing else. Three classes of damage survive it. First, anything the manifest never described: rows a job wrote, files on a volume, an entry in an external system, a queue already drained. Second, objects Helm declines to revert — CustomResourceDefinitions installed from a chart's `crds/` directory are never upgraded and never deleted, so a schema change persists, and resources annotated `helm.sh/resource-policy: keep` are left alone. Third, and most dangerous, things the rollback *does* revert that you did not want reverted: a chart that renders a Secret from a values key rewrites that Secret with the old value, silently un-rotating a credential. In Helm 4, note also that the rollback did not wait for the restored workloads unless `--wait` was passed — the default strategy waits only for hooks — so a `deployed` status means the manifest was accepted, not that the old pods are serving.
code
bash · 6 lines# exactly which objects the rollback will add, remove or change
diff <(helm get manifest indexer-platform -n search --revision 47) \
<(helm get manifest indexer-platform -n search --revision 46)
# roll back and actually wait for the restored workloads (Helm 4)
helm rollback indexer-platform 46 -n search --wait=watcher --timeout 10mgo deeper
Know the boundary: a rollback restores the Kubernetes objects the release describes, and nothing about data, files or external systems. If asked, say what a manifest contains and reason from there.
Be able to enumerate the carve-outs — CRDs installed from crds/, resources annotated to be kept, objects that exist only in the newer revision and are therefore deleted — and explain what the release status does and does not assert.
Show incident judgement: diff the two revisions' stored manifests first, name the side effects that live outside them, and call out the reverted-credential trap before it bites. Verify workload health yourself rather than trusting the command's exit.
Own the design consequence: if rollback is your stated recovery path, the platform has to keep irreversible work out of the release and secrets out of rendered output. Decide where roll-forward is the honest answer and make that explicit in the runbook.
## The one-line model A Helm rollback re-applies a stored manifest. Everything it can restore is inside that YAML; everything else is out of scope, whether or not the incident depends on it. On a large release — a 41-service platform chart, say — that gap is where the second half of an outage lives, because the rollback genuinely succeeded and `helm history` genuinely reads `deployed`. ## Class 1: state the manifest never described A manifest declares objects, not their contents. A document-indexing job that ran under the new revision and rewrote index documents in a shared store has changed data that no rollback touches; the old code now reads records written by the new code. The same holds for files on a PersistentVolume, an external configuration service the app wrote to, a message queue whose offsets moved, a third-party system that was told about the new version. If a revision had side effects beyond the API server, rolling the release back leaves every one of them in place. This is the reason "we can always roll back" is a weaker guarantee than it sounds, and it is why irreversible steps deserve to be sequenced deliberately rather than trusted to a rollback. ## Class 2: objects Helm will not revert Two carve-outs matter here. **CRDs from `crds/`.** Custom resource definitions placed in a chart's `crds/` directory are installed once, before anything else, and are then never templated, never upgraded and never deleted by Helm. A rollback does not put an earlier schema back. If the newer revision widened a CRD and your custom resources were written in the newer shape, rolling back the controllers leaves them facing objects they may not understand. Nothing in the rollback output hints at this. **Kept resources.** Anything annotated `helm.sh/resource-policy: keep` is deliberately excluded from Helm's deletion path, so a rollback that would otherwise remove it leaves it in place — usually correct for stateful things, occasionally surprising. Also note the direction of change on ordinary resources: a rollback deletes objects that exist only in the newer revision and recreates ones the newer revision had removed. Deleting an object is not always cheap. If the newer revision introduced a claim or another stateful object, the rollback removes it — and whatever lived behind it goes with it or is orphaned, depending on how it was declared and what reclaim policy applies. ## Class 3: reverted things you wanted kept This is the class people never anticipate. A chart that renders a Secret from a values key stores that rendered Secret in every revision's manifest. If someone rotated the credential by upgrading with a new value, a rollback to a revision from before the rotation rewrites the Secret back to the **old** credential. Helm did exactly what it was asked; the effect is a silent un-rotation, and consumers holding the new credential start failing minutes later for reasons that look unrelated to the rollback. Worse, pods that took the value as an environment variable keep the value they started with, so half the fleet has one credential and half the other. The same shape applies to anything else that changed forward between the two revisions but is rendered from values: a replica count someone raised during the incident, a feature flag, an allow-list. A rollback is a *whole-release* operation; it cannot revert one service of forty-one and leave the rest alone. ## What "success" actually asserted In Helm 4 the wait strategy defaults to waiting for hooks only, so unless `--wait` was passed the command returned once the API server accepted the manifest. `deployed` therefore means "written", not "running": the restored Deployments may still be pulling images, and the old version may fail to start for the same environmental reason the new one did. Helm 3's separate opt-in `--wait` behaved comparably in that omitting it did not wait for workloads. Either way, the health check after a rollback is on you. ## Working the incident Diff the two revisions' stored manifests before and after rolling back — that tells you precisely which objects the rollback added, removed and mutated, and it is the shortest route to spotting a Secret or a claim in the change set. Then enumerate, deliberately, what the newer revision did that was *not* in that diff: data written, external calls made, CRDs installed. Those are the items that need their own remediation, and they are the ones that keep an incident open after the release ledger says it is closed.
- How would you stop a rollback from un-rotating a credential rendered from values?Take the credential out of the chart's rendered output. Have the chart consume an existing Secret by name — the value is then owned outside the release and no revision carries it — or have an in-cluster operator populate it from an external store. Rotation then lives on its own timeline and is unaffected by which revision the release is on. If the chart must render it, at minimum make rotation a values change that is applied forward again immediately after any rollback.
- The newer revision added a CustomResourceDefinition from the chart's crds/ directory. What does rolling back do to it?Nothing. Resources in `crds/` are installed once, before the rest of the release, and Helm never templates, upgrades or deletes them afterwards — a rollback included. The definition and any custom resources written against it survive, so a controller restored to its older version may face objects in a shape it does not expect. Reverting a CRD is a manual, deliberate operation, and one worth rehearsing before you need it.
- helm history shows the rollback revision as deployed, but users still report errors. What do you check first?Whether the restored workloads are actually running: `deployed` records that the manifest was accepted by the API server, and unless a wait strategy was requested the command did not watch the rollout. Check pod readiness and events for the restored version, and confirm the running image digests match what that revision expected — a mutable tag can leave you on the newer build despite a perfect rollback.
saying these in an interview costs you the question
- Treating rollback as a full undo of everything
- Assuming migrated or written data comes back
- Expecting CRDs from crds/ to be reverted
- Not realising a rollback rewrites rendered Secrets
- Reading deployed status as workloads healthy
- Forgetting that a rollback deletes newer-only objects