skip to content

In a Helm estate holding persistent data and CRDs, how do you make uninstall predictable?

level: principalimportance: should knowfreq 38%

answer

  1. Two phases, two owners, two reversibilities
  2. The dangerous half needs a waiting period
  3. Decide survivors when authoring, not during incidents
  4. Blast radius crossing tenants stays manual
  5. A zero exit code proves nothing

basics

~20 s

Decide at chart-authoring time which resource classes may survive an uninstall, split teardown into deleting the release and a separate owned reclaim step, keep cluster-scoped deletion manual, and verify the namespace afterwards instead of trusting the exit code.

solid answer

~50 s

Treat decommissioning as a two-phase operation with different owners. Phase one is `helm uninstall`, which is safe to automate because its blast radius is exactly the stored manifest. Phase two is reclaiming what survives — orphaned claims, `keep`-annotated objects, CRDs from `crds/`, the namespace — which is a deliberate, logged act with a named owner and a waiting period, never a tail appended to the pipeline. Make the survivor set a chart-authoring standard rather than a per-release judgement: only unrecoverable state may carry `helm.sh/resource-policy: keep`, and that list is reviewed. Differentiate by environment — ephemeral preview namespaces should keep nothing, production keeps everything data-bearing. Cluster-scoped deletion, above all CRDs, stays manual because its blast radius crosses tenants. Finally, verify: assert the namespace is empty against the manifest you captured before teardown rather than trusting a green exit code.

code

bash · 8 lines
bash
# Phase one: bounded, automatable
helm get manifest platform-core -n team-alpha > teardown/delete-set.yaml
helm uninstall platform-core -n team-alpha

# Reconcile: anything here that is not in the capture is a survivor
kubectl get pvc,secret,configmap -n team-alpha -o name

# Phase two runs later, by a named owner, and is not in this script

go deeper

for a junior

Take away one habit: after a teardown, look at the namespace rather than trusting the command's success message, and ask someone before deleting anything that holds data.

for a middle

Be able to explain why teardown splits into a bounded phase and an irreversible one, and which resource classes fall on each side. Knowing what survives an uninstall is the prerequisite for having any policy at all.

for a senior

Show that you design the teardown as an operation with evidence: a captured delete set, a reconciled survivor inventory, and a reclaim step with an owner. Be ready to argue why the cluster-scoped part stays manual.

for a principal

Own the standard rather than the script. Decide which resource classes may carry keep, how policy differs by environment, who reclaims orphans on what clock, and how the ongoing cost of survivors gets attributed to a team.

### Why this is a design question and not a command At one chart, uninstall leftovers are trivia. At an 18-chart umbrella owned by an infrastructure team and installed into dozens of namespaces, they are a policy problem: one `helm uninstall` spans 18 subcharts' worth of residue, and the person who ran it knows the residue of maybe two of them. The failure is not that Helm behaves surprisingly — it behaves exactly as documented — but that nobody owns what documented behaviour leaves behind. ### Phase the teardown, and give the phases different owners Split decommissioning in two. **Phase one — remove the release.** `helm uninstall` deletes the objects in the stored manifest and nothing else. Its blast radius is bounded and inspectable in advance, so it is legitimately automatable: a pipeline may run it. **Phase two — reclaim the survivors.** Orphaned claims, `keep`-annotated objects, CRDs, and the namespace itself. This phase is where irreversible loss lives, and it should be a separate, logged action with a named owner, run after a deliberate delay — long enough that the team that asked for the teardown can say "actually, put it back". The delay is the whole point: phase one is reversible in shape, phase two is not. A good tell in an interview is a candidate who notices that combining the phases is what makes teardown scary, and that separating them is what makes phase one boring enough to automate. ### Decide the survivor set at authoring time `helm.sh/resource-policy: keep` is written by a chart author months before any incident. Make that a reviewed standard, not a habit: - **May carry keep:** objects whose loss is unrecoverable — a chart-rendered claim holding data, a CRD rendered under `templates/`, a credential generated once and not reproducible. - **May not:** anything re-renderable. Deployments, Services, ConfigMaps, ServiceAccounts. Keeping those makes uninstall a lie: the release vanishes from `helm list` while the workload keeps serving traffic, and the next engineer has to reconstruct by hand what it consisted of. Every `keep` also has a second cost to price in: the survivor blocks a later install of the same chart, because Helm will not adopt an object whose ownership metadata points elsewhere. So each annotation buys data safety and sells redeploy smoothness. Charts should be able to justify the trade per object. ### Differentiate by environment One policy for all environments is the wrong answer. Ephemeral namespaces — preview environments, CI runs — should keep nothing at all, and their teardown should delete the namespace outright, because the cost of losing their state is zero and the cost of leaking storage across thousands of runs is not. Production is the opposite: data-bearing objects keep, and phase two runs on a human clock. If the same chart serves both, the difference belongs in values, not in two forks of the chart. ### Cluster-scoped deletion never goes in the pipeline CRDs are the sharpest case. Helm never deletes the ones installed from `crds/`, and that exemption exists because a CRD is cluster-scoped: removing it removes every custom resource of that kind everywhere, including objects belonging to teams that have nothing to do with this release. Automating that cleanup re-introduces exactly the blast radius Helm deliberately declined. The standard is: no automated deletion of cluster-scoped objects, ever; removal requires a check that no instances exist anywhere and a human decision recorded somewhere durable. ### Make emptiness provable A teardown that ends with a zero exit code has proved nothing about the namespace. Capture the delete set (`helm get manifest`) before phase one, then after it reconcile the namespace against that capture — what is present and was not in the capture is your survivor inventory, and it should match what the chart's `keep` annotations and storage design predicted. When it does not, that is the finding: either the chart changed or someone created objects outside the release. Make that reconciliation an artefact of the teardown, not an ad-hoc `kubectl get` someone runs from memory. ### Own the cost side Orphaned storage bills monthly and is invisible in every dashboard organised by release, because the releases are gone. Whatever the reclaim policy is, it needs a periodic sweep that finds claims in namespaces with no live release and asks the owning team to confirm — with an escalating clock rather than a silent one. Retained release records deserve the same treatment: `--keep-history` is an audit decision with an ongoing price, namely an occupied release name and a stale record, and something should eventually prune the ones nobody consults. ### The shape of a strong answer Phases with different owners and different reversibility; the survivor set decided in the chart and reviewed; environment-differentiated policy; cluster-scoped deletion kept manual on blast-radius grounds; and verification that produces evidence. A candidate who reaches for a cleanup script that deletes leftovers automatically has optimised the wrong variable — the difficulty here was never running the deletes, it was deciding which ones are safe.

  • A platform team proposes a cleanup job that deletes every leftover after an uninstall, CRDs included. What is your objection?
    The hard part was never issuing the deletes, it was deciding which are safe. A CRD is cluster-scoped, so deleting it removes every custom resource of that kind across the cluster, including other tenants' objects — that is precisely the blast radius Helm declined to take on. Automate the bounded phase; keep the irreversible, cross-tenant phase human and logged.
  • How does your teardown policy differ between a preview namespace and production?
    Preview namespaces keep nothing and are torn down by deleting the namespace outright: their state is worthless and leaked storage across thousands of runs is not. Production keeps every data-bearing object and reclaims on a human clock with a waiting period. If one chart serves both, that difference lives in values, not in a forked chart.
  • How do you stop orphaned storage from accumulating invisibly across the estate?
    A periodic sweep that finds claims in namespaces with no live release, attributes them to a team, and escalates on a clock rather than notifying silently. Dashboards organised by release cannot see them by definition — the release is gone — so the sweep has to key off the namespace and the absence of a release, and it should propose deletion rather than perform it.

Moving out of a flat: handing back the keys is one action, and clearing the storage locker in the basement is another. Doing them in one motion is how someone else's boxes end up in the skip.

saying these in an interview costs you the question

  • Automates cluster-scoped deletion at teardown
  • Treats a zero exit code as proof of an empty namespace
  • Applies one teardown policy to preview and production alike
  • Annotates everything keep to be safe
  • Ignores the ongoing cost of orphaned volumes
  • Puts irreversible reclaim in the same step as uninstall

context