skip to content

Across 400+ Helm releases in one cluster, how do you decide storage driver, retention and recovery for release state?

level: principalimportance: should knowfreq 38%

answer

  1. Ask what the real source of truth is
  2. Keep state beside the workloads unless forced
  3. Retention comes from a detection window
  4. One enforcement point, including break-glass
  5. Lost records orphan, they do not delete

basics

~20 s

Treat release records as derived state: chart and values in version control are the truth. Keep the default in-cluster backend unless size or scale forces otherwise, set retention centrally from your detection window, and govern who can read stored values.

solid answer

~50 s

Start from what the records are: one object per revision per release, holding the chart, the supplied values and the rendered manifest, living in each release's namespace. That gives three decisions. **Backend** — keep the default so state travels with the namespace, shares its RBAC and its backup, and adds no dependency to the rollback path; move to `sql` only when a record cannot fit or the accumulated records are a measured load, accepting that release state then leaves the cluster it describes. **Retention** — set `--history-max` from how long a bad deploy takes you to detect, and enforce it in the one wrapper every path uses, including break-glass, because the flag is per invocation. **Recovery** — decide whether losing records is recoverable: with chart and values in Git a lost release is re-adoptable; otherwise the record is an unbacked-up system of record. Govern reads too: supplied values are stored verbatim.

code

bash · 2 lines
bash
kubectl get secret -A --field-selector type=helm.sh/release.v1 \
  -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name --no-headers | wc -l

go deeper

for a junior

Focus on the ground truth behind this discussion: release state is stored per revision in the cluster and Helm depends on it for history and rollback. You are not expected to set fleet policy yet.

for a middle

Be able to explain the levers being traded — which backend, how many revisions retained, what a lost record costs — and why the chart and values living in version control changes the answer to all three.

for a senior

Show that you would measure before moving: count records and sizes, derive retention from a detection window, enforce it in one path, and rehearse recovering a release whose records are gone rather than improvising it during an incident.

for a principal

Own the framing that release records are derived state, and defend it: what makes them re-creatable, who may read namespaces where values carry credentials, why the storage backend is a fleet decision rather than a per-chart fix, and what you accept losing in a cluster restore.

### Frame it: derived state or system of record? Every other decision follows from this one. Helm's release records hold the chart, the values and the rendered manifest for each revision. If the chart version and the values that produced a release are reproducible from version control, the records are a *cache* — convenient, load-bearing for rollback, and re-creatable in principle. If people deploy with ad-hoc `--set` flags from laptops, the records are the only place the deployed configuration exists, and you are running an unbacked-up system of record inside the cluster it manages. Most estates are somewhere in between and have never said which. Saying it out loud changes the recovery plan, the retention number and how much you care about backups, so it is the first thing to settle. ### Backend The default keeps a Secret per revision in the release's own namespace. The properties that gives you are worth naming, because they are why the default usually wins: state is namespaced, so tenant boundaries apply to it automatically; it inherits the namespace's RBAC and whatever the cluster does to protect Secrets; it is included in any backup that covers the cluster; and there is no second system that can be unavailable at the moment you need to roll back. Moving to `sql` is justified by exactly two constraints — a chart whose record does not fit in one object, and record volume that is a real load on the cluster datastore. At 400 releases with default retention you are talking about a few thousand objects; sizes vary hugely, so measure before assuming. What you pay is real: release state now lives outside the cluster it describes, so a cluster restore no longer restores release history; every install, upgrade and rollback depends on a reachable database; the connection string is a credential every deploying identity holds; and your tenant isolation is no longer namespace-shaped. I would exhaust cheaper levers first — retention, and shrinking what charts carry — and adopt `sql` only for the specific estate that genuinely needs it, not fleet-wide because one chart is large. The corollary is discipline about the variable: because the backend is read from the environment per invocation, a non-default choice must live in a shared wrapper or base image. A split where CI uses one backend and engineers use another produces two groups who disagree about which releases exist, and an `upgrade --install` that tries a fresh install over live objects. ### Retention Retention is a rollback horizon expressed as a count, and it should be derived from a time. Ask how long a bad change typically survives before someone notices — the answer for a telemetry ingest gateway that deploys twelve times a day is very different from a batch service that deploys monthly. Multiply detection time by deploy rate, add margin, and that is your number; the default 10 is a starting point, not a decision. Enforce it in one place. The flag is per invocation, so a single manual upgrade without it silently trims the release back to 10 — at exactly the moment somebody is deploying by hand during an incident and might want a deep rollback. Put it in the deploy template, the wrapper script and the break-glass runbook, not in each pipeline's environment block. And resist using retention as a storage-pressure fix for the wrong problem: it bounds the number of records, never the size of one. A record too large to write is a chart problem or a release-splitting problem. ### Recovery Write down what happens when a namespace's records are lost — deleted by a cleanup job, dropped by a partial restore, or stranded by a backend switch. The workloads keep running; Helm simply forgets them. Recovery means re-establishing a release over live objects that still carry the previous release's ownership metadata, which is a deliberate adoption step rather than a re-install, and it goes far more smoothly when the chart version and values are recoverable from Git than when they were typed into a terminal. So the concrete asks are: chart and values in version control for every production release; deploys through a path that records what it passed; backups that include the release records or a documented acceptance that they are re-creatable; and a rehearsed procedure for adopting live objects back into a new release record. Rehearsed matters — the first time anyone tries this should not be during the incident that caused it. ### Governance of reads One more thing a lead owns: whatever values were supplied are stored in the record. A password passed to a chart is readable by anyone who can read Secrets in that namespace, whether or not the chart rendered it anywhere. In a multi-tenant cluster that makes namespace read access to release records a real privilege boundary, and it argues for charts that reference externally-managed credentials rather than accepting them as values — a chart-design decision driven entirely by how release state is stored. ### What I would land on Default backend; retention set from detection windows and enforced centrally; chart and values in Git so records are genuinely derived state; explicit read governance on namespaces where credentials pass through values; and `sql` reserved as a targeted answer to a measured problem, never as a fleet-wide default.

  • Would you back up the release records, or accept re-creating them?
    It depends on whether they are derived. If every production release's chart version and values come from version control, I accept re-creation and invest instead in a rehearsed adoption procedure. If deploys happen with ad-hoc flags, the records are the only copy of the deployed configuration and must be backed up — but I would treat that as a defect to remove rather than a property to protect, because a backup does not fix the underlying loss of reproducibility.
  • A team wants the sql backend fleet-wide after one chart hit the object-size ceiling. What do you say?
    That the ceiling is one chart's problem and the backend is everyone's. I would fix the chart — strip what it carries, or split it into separate releases — and keep the default elsewhere. Fleet-wide `sql` puts a database in the rollback path for 400 releases, removes release history from cluster backups, and hands every deploying identity a connection string, to solve a problem that exists in one place.
  • How does release storage shape how you want charts to take credentials?
    Values supplied at install are stored verbatim in every revision record, so a chart that accepts a password as a value writes that password into a Secret per revision, readable by anyone with namespace read access, and retained for as long as that revision survives. I would prefer charts that reference an externally-managed credential by name and let the platform deliver it, so release records stay free of material nobody intended to persist.

saying these in an interview costs you the question

  • Adopts an external storage backend fleet-wide to fix one chart
  • Treats release history as an audit trail or a backup
  • Ignores that stored values are readable by namespace readers
  • Sets retention per pipeline instead of one enforcement point
  • Assumes losing release records deletes running workloads
  • Never states what the actual source of truth is

context