skip to content

A helm upgrade fails saying the release Secret is too long. What is Helm storing there, and how do you fix it?

level: seniorimportance: should knowfreq 33%

answer

  1. The render succeeded; only the write failed
  2. One object holds the whole chart, every revision
  3. Compressed payload against a per-object ceiling
  4. Non-template files count against you
  5. Trimming history is the wrong lever here

basics

~20 s

The record embeds the whole chart plus supplied values and rendered manifest, gzipped into one Secret the API server rejects past roughly 1 MiB. Shrink what the chart carries, split the release, or change backend; trimming history will not help.

solid answer

~50 s

Helm stores each revision as a single object holding the entire chart — every file in it, subcharts included — plus the values supplied and the manifest rendered, JSON-encoded and gzipped. When that compressed payload passes the roughly 1 MiB limit the API server enforces on a Secret's data, the write is rejected and the upgrade fails, typically with a message naming `sh.helm.release.v1.<name>.v<rev>` and saying its data is too long. The tell is that `helm template` and `--dry-run=server` still succeed: rendering was never the problem, only the record write. Fixes attack what goes into the record: exclude bulk files with `.helmignore`, stop vendoring large static assets into the chart, split an umbrella chart into separate releases per component, or move to the `sql` storage backend, which is not bound by an object-size limit. `--history-max` does not help — that limit is per record.

code

bash · 3 lines
bash
helm template monitoring ./monitoring-stack -n observability > /dev/null && echo 'render ok'
helm package ./monitoring-stack && ls -l monitoring-stack-*.tgz
helm upgrade monitoring ./monitoring-stack -n observability

go deeper

for a junior

Take away the core fact: the whole chart, not just the rendered output, is stored with every revision, so a chart carrying large files eventually cannot be written at all. You are not expected to have diagnosed this in production.

for a middle

Explain what goes into the record and that it is gzipped into one object with a hard size limit. Be able to say why the same chart renders fine, and name .helmignore as the first thing to reach for.

for a senior

Show the diagnosis path — render works, dry run works, the write fails — and rank the fixes: strip what the chart carries, move bulk data out, split into separate releases, change backend last. Separate this cleanly from history retention.

for a principal

Own the design question: how large a chart is allowed to get before it becomes several releases, where that is checked so it is caught in review rather than during an incident, and what you are willing to pay in operational coupling if release state moves out of the cluster.

### The failure A monitoring-stack chart with three subcharts installs fine for months, someone adds a directory of dashboards and recording rules to it, and the next `helm upgrade` dies before anything reaches the cluster. The message names the release record — `sh.helm.release.v1.monitoring.v8` — and complains that its data is too long. No workload changed, no template is broken, and running `helm template` on the same chart produces perfectly good YAML. That asymmetry is the diagnosis. `helm template` renders without writing a record. `--dry-run=server` sends the manifest for validation without writing a record. Anything that skips the record write succeeds; anything that performs it fails. So the fault is not in rendering, in the manifest, or in the cluster's admission path — it is that the record Helm wants to persist will not fit in a single object. ### Why the record is that big The release record is not a diff and not a pointer. Every revision embeds: - the chart itself, file by file — `Chart.yaml`, `values.yaml`, every template, every file under any other directory in the chart, and each resolved subchart in full; - the values the caller supplied; - the manifest Helm rendered for this revision; - the hook manifests, notes and status. That whole structure is marshalled to JSON, gzipped, base64-encoded, and stored under one key of one Secret. The API server enforces roughly 1 MiB on a Secret's data, so that is the wall. Two things follow. First, the budget is a *compressed* budget: rendered Kubernetes YAML is repetitive and compresses very well, so a chart producing several megabytes of manifest can still fit comfortably. Second, and less obviously, non-template files count. A directory of dashboards, a vendored `.tgz`, sample data, screenshots in a docs folder — none of it renders into anything, and all of it is copied into every revision record, where it compresses far worse than the YAML does. Charts usually cross this line by carrying things, not by rendering things. Umbrella charts cross it faster because subcharts are stored in full too, so the record grows with the sum of everything the parent pulls in. ### Fixes, in the order worth trying **Stop shipping what does not need to ship.** `.helmignore` excludes paths when the chart is loaded and packaged, so excluded files never enter the record. Docs, test fixtures, CI configuration, raw dashboards, stray archives — this alone often reclaims most of the space, and it is the change with no operational consequence. **Move bulk data out of the chart.** Large static payloads embedded so a template can inline them into a ConfigMap are the classic cause. Have the workload fetch them at runtime, or manage them as their own objects with their own lifecycle, and the chart stops carrying them once per revision forever. **Split the release.** An umbrella that installs several substantial components as one release keeps one record for all of them. Installing each component as its own release gives each its own record, its own revision chain and its own rollback — usually a better operational shape than a single record straining against a hard ceiling, at the cost of coordinating versions across releases yourself. **Change the storage backend.** `HELM_DRIVER=sql` persists records to an external PostgreSQL database, where the per-object ceiling does not apply. It genuinely solves the problem, and it is the last resort: release state leaves the cluster it describes, every install and rollback gains a hard dependency on a reachable database, and a cluster restore no longer restores release history. ### The trap: trimming history The reflex is `--history-max`, and it is the wrong lever. Retention controls how *many* records exist; this failure is one record that is individually too large. Trim to three, trim to one, and revision 8 still will not fit. The two levers look adjacent because both are described as "release storage", and separating them cleanly is most of what a good answer to this question demonstrates. Retention is for estate-wide datastore pressure; the size ceiling is for a specific chart that carries too much. ### Preventing the recurrence The practical guard is to know the number before the cluster tells you. Package the chart and look at what went in, and treat a chart tarball growing by a large fraction in one merge as a review question rather than a detail. A `.helmignore` written once at chart-creation time, covering docs, fixtures and archives, prevents most of these outright. And when a chart is genuinely large because the system is large, prefer splitting it into releases early — the alternative is discovering the ceiling during an incident, when the release you cannot upgrade is also the release you cannot roll back, because a rollback writes a new record too.

  • Why does helm template succeed on the same chart that fails to upgrade?
    Because `helm template` renders locally and writes no release record. The failure is in persisting the record, not in producing the manifest — the same is true of `--dry-run=server`, which validates against the API server but stores nothing. That contrast is the fastest way to prove the problem is the record write rather than a broken template or a rejecting admission path.
  • Would setting --history-max 3 resolve it?
    No. Retention bounds how many revision records exist; this failure is a single record that individually exceeds the per-object size limit. With a cap of three the same revision is still too large to write. Retention is the lever for total datastore pressure across many releases; the ceiling is fixed by shrinking what the chart carries, splitting the release, or changing the storage backend.
  • A colleague suggests deleting files from the chart directory just before the upgrade. Why is that a bad fix?
    Because whatever is present when the chart is loaded ends up in the record, so it works exactly once and only for whoever remembered. The durable version of the same idea is `.helmignore`, which excludes paths at load and package time for every caller and every pipeline. If the files are genuinely needed at runtime, deleting them locally also means the next person's render differs from yours, which is a worse problem than the one being fixed.

saying these in an interview costs you the question

  • Blames the rendered manifest size rather than the stored record
  • Reaches for --history-max to fix a single oversized record
  • Thinks the record stores only a diff from the previous revision
  • Forgets that non-template files in the chart are stored too
  • Assumes subcharts are referenced rather than embedded in the record
  • Says compression makes the limit effectively unreachable

context