skip to content

A CI job's `helm upgrade` resolves a chart version older than the one published an hour ago — why, and how do you make it deterministic?

level: seniorimportance: should knowfreq 46%

answer

  1. Version resolution happens locally first
  2. Something on disk was written once
  3. Nothing refreshes it on a timer
  4. One command re-downloads the index
  5. An oci:// reference has no index

basics

~20 s

Helm resolves a repo/chart reference against the repository index cached on that machine, not against the server. Nothing refreshes that cache on a timer, so a runner keeps whatever it downloaded until helm repo update runs.

solid answer

~50 s

A reference like `billing-charts/billing-cron` is resolved locally. When a repository is added, Helm downloads its index into the repository cache; from then on `helm install`, `helm upgrade`, `helm search repo` and dependency resolution all read that cached copy. There is no TTL and no background refresh — only `helm repo update` (optionally for a single repository) re-downloads it. A CI runner with a warm cache volume, or a pre-baked image, therefore keeps resolving the chart versions that existed when that cache was written. Two symptoms follow: without `--version` Helm silently picks the newest version *it knows about*; with `--version` it errors that no such version exists in that repository's index and suggests running `helm repo update`. Fix it by running `helm repo update <repo>` in the job and pinning `--version`, so a stale cache fails loudly instead of deploying yesterday's chart.

code

bash · 8 lines
bash
# Deterministic deploy step: refresh the one repository, then pin the version
helm repo update billing-charts
helm upgrade --install billing-cron billing-charts/billing-cron \
  --version 4.7.3 \
  -n billing -f prod-values.yaml

# Stale cache or a publish that never landed?
helm search repo billing-charts/billing-cron --versions | head -5

go deeper

for a junior

Remember that a repo/chart reference is resolved from a local copy of the repository's index, and that helm repo update is the command that refreshes it. Run it before installing something newly published.

for a middle

Explain both symptoms — silently installing an older chart when unpinned, and a not-found-in-index error when pinned — and say why search repo agrees with the stale cache rather than correcting it.

for a senior

Diagnose it from the pipeline outward: warm cache volumes and pre-baked images, then a fix that refreshes explicitly and pins the version so the next stale cache fails loudly instead of deploying yesterday's chart.

for a principal

Own the reproducibility rule for the estate: what is allowed to be resolved at deploy time at all, whether charts are referenced by version or by digest, and what the build cache may retain without making deploys non-deterministic.

### Where version resolution actually happens A chart reference of the form `<repo>/<chart>` is not a network address. `helm repo add` records the repository's URL in the client's repository configuration and downloads its index into the repository cache directory (the one `helm env` prints as `HELM_REPOSITORY_CACHE`). Every later command that has to turn `<repo>/<chart>` into a concrete chart — `install`, `upgrade`, `pull`, `template`, `search repo`, and dependency resolution for a chart that lists that repository — consults that **local file** to learn which versions exist and where the archive for each one lives. Only once a specific version has been chosen does Helm fetch the archive itself. That design is why the tool is fast and why it is wrong so often. The cached index is a snapshot. Nothing expires it, nothing refreshes it in the background, and no command refreshes it as a side effect. `helm repo update` — optionally naming a single repository rather than all of them — is the only thing that re-downloads it. ### The two symptoms They look nothing alike, and only one of them is loud. **Unpinned.** `helm upgrade --install billing-cron billing-charts/billing-cron` with no `--version` asks for the newest version, and Helm answers with the newest version *in the cached index*. If that snapshot is nine days old, the deploy quietly ships a nine-day-old chart. Everything succeeds. The release gets a new revision. Nothing in the output says the word "stale". **Pinned.** The same command with `--version 4.7.3` fails, because 4.7.3 is not in the snapshot: Helm reports that no chart matching that version was found in that repository's index and points you at `helm repo update`. This is the better failure by a wide margin — it is immediate, specific and safe. The same staleness makes `helm search repo` lie in the same direction, which is worth knowing because search is the command people reach for to "check" whether the new version exists. It reads the cache, so it confirms the wrong answer. ### Why CI is where this bites On a laptop the cache is refreshed accidentally — you add a repository, you run an update, you install something else. Automation has none of those accidents: - A **pre-baked runner image** ran `helm repo add` at image build time. Every job it ever runs starts from that image's snapshot until the image is rebuilt. - A **cached workspace volume** keyed on something stable restores the same repository cache into every run, deliberately, to save a download. - A **short-lived container** that adds the repository fresh each run is immune — which is why the same pipeline can be correct for months and then break the day someone adds caching to make it faster. ### Worked example A subscription-billing cron ships as a chart guarded by a `values.schema.json` contract, deployed with an 11-value production override file. Version 4.7.3 is published at 09:14. The 09:52 pipeline deploys 4.6.9. The release succeeds, `helm history` shows a fresh revision, and the schema accepts the override file exactly as before — so nothing in the deploy looks wrong; only the running image tag disagrees with the release notes. On the runner, `helm env` shows a repository cache restored from a warm volume, and `helm search repo billing-charts/billing-cron --versions` tops out at 4.6.9. One `helm repo update billing-charts` and the same search shows 4.7.3. ### Making it deterministic 1. **Refresh explicitly in the job**, naming the repository so you are not paying to update a dozen unrelated ones. 2. **Always pin `--version`.** This converts the silent-old-chart failure into a loud one. The pin does not remove the need to refresh; it makes forgetting to refresh detectable. 3. **Decide deliberately what the cache is for.** Caching downloaded chart archives is a genuine saving; caching index files buys milliseconds and costs correctness. If the pipeline caches Helm directories, exclude or refresh the index. 4. **Consider an `oci://` reference.** A chart pulled from an OCI registry by reference has no index in the middle: the tag is resolved against the registry at pull time, so there is no cached listing to go stale. It is not magic — a mutable tag can still be moved under you, and the right defence there is a digest or an immutable tag — but the specific failure mode of "my local catalogue is out of date" does not exist. ### Distinguishing stale from genuinely missing When the pinned install fails, one command separates the two causes: run `helm repo update <repo>` and search again. If the version now appears, the cache was stale and the pipeline is at fault. If it still does not, the publish never landed — and the investigation moves to whatever pushes charts into that repository, which is a different team's problem and a different conversation.

  • Pinning `--version` does not refresh anything. Why is it still part of the fix?
    Because it changes the failure mode. Unpinned, a stale index produces a successful deploy of an older chart, which nobody notices until an incident. Pinned, the same staleness produces an immediate error naming the version that could not be found and suggesting a repository update. You still need the refresh in the job; the pin is what guarantees you find out when the refresh did not happen.
  • The pipeline caches the whole Helm cache directory between runs to save time. What would you change?
    Keep the downloaded chart archives, refresh the index. Archives are content that never changes for a given version, so caching them is a real saving; index files are a listing whose whole purpose is to be current, and caching them buys almost nothing while making resolution non-deterministic. Concretely: run `helm repo update` in the job even with a warm cache, or exclude the index files from the restored cache.
  • How do you tell a stale cache apart from a chart version that was never actually published?
    Refresh that one repository and search again. If the version appears afterwards, the cache was stale and the fault is in your pipeline. If it still does not appear, the publish failed and the investigation moves upstream to whatever pushes charts to that repository. Doing the refresh first stops you escalating a local cache problem to another team.

The cached index is a printed catalogue. The shop restocks daily, but you keep ordering from the copy you picked up last month — and it never occurs to you that the catalogue itself is out of date.

saying these in an interview costs you the question

  • Thinks helm install queries the repository server every time
  • Believes the cached index expires on a timer
  • Runs helm repo update only when adding a repository
  • Uses helm search repo to confirm a version exists without refreshing
  • Blames the chart publisher before refreshing the local index
  • Deploys without --version and calls the result reproducible

context