skip to content

What does helm repo index --merge do that regenerating the index alone does not?

level: middleimportance: nice to knowfreq 34%

answer

  1. The command has no memory of before
  2. It describes one directory, nothing else
  3. Something must carry old entries across
  4. Fold the live index into the generated one
  5. Read-modify-write on a shared file races

basics

~20 s

helm repo index builds index.yaml from only the packages present in the directory. --merge folds an existing index.yaml into that result, so versions whose tarballs are not in that directory keep their entries instead of disappearing.

solid answer

~50 s

`helm repo index <dir>` is a full regeneration: it lists the `.tgz` files actually in `<dir>` and writes an `index.yaml` describing exactly those. That is fine when the directory holds the whole published catalogue, and catastrophic when it does not - a CI workspace usually contains only the package just built, so regenerating there and uploading the result silently drops every previously published version from the catalogue, even though the tarballs are still sitting on the host. `--merge <existing-index.yaml>` folds a copy of the live index into the generated one so those older entries survive. The two workable patterns are therefore: keep the full catalogue on the publish host and re-index everything, or fetch the live `index.yaml` first and merge into it. `--url` is orthogonal - it sets the base URL written into new entries' `urls`.

code

bash · 7 lines
bash
helm package ./digest-builder      # -> digest-builder-2.7.3.tgz
curl -fsSL https://charts.internal.example/charts/index.yaml -o previous-index.yaml
helm repo index . --merge previous-index.yaml \
  --url https://charts.internal.example/charts

# upload digest-builder-2.7.3.tgz first, then the regenerated index.yaml
grep -c 'digest-builder-' index.yaml

go deeper

for a junior

Remember one sentence: the generated index describes only the packages sitting in the directory you point at. If you have ever wondered why old versions vanished after a publish, that sentence is the whole answer.

for a middle

Explain the regeneration model and both publishing patterns - re-index the whole catalogue, or fetch and merge the live index - and say what each costs. Knowing that the flag takes a path to an existing index file, not a directory or URL, is the detail that shows you have used it.

for a senior

Talk about the index as shared mutable state: the lost-update race between concurrent publishes, ordering the tarball upload before the index upload, and the difference between a version missing from the catalogue and a catalogue entry whose file is gone.

for a principal

Decide whether the catalogue should be derived state at all. Re-indexing from the stored packages makes the files the source of truth and removes a class of race, at the cost of touching the whole catalogue per publish - the same choice you make in any cache-versus-recompute design.

This flag exists because `helm repo index` has no memory. It is a pure function of one directory: open every chart package in it, copy each chart's metadata into an entry, write `index.yaml`. Nothing about the resulting file is incremental. If a version's `.tgz` is not in the directory when you run the command, that version is not in the file you just produced. ## The failure this causes The classic incident looks like this. A chart called `digest-builder` - a single-service chart with an Ingress and a HorizontalPodAutoscaler for an email-digest builder - has 63 published versions living in an object-store bucket, listed in a 412 KiB `index.yaml`. The publish job runs `helm package`, producing one tarball in the job's empty workspace, then runs `helm repo index .` there and uploads both files. The upload succeeds. The bucket still holds all 63 tarballs. But the catalogue now lists exactly one version, so `helm search repo` shows one version, `--version 2.6.1` fails to resolve for every consumer that refreshes, and any pipeline that installs a pinned older version breaks the moment its cached index expires. Nothing was deleted; the map to the data was replaced with a map of one house. ## What --merge actually does `helm repo index <dir> --merge <path-to-index.yaml>` generates the index from the directory as usual, then folds in the entries of the index file you point at, so chart versions that exist only in that older index are carried into the output. The result is a catalogue that is the union of what you just built and what was already published. The flag takes a *file path*, so the normal shape in a pipeline is to download the live index first: ```bash curl -fsSL https://charts.internal.example/charts/index.yaml -o previous-index.yaml helm repo index . --merge previous-index.yaml \ --url https://charts.internal.example/charts ``` ## The two publishing patterns **Re-index the whole catalogue.** Keep every published tarball in one place - typically syncing the bucket down, or generating the site from a directory that holds all of them - and run a plain `helm repo index` over it. No merge needed, and the index is always a faithful description of what is actually there. This is the more robust option and it is what static-site style chart repositories usually do; the cost is transferring the whole catalogue on every publish. **Merge into the live index.** Fetch `index.yaml`, generate from a directory containing only the new package, merge, upload the new tarball and the new index. Cheap and fast, and the option most CI jobs reach for. The cost is that the index becomes state you edit blind: it can now describe versions whose files are gone, and it is exposed to a lost-update race. ## The race worth naming Merging is read-modify-write against a shared file with no locking. Two publish jobs that both download `index.yaml` at the same moment each merge their own package into their own copy; whichever uploads second overwrites the first, and that version vanishes from the catalogue while its tarball sits perfectly healthy in the bucket. It is a lost update, and it is invisible until someone tries to install the missing version. Serialise publishes for a given repository - a single-concurrency job or a lock - or use the full re-index pattern, where the source of truth is the set of files rather than a snapshot of a file. ## What --merge does not do It does not upload anything - publishing remains a file copy you arrange yourself. It does not repair entries in the old index: an entry that points at a wrong or dead URL keeps pointing there, because merged entries are carried across as they are rather than recomputed. And it is not a substitute for keeping the packages: a merged index entry whose tarball has been deleted is a 404 waiting to happen, and it will surface as a download failure at install time rather than as a search-time absence, which is a considerably more confusing report to receive. ## The neighbouring flag `--url` is often confused with `--merge` because they are usually typed together. `--url` sets the base URL prefixed onto the package filenames in the generated entries' `urls` field. You need it whenever the index is generated somewhere other than where the files are served from - which, in a pipeline, is always. Entries carried in by `--merge` keep the URLs they already had, so a change of hosting requires regenerating the whole catalogue rather than merging on top of it.

  • Your publish job downloads the live index, merges, and uploads it. Two jobs run at the same time - what goes wrong?
    A lost update. Both jobs read the same index, each merges its own package into its own copy, and the second upload overwrites the first - so one version disappears from the catalogue while its tarball sits untouched on the host. Fix it by serialising publishes to a repository, or by re-indexing the full set of tarballs from the storage location instead of merging a snapshot.
  • What does the --url flag change in the generated index, and why do pipelines always need it?
    It sets the base URL prefixed onto each generated entry's `urls` field, which is what clients actually fetch. A pipeline generates the index in a workspace that has nothing to do with the public hosting path, so without `--url` the entries do not point at where the packages are served from. Entries brought in by `--merge` keep the URLs they already carried.
  • Can --merge resurrect a version whose tarball you deleted from the host?
    It will happily carry the entry across, which is worse than useless: the catalogue advertises a version whose file is gone, so consumers get a download failure at install time instead of a clean "no such version". If a version is genuinely withdrawn, remove it from the index too - and never republish that number with different content.

saying these in an interview costs you the question

  • Assumes re-indexing preserves versions not in the directory
  • Runs helm repo index in a workspace holding one tarball
  • Thinks --merge uploads or publishes anything
  • Merges a stale index copy while another job publishes
  • Expects --merge to fix wrong URLs in old entries
  • Believes helm repo index reads chart sources, not packages

context