skip to content

How would you make review of third-party Helm chart upgrades repeatable across teams?

level: principalimportance: nice to knowfreq 32%

answer

  1. Nobody re-reads twelve thousand lines twice
  2. Shrink the unit of review, not the standard
  3. Same values, two chart versions, one diff
  4. Write down what forces a human to stop
  5. A cluster-side control is the backstop, not the review

basics

~20 s

Make the review artefact a rendered diff, not a chart. Pin exact chart versions, render the old and new version with the same production values, and have a human read the delta — new hooks, new cluster-scoped objects, changed images.

solid answer

~40 s

Nobody re-reads twelve thousand lines of vendor YAML per upgrade, so a policy that asks for it produces rubber stamps. Design around the delta. Pin every third-party chart to an exact version in the repository that deploys it, so an upgrade is a commit. In CI, render the pinned version and the proposed version with the same production values (`--include-crds`), diff, and publish that diff as the review artefact. Then define what forces a human read: a new or changed `helm.sh/hook` manifest, a new cluster-scoped object, a changed image reference, a new subchart. Everything else passes on a scan. Be explicit that this proves nothing about what the images do, and keep a cluster-side control as the backstop. Name an owner per vendor chart; treat 'no owner' as a reason not to adopt it.

code

bash · 3 lines
bash
helm template checkout ./charts/checkout-platform-4.7.2 -f prod-values.yaml --include-crds > old.yaml
helm template checkout ./charts/checkout-platform-5.0.0 -f prod-values.yaml --include-crds > new.yaml
diff -u old.yaml new.yaml

go deeper

for a junior

Know that upgrading a third-party chart changes what gets applied, and that comparing the rendered output of the old and new versions is how teams see that change.

for a middle

Be able to produce the diff: render both chart versions with the same values file and --include-crds, then explain which kinds of change in that diff matter more than others.

for a senior

Show how you would run this per environment in a pipeline, and which classes of change — hooks, cluster-scoped objects, image references, new subcharts — you would make blocking.

for a principal

Own the tradeoffs: mirror versus upstream, fork versus configure, what the review can never prove, and what has to be enforced cluster-side instead. Name an owner per vendor chart and a cadence, or the process decays into version bumps.

### The failure this is fixing The honest state of third-party chart review in most estates is: one engineer read the render once, at adoption, and every upgrade since has been a version bump in a values repo. That is not a discipline problem. A vendor chart renders thousands of lines; a minor bump changes forty of them; asking a reviewer to re-read the whole thing every time guarantees they read none of it. Any workable answer has to make the *unit of review* small enough that reading it is realistic. ### Make the artefact a diff The mechanism Helm gives you is that a render is a pure function of chart plus values. If both versions are rendered with the same production values, the difference is exactly what the upgrade does to your cluster. ``` helm template checkout ./charts/checkout-platform-4.7.2 -f prod-values.yaml --include-crds > old.yaml helm template checkout ./charts/checkout-platform-5.0.0 -f prod-values.yaml --include-crds > new.yaml diff -u old.yaml new.yaml ``` Three design decisions make that reliable. **Pin exact versions**, never a range — an upgrade must be a commit someone can point at, and the same input must render the same output tomorrow. **Render with real values per environment**, because the same chart bump can be inert in staging and add an object in production. **Include `--include-crds`**, since CRDs are the one class Helm installs and then never upgrades or deletes; a CRD change arriving in a chart bump is exactly the change you cannot let slide past. ### Define what forces a human read A diff still needs a reading policy, or it becomes the same rubber stamp at a smaller size. The classes worth stopping on are the ones whose blast radius exceeds the namespace or whose effects are not reversible: a new or changed manifest carrying a `helm.sh/hook` annotation, especially one that touches a database; any new cluster-scoped object; a changed or added `image:` reference; a new subchart appearing under `charts/`; a new initContainer or a pod spec gaining privileged, hostPath or hostNetwork. Everything else — replica counts, labels, a new optional annotation — can pass on a scan. Writing that list down is what converts "review the chart" from an aspiration into a check a reviewer can actually complete in ten minutes. ### Where the review does not reach Be explicit with yourself about the limits, because overclaiming here is how a control quietly becomes theatre. Reading a rendered manifest tells you what objects will exist and which images they run; it tells you nothing about what those images do once running. It is a point-in-time check on a commit, so it does not constrain anything installed by hand outside that path. And it is advisory by construction — the reviewer can approve anything. The complement is a cluster-side control that enforces your non-negotiables regardless of which chart, which team or which install path produced the manifest; an admission policy engine is that backstop, and the review is what catches the specific, judgement-laden things a general rule cannot express. ### The organisational half The tooling is the easy part. The decisions a lead actually owns are: **who owns each vendor chart** — a chart with no named owner has no reviewer, and "no owner" is a good reason to decline adoption rather than a state to live with; **mirror or not** — copying charts into an internal repository buys you reproducibility and a chokepoint, at the cost of a sync job and staleness nobody notices until an upgrade is urgent; **fork or configure** — vendoring a chart to change what values cannot reach gives total control and a permanent merge obligation, so it should be a decision with an owner, not a workaround an engineer reaches for at 6pm; and **how upgrades get scheduled**, because a review process that only runs when someone feels like upgrading produces charts three years stale, where the diff is now enormous and the review is impossible again. Cadence is part of the control. ### How to argue it in an interview The strong answer names the tradeoff rather than a tool: you are buying determinism and a small, readable delta at the cost of pipeline machinery, pinned versions that must be actively maintained, and a review policy someone has to keep honest. The weak answer is "we review charts before installing them", which is what everyone says and nobody does past the first quarter.

  • Why pin an exact chart version rather than a version range for a third-party chart?
    Because a review is only meaningful if the thing reviewed is the thing installed. A range lets the resolved chart move between the review and the deploy, and between two deploys of the same commit, so the render you approved is not reproducible. Pinning makes every upgrade an explicit commit with a diff attached, which is also the only way the cadence is visible.
  • Your render diff is clean but the upgrade still broke production. What did the review structurally miss?
    A diff compares manifests, so it is blind to everything that is not a manifest: what the new image actually does, a hook Job's runtime behaviour against real data, and any drift between the stored release and the live objects. It also renders with the values in the repository, so an out-of-band `--set` used by a human at install time never appears in it.
  • When is vendoring a third-party chart into your own repository the right call?
    When you need changes values cannot reach — a hardcoded image, a hook you must remove, an object you must not create — and the alternative rewriting step would be more obscure than a fork. It is a decision with a standing cost: you now merge every upstream release. Assign it an owner, or the fork silently becomes a three-year-old chart nobody can upgrade.

saying these in an interview costs you the question

  • Requiring a full re-read of every render each upgrade
  • Reviewing charts pinned to a floating version range
  • Rendering with default values instead of per-environment ones
  • Claiming manifest review proves the images are safe
  • Leaving crds/ out of the diff that gets reviewed
  • Adopting vendor charts with no named owner

context