Before a production `helm upgrade`, what must a change preview prove, and where does it stop helping?
answer
- three checks, not one
- delta, acceptability, drift
- a noisy diff is an unread diff
- gate on deletions and replacements
- preview is consent, not correctness
basics
~20 sA preview should prove three separate things: the delta against what Helm last applied, that the cluster would accept the objects, and that nothing drifted since the last upgrade. It cannot prove the rollout will be healthy, and it never prevents the next out-of-band edit.
solid answer
~40 sDecide what each preview step buys and price it. A plugin diff against the stored manifest shows the delta a reviewer approves; a server-side dry run proves the API server and admission accept the objects; a stored-manifest-versus-live diff proves the cluster still matches the record. Then set the noise budget: charts that embed timestamps or random values make every diff differ, so previews stop being read, and the authoring rule is part of the gate. Tier the response — block automatically on object deletion or an immutable-field replacement, treat the rest as review material — because gating every routine upgrade on human diff reading trains rubber-stamping. Be explicit about the limits: previews say nothing about rollout health, never cover `crds/`, go stale between review and apply, and are no substitute for a rollback plan.
code
bash · 11 linesset -euo pipefail
REL=reranker-api NS=reco CHART=./reranker-api VALS=prod-values.yaml
# 1. Baseline check: does the cluster still match what Helm recorded?
helm get manifest "$REL" -n "$NS" | kubectl diff -f - || echo "DRIFT: review before upgrading"
# 2. Acceptability: validation, defaulting and admission on the target cluster
helm upgrade "$REL" "$CHART" -n "$NS" -f "$VALS" --dry-run=server > /dev/null
# 3. The delta a human approves
helm diff upgrade "$REL" "$CHART" -n "$NS" -f "$VALS"go deeper
Learn the habit before the policy: never run an upgrade against production without looking at what would change first, and know which command produces that output.
Be able to say what each preview step actually proves, and to name at least two things none of them prove — rollout health and anything installed from crds/ are the easy ones.
Show that you have operated this: how you keep the diff readable, what you gate on automatically, what happens when the drift check finds a production edit, and how the preview interacts with your rollback plan.
Own the whole shape: which checks are mandatory for which blast radius, the chart-authoring standard that keeps previews meaningful, who gets production credentials for a server dry run, and when the honest answer is that this team needs continuous reconciliation rather than a better preview.
## Name the three questions before designing the gate Teams say "we preview our Helm upgrades" and mean one of three unrelated checks. Separating them is most of the leadership value here. 1. **The delta.** A plugin diff renders the new chart and compares it with the manifest stored in the current release revision. This is the artefact a human reviews, and the only one that answers "what does this change actually do?" 2. **Acceptability.** `--dry-run=server` pushes the rendered objects through schema validation, defaulting and admission on the target cluster without persisting anything. This catches a removed API version, a schema violation on a custom resource, and a policy webhook that will deny the object. 3. **Drift.** Piping `helm get manifest` into a server-side diff compares the record with reality. Neither of the first two does this: both take Helm's record as their notion of the world. A gate that runs only (2) and calls itself a preview will happily let through a change whose diff nobody looked at. A gate that runs only (1) will pass a manifest the cluster refuses. ## The noise budget decides whether anyone reads it A preview is a human artefact, and its value collapses the moment it is unreadable. Take `reranker-api`, a stateless service running a recommendation re-ranker, whose chart carries the standard `checksum/config` annotation on its pod template so that a config change rolls the pods. That is good chart design and it also means every config edit shows a pod-template change in the diff — legitimate but recurring. Now add a template that embeds a render timestamp or a random suffix, and every preview differs on every run; within a month reviewers approve without reading, and the gate is theatre. So the preview policy has an authoring clause: charts that a preview gate protects must render deterministically from the same inputs. That is a platform standard, not a per-team preference, and it is enforceable in review because non-determinism is visible as a diff on a no-op run. ## Tier by blast radius, not by ceremony With 34 revisions on that release in a quarter, most upgrades are routine image bumps. Requiring a human to read every diff at that cadence guarantees they stop reading. A workable policy blocks automatically on the categories that actually hurt — an object disappearing from the render, a change to an immutable field that forces a replacement, a change to a StatefulSet's volume claims, a namespace or name change that would orphan the old object — and treats everything else as posted-for-information. That way the rare dangerous diff arrives on a clean desk. ## Costs you are actually signing up for A plugin diff is a third-party CLI plugin in the release path: someone owns its version, its provenance and its behaviour when it breaks. A server-side dry run needs credentials against the production cluster at *review* time, which is a genuine widening of who can talk to production, and it runs admission webhooks for real, so a webhook that is not dry-run-safe can block changes that would apply fine. Both add latency to every deploy. None of that argues against previews; it argues for deciding deliberately rather than accumulating steps. ## Where preview stops Be blunt about the limits, because this is where principals are separated from enthusiasts. - **Health is not previewable.** Every object can be accepted and the diff can be exactly what you intended, and the new pods can still fail readiness. The preview is about the change; the wait strategy, the readiness gates and the rollback path are about the outcome. - **`crds/` is out of scope.** Helm installs those files once and never updates or deletes them, so no preview of the templates covers a CRD change; that migration is a separate, manual plan. - **Previews go stale.** The diff and the dry run described the cluster at the moment they ran. If review takes hours, another actor can move the world underneath the approval. Narrow the window, or re-run the checks at apply time and fail on a changed baseline. - **Preview does not prevent drift.** It detects it, once, when someone runs it. Continuous enforcement of a desired state is a different tool class with a different operating model — an in-cluster reconciler — and choosing it is a delivery-model decision, not an improvement to your preview script. Say that explicitly rather than trying to build a reconciler out of a scheduled diff. - **Preview is not a rollback plan.** The complement to a good preview is a cheap reversal: bounded revision history so the target revision still exists, a wait strategy that fails fast, `--rollback-on-failure` where an automatic reversal is genuinely safe, and a rehearsed manual `helm rollback` where it is not. The honest summary to give an interviewer: previews buy you *informed consent* about a change. They do not buy correctness, they do not buy convergence, and a team that believes otherwise will be surprised by the first field they never templated.
- Would you require a human to approve every production upgrade diff?No. At a routine cadence that trains rubber-stamping. I would auto-approve diffs confined to expected shapes — an image tag, a replica floor, a config value — and require a named approver only for the categories with real blast radius: an object disappearing, an immutable field forcing a replacement, a change to persistent-volume claims, or anything touching a shared resource. The rule should be mechanical enough that people trust it.
- A team argues the drift check is pointless because the next upgrade overwrites everything anyway. What do you say?That it is not true and the belief is the danger. Whether an upgrade reasserts a field depends on the apply path and on whether the chart templates that field at all; a hand-set value on a field the chart never mentions survives indefinitely. More importantly, the drift check tells you a human changed production out of band, which is a process finding worth having regardless of whether the value gets overwritten later.
- How do you keep a preview from becoming stale between review and apply?Shorten the window and re-verify at the end of it. Run the drift check and the server dry run again immediately before the apply and fail if the baseline moved, rather than trusting a diff approved hours earlier. Pin the exact chart artefact and values that were previewed so the apply cannot differ from the review, and keep the approval attached to that pinned pair rather than to a branch.
saying these in an interview costs you the question
- Treats a passing server dry run as a preview of the change
- Requires human diff review on every routine upgrade
- Believes a preview predicts a healthy rollout
- Ignores that non-deterministic charts make diffs unreadable
- Assumes previews cover resources installed from crds/
- Offers preview as a substitute for continuous reconciliation