Piping `helm get manifest` into `kubectl diff` shows a large diff on a release nobody touched — why?
answer
- an empty diff is not the expectation
- ask who owns each changed field
- controllers and webhooks write too
- hooks and crds are out of scope
- a stale config checksum means no restart
basics
~20 sMost of that diff is expected: fields other controllers own, server-assigned bookkeeping metadata, and objects the stored manifest never covered. Real drift is what a human or an unexpected actor wrote, so triage the diff by asking who owns each changed field.
solid answer
~50 sThe stored manifest is what Helm rendered, and the live object is that plus everything the cluster did to it since. Expect differences from fields another controller owns — an autoscaler rewriting `spec.replicas`, a mutating webhook injecting a sidecar or volume, a controller stamping labels — plus server-assigned bookkeeping metadata. Also expect gaps rather than noise: hook-annotated manifests are not in `helm get manifest` output, and anything installed from `crds/` is never touched by upgrades, so the diff cannot speak to it. Triage by asking, per changed field, *who wrote this*. A field the chart templates that now holds a different value is real drift; a field the chart never mentions is usually someone else's business. On the release I would then check whether an edited ConfigMap's contents match the `checksum/config` annotation the pod template carries — if they do not, the config changed and the pods never restarted.
code
yaml · 9 linesapiVersion: apps/v1
kind: Deployment
metadata:
name: reranker-api
spec:
template:
metadata:
annotations:
checksum/config: {{ include (print $.Template.BasePath "/configmap.yaml") . | sha256sum }}go deeper
Understand that the stored manifest is only what Helm wrote, and the live object has been touched by the cluster since, so some differences are normal. Know the command that produces the comparison.
Be able to name the categories of expected difference — controller-owned fields, injected sidecars, server bookkeeping — and explain why hook manifests and crds/ resources never appear in the comparison at all.
Walk the triage: classify each changed field by owner, isolate the one that a human wrote, check whether a config change ever reached the pods, and decide whether the chart or the cluster is the thing to change.
Own the signal-to-noise problem. Decide what a fleet-wide drift report is allowed to alert on, how charts must be authored so the report stays readable, and what your policy is when the cluster is repeatedly right and the chart is wrong.
## Why a diff is noisy by construction `helm get manifest` prints the text Helm rendered and applied. The live object is that text, plus everything the API server and every controller has done to it since: defaulting, status, ownership metadata, and fields other actors legitimately own. Piping one into a server-side diff therefore never produces an empty result on a busy cluster, and a senior answer is about **classifying** the output rather than expecting silence. The useful triage question, applied field by field, is *who wrote this?* **The chart wrote it, and it changed.** This is real drift. A `kubectl edit`, a `kubectl scale`, a quick incident fix, or an out-of-band apply by another pipeline. These are the lines you care about. **Another controller owns it.** An autoscaler owns `spec.replicas`; a mutating admission webhook may inject a sidecar container, a volume and environment variables into every pod template; controllers add their own labels and annotations. These show up on every diff forever and are not drift in any actionable sense — they are the cluster doing its job. The fix is not to chase them but to know which ones your platform produces, so a reviewer can skip them by name. **The server wrote it.** Bookkeeping metadata — resource versions, generation counters, managed-field records, creation timestamps and status content — belongs to the API server. A dry-run-based diff defaults your input before comparing, which removes a lot of this class, but not all of it. **Nothing wrote it, because it was never in scope.** Two gaps matter. Hook manifests, annotated with `helm.sh/hook`, are not part of what `helm get manifest` prints, so a Job left behind by a hook produces no diff line. And resources installed from `crds/` are never updated or deleted by Helm, so a CRD that has drifted — or that has been changed deliberately — is invisible to any diff of the templates. ## A worked example Take `reranker-api`, a stateless API serving a recommendation re-ranker, on revision 34 of its release. The chart is small: a Deployment, a Service, a ServiceAccount, a ConfigMap of scoring parameters, and a HorizontalPodAutoscaler. `helm get manifest reranker-api -n reco | kubectl diff -f -` returns 41 changed lines across three objects. Sorting them: - `spec.replicas` on the Deployment reads 11 live against the chart's 4. The chart also ships an autoscaler for this workload, so that field is owned elsewhere. Not drift — and worth fixing at the source by not templating `replicas` when an autoscaler owns it. - The pod template has an extra container, a projected volume and two environment variables that appear in no template in the chart. That is a mutating webhook on the namespace. Not drift. - The ConfigMap's `scoring.yaml` key differs: a threshold reads `0.63` live against `0.55` in the stored manifest. Nobody's controller writes that. **This is the real finding.** The chart carries the standard restart-on-config-change idiom — a `checksum/config` annotation on the pod template, computed from the ConfigMap file at render time. Because the live edit was made with `kubectl edit` on the ConfigMap and not through Helm, the Deployment's annotation was never recomputed, so no rollout happened. That gives you the second half of the finding: not only did the configuration change out of band, but depending on how the application reads it, the running pods may still be using the old value — or may have picked it up silently through the mounted file, which is worse, because now the record, the pod template annotation, and the behaviour of the process all disagree. ## Closing it out Decide which side is right. If `0.63` was a deliberate tuning change, it belongs in the chart values and a normal upgrade should put it there, which also recomputes the annotation and rolls the pods deliberately. If it was an experiment, re-running `helm upgrade` with the current chart and values reasserts the stored value; you can preview exactly that with a diff first. Either way the goal is to get the record and the cluster back into agreement, because every future preview takes the record as its baseline and silently inherits the discrepancy until you do. Finally, make the check repeatable. A scheduled job that runs the stored-manifest diff per release and reports only the objects and fields your chart actually templates — filtering the known controller-owned ones — turns a wall of expected noise into a short, trustworthy signal. A drift check nobody trusts is a drift check nobody reads.
- How do you stop `spec.replicas` showing up in every drift check for an autoscaled workload?Stop templating the field. If an autoscaler owns the replica count, the chart should omit `replicas` from the Deployment when autoscaling is enabled, which is what the scaffolded chart does. Then Helm never asserts a value, the field has one owner, and the diff goes quiet. Filtering the field out of the report instead leaves a chart that fights the autoscaler on every upgrade.
- A ConfigMap in the release was edited by hand. Why might the pods still be serving the old values?Because a ConfigMap edit does not restart anything. If the chart ties a checksum of the config into the pod template annotation, only a Helm-driven change recomputes it and rolls the pods; an out-of-band edit leaves the annotation untouched. Whether the process picks the change up depends on how it reads the data — a mounted file updates in place after a delay, an environment variable never does.
- Which parts of a release can this diff never cover?Hook manifests, which are stored separately from the release manifest, and anything installed from `crds/`, which Helm installs once and then never updates or deletes. Objects the chart deliberately renders empty are also absent. Add resources created at runtime by an operator the chart installed — they were never in the manifest, so their absence from the diff means nothing about their state.
saying these in an interview costs you the question
- Expects the diff to be empty on a healthy release
- Calls every difference drift without asking who owns the field
- Templates replicas while an autoscaler owns the field
- Thinks editing a ConfigMap restarts the pods that mount it
- Assumes hook Jobs and crds/ resources appear in the manifest diff
- Fixes drift by editing the cluster again instead of the chart