In Helm, how do you decide what belongs in chart defaults versus a per-environment values file?
answer
- Sort the differences before shrinking them
- Some divergence is size, some is behaviour
- Same keys everywhere, different values
- Constant everywhere means it is a default
- Move the version or the values, not both
basics
~20 sChart defaults carry everything true of every environment. The per-environment file carries only the axes that legitimately differ — sizing, wiring and identifiers. A difference in behaviour, such as a feature enabled only in production, is debt to budget rather than configuration.
solid answer
~50 sI sort every per-environment key into three buckets. **Sizing** — replicas, requests, autoscaler bounds — is legitimate and expected to differ. **Wiring** — hostnames, storage classes, external endpoints — is legitimate but must use the *same keys* in every environment, so the three files diff cleanly. **Behaviour** — a feature flag, an extra sidecar, a subchart enabled only in production — is the dangerous bucket, because production then runs code paths staging never exercised; I treat each of those as debt with an owner and an expiry. Anything constant goes back into the chart so it is written once. The guardrails are a values schema plus a CI job that renders all three environments and reports keys present in production but absent from staging. And I keep the two promotion axes separate: a deploy moves the chart version or the values, never both.
code
bash · 3 lineshelm get values fraud-platform -n fraud-staging > staging-values.yaml
helm get values fraud-platform -n fraud-prod > prod-values.yaml
diff -u staging-values.yaml prod-values.yamlgo deeper
You are unlikely to be asked this, but know the basic rule you will be held to: put anything that is the same everywhere in the chart's defaults, and put only real differences in your environment's file.
Be able to classify a given key — sizing, wiring or behaviour — and say why a behaviour difference between staging and production is worse than a replica-count difference. Know why identical values repeated in three files belong in the chart.
Show the guardrails you would actually put in: a values schema fixing the shape, every environment rendered on every change, a report of production-only keys, and a rule that a deploy moves the chart version or the values but not both.
Own the policy and its cost. Decide how much environment divergence the organisation buys, who owns chart defaults versus environment files, when a production-only difference is legitimate rather than debt, and when one chart should stop trying to express two architectures.
## The question behind the question Every environment file starts small and grows. For an umbrella chart bundling a fraud-scoring API and a worker, the pattern is familiar: `values/dev.yaml` at 38 lines, `values/staging.yaml` at 74, and `values/prod.yaml` at 214, with a third of production's keys appearing in no other file. The interesting question is not how to reduce the line count, it is which of those differences are *supposed* to exist. ## A taxonomy worth having **Sizing.** Replica counts, CPU and memory requests, autoscaler minimum and maximum, queue concurrency, cache sizes. These differ by design and always will; nobody runs eleven API replicas in dev. Sizing divergence is free: it changes how much of the same behaviour runs, not which behaviour runs. The only discipline needed is that the keys exist in every file, so the difference is a number rather than an absence. **Wiring.** Ingress hostnames, storage class, the address of the payments system, the name of the Secret holding credentials, the namespace of a dependency. These differ by necessity. The rule here is uniformity of *shape*: every environment sets the same key, even when dev's value is a stub. A key that exists only in production is a key no other environment has ever exercised. **Behaviour.** A feature flag on in production only, a sidecar added for production, the worker subchart enabled in production and disabled in staging, a template branch reachable only under one environment's values. This is the bucket that causes incidents, because it means the artifact you promoted was never run in the configuration you promoted it into. Staging's whole value proposition is that it ran this chart with this behaviour; behaviour divergence spends that value. Chart defaults, meanwhile, should hold everything in none of those buckets: the probe paths, the port, the label conventions, the config-file structure, the sane starting numbers. If a value is identical in all three files, it is not environment configuration — it is a default that has been copied three times and will eventually be updated in two of them. ## Budgeting divergence Treat behaviour divergence like any other debt: allowed, counted, owned, and expiring. A production-only flag needs a name, a reason, and a date by which staging runs it too. The cheap instrument is a CI check that compares the key sets of the environment files — or better, the values Helm actually used, via `helm get values` per release — and reports keys present in production and missing from staging. It should not necessarily fail the build; it should be visible in the pull request, because the point is that someone decides rather than that nobody notices. The strongest structural control is the chart's `values.schema.json`, which fixes the shape all three files are instances of and, with `additionalProperties: false`, makes a key that exists in only one environment a deliberate schema change rather than a typo nobody sees. ## Things that look like configuration and are not An environment-name conditional inside a template — branching on `.Values.environment` to decide whether an object exists — is behaviour divergence that has been hidden from the place people review it. The values files will look reassuringly similar while the rendered manifests differ. Express the difference as a named value with a default, so it appears in the diff. Equally, per-environment differences that are really topology differences — production needs an extra component, or a different one — are a signal to ask whether the thing should be a subchart toggle, a separate release with its own values, or genuinely a different chart. Bending one chart until it can express two architectures produces templates nobody can read. ## Two axes, one at a time Promotion has exactly two variables: the chart version and the environment's values. A deploy that moves both leaves an incident review with no way to attribute the failure. The policy worth holding is that a promotion moves the version with the values untouched, and a configuration change moves the values with the version pinned. It costs an extra deploy and buys attribution. ## Ownership Finally, decide who owns which file. Chart defaults are usually the platform or chart author's, and changing one changes everybody; environment files are the operating team's, and changing one changes one place. That split is the reason to push a value into the chart when it is universal and the reason to resist pushing a value into the chart just because production wants it — a default is a decision made on behalf of every consumer, including the environments that have not been created yet.
- How would you measure divergence between two environments rather than argue about it?Compare what Helm actually used, not what the repository claims: `helm get values` for each release, or render both with `helm template -f` and diff the results. Report the count of keys present in production and absent from staging, and surface any increase in the pull request. A number that trends up is a conversation; an opinion about drift is not.
- When is an environment-name conditional inside a chart template acceptable?Almost never. It hides a behaviour difference in the template body, where a reviewer diffing three short values files will not see it, and it hard-codes your environment names into a chart that should be installable by anyone. Express the difference as a named value with a default; if a whole capability is production-only, gate it on an explicit feature key so the difference appears in the environment's file.
- When would you stop layering values files and give production its own chart?When the difference is topology rather than configuration — production runs different components, not larger ones — and the templates have become a maze of conditionals to express both. Before splitting, check whether a subchart toggle or a second release of the same chart covers it, because two charts means two things to review, version and promote, and the staging-ran-this guarantee is gone.
Sizing differences are like running the same play with a bigger squad; behaviour differences are running a play the reserves have never rehearsed and finding out on match day.
saying these in an interview costs you the question
- Says production is special so its file can be anything
- Branches on the environment name inside templates
- Keeps keys only in the environment that needs them
- Copies production's file to create a new environment
- Moves chart version and values in one production deploy
- Duplicates identical values across all three files