skip to content

Would you standardise 43 internal services on Helm charts or Kustomize overlays, and how would you decide?

level: principalimportance: should knowfreq 41%

answer

  1. The estate already runs one of them anyway
  2. Ask who installs it, not which is nicer
  3. Does anyone need one-command reversal?
  4. Values differences or structural differences?
  5. Two toolchains is the recurring tax

basics

~20 s

Decide on who consumes the output and what lifecycle you need. Charts earn their cost when a versioned, parameterised artifact ships to consumers you cannot redeploy, or when rollback, ordered hooks and uninstall matter. Otherwise overlays keep first-party YAML readable.

solid answer

~50 s

I would not answer with a preference; I would answer with two axes. First, **audience**: a chart is a versioned, parameterised package with a public interface — `values.yaml`, optionally constrained by `values.schema.json`, distributed as a `.tgz` through a chart repository or an `oci://` reference. That interface is worth its cost when consumers exist whom you cannot redeploy yourself; for an order-checkout API deployed only by its own team into six environments that differ in a handful of fields, it is overhead. Second, **lifecycle**: if you need release history, one-command rollback, ordered install hooks or a real uninstall, Helm supplies them and a plain apply does not. I would also note the estate already runs Helm as a consumer of third-party charts, so the real question is only what first-party services use. I would set one default, allow a documented exception, and forbid quietly maintaining both for the same service.

code

json · 16 lines
json
{
  "$schema": "https://json-schema.org/draft-07/schema#",
  "type": "object",
  "required": ["image", "replicaCount"],
  "properties": {
    "replicaCount": { "type": "integer", "minimum": 2, "maximum": 24 },
    "image": {
      "type": "object",
      "required": ["repository", "tag"],
      "properties": {
        "repository": { "type": "string" },
        "tag": { "type": "string", "pattern": "^[0-9]+\\.[0-9]+\\.[0-9]+$" }
      }
    }
  }
}

go deeper

for a junior

You are unlikely to own this call, but know the two words that decide it: distribution and lifecycle. Charts are versioned packages other people install; overlays are your own YAML, patched.

for a middle

Be able to translate a requirement into the right model — an external consumer or a needed rollback points at Helm; six environments differing in a few fields, and patching manifests you did not write, points at overlays.

for a senior

Show you would gather evidence before deciding: who installs each service, what recovery actually looks like today, whether variation is values-shaped or structural, and how many vendor charts the team already operates.

for a principal

Own the organisational cost and the migration. Set one default with written exception criteria, forbid maintaining both for the same service, and be explicit that ownership metadata makes moving in either direction a project rather than a flag.

### Reject the framing first The question is usually posed as a tooling preference, and the strongest answer refuses that framing. Almost every Kubernetes estate already runs Helm whether or not it writes charts, because vendors ship charts — an ingress controller, a database, a metrics stack. So the decision is never "do we use Helm"; it is "what do our own 43 services use, and what do we pay for that". Say that out loud, then decide on axes rather than taste. ### Axis one: who consumes the artifact A chart is a **distribution format**. `helm package` yields a versioned tarball with a declared interface: `values.yaml` as the knob surface, `values.schema.json` to constrain it, `Chart.yaml` carrying `version` and `appVersion`, and a `.prov` file if you sign it. It is served from a classic repository with an `index.yaml` or pushed to an OCI registry and installed with an `oci://` reference. All of that machinery exists so that somebody you cannot reach can install version 2.4.1, read the documented values, and upgrade on their own schedule. If that somebody exists — another business unit, a customer running your software in their cluster, an installer you ship — a chart is the right answer and there is little to debate. If the only consumer is the team that wrote the service, and the environments differ in replica counts, a few resource limits, a hostname and an image tag, then you are paying for a public interface with no public. Worse, the values surface becomes a contract you must keep working for yourself: rename a key and every environment breaks at once. ### Axis two: what lifecycle you need Helm brings state. Each install and upgrade is a revision recorded in a Secret in the namespace, which gives `helm list`, `helm history`, `helm get manifest`, `helm rollback`, deletion of resources dropped from the chart, and `helm uninstall` that knows exactly what belongs to the release. It brings ordering too: hooks with `helm.sh/hook-weight` run a migration before the workload rolls, and `helm test` exercises a real release afterwards. A flat apply of overlay output brings none of that and does not pretend to. If your services genuinely need pre-upgrade migrations sequenced against the rollout, or if your incident procedure is "put the last known good back in one command", Helm is doing real work. If deployment is a pipeline that re-applies from a commit and recovery is reverting the commit, you may prefer having exactly one source of truth rather than two that can disagree. ### Axis three: what kind of variation you have Be concrete about the differences between your environments. If they are values — replicas, limits, a hostname, an image tag — both models handle them and it is a wash. The tie-breaker is structural variation. If production needs objects that staging does not have at all, generation with `if` handles it cleanly and patching has to reach for other means. If, instead, what you need is to change fields on manifests you did not author, patching wins outright, because a chart can only expose knobs its author thought of. ### Axis four: what the organisation can carry This is where a principal answer separates itself. Two toolchains for one estate means two ways to render, two ways to diff, two review habits, two on-call runbooks. That is a real, recurring tax, and it is usually larger than the difference between the tools. So: pick one default, write down the exception criteria ("we publish a chart when the artifact leaves this organisation, or when the service needs ordered install hooks"), and require a decision record for exceptions. Forbid the worst outcome — a service maintained as both a chart and an overlay set, where nobody knows which one produced the running pods. Factor migration in honestly. Moving 43 services is not free in either direction, and moving away from Helm has a specific sharp edge: objects created by a release carry ownership metadata, and moving them under a different tool without cleaning that up leaves manifests claiming a release that no longer exists. Moving toward Helm has the mirror problem — Helm refuses to adopt resources lacking that metadata unless told to take ownership. ### The shape of a good answer A defensible recommendation for 43 first-party services with no external consumers is: overlays as the default for application services, Helm charts for the handful that ship outside the team or need ordered hooks, Helm as the consumer of every third-party chart regardless, and a post-renderer where a vendor chart must be patched — so the release record survives. A defensible answer in the other direction is: charts everywhere, because one lifecycle model across first-party and third-party workloads is worth more than the templating cost, and a schema on the values file keeps the interface honest. Both are correct. What is not correct is an answer that names a favourite tool, ignores who consumes the artifact, ignores whether anybody needs rollback, and never mentions the cost of running both.

  • Your platform standardises on overlays, but a vendor chart needs a field it never exposed. What do you do?
    Install it with Helm and patch the render with a post-renderer, so the release record, hooks and rollback survive the customisation. Rendering the chart out and applying it by hand solves the same field and quietly transfers the whole lifecycle to us. I would also open an upstream request for the value, because the post-renderer is a patch we maintain forever otherwise.
  • What is the strongest argument against making every internal service a chart?
    You take on a public interface nobody outside the team consumes. Every environment difference has to be anticipated as a value, the values file becomes a contract you must not break, and the source stops being readable Kubernetes YAML. For a service its own team deploys, that cost buys mostly the lifecycle features — so the honest question is whether you need those.
  • How would you stop the exception policy from becoming 'everyone picks their own'?
    Make the default the path of least resistance — a scaffold, a pipeline template and a runbook for it — and make exceptions require a short written record naming which criterion is met. Then measure: if most services claim the exception, the default was wrong and should change, rather than being enforced harder.
  • Does adopting a GitOps agent change this decision?
    Less than people expect. The agent decides who applies and how often; it does not decide whether your source is templated or patched, since agents can drive either. The relevant question is where the lifecycle lives, and that is the same argument I made without an agent in the picture.

saying these in an interview costs you the question

  • Picks a favourite tool without asking who consumes the output
  • Ignores that vendor charts already put Helm in the estate
  • Assumes templating is the only difference between the two
  • Ignores the ongoing cost of maintaining two toolchains
  • Treats migration between the two as a free rewrite
  • Argues from popularity rather than lifecycle requirements

context