How would you design the pre-merge checks for a Helm chart many teams install?
answer
- Combinations grow exponentially, promises do not
- The supported matrix is a published contract
- Cheapest checks first, credentials last
- Generic invariants outlive golden files
- Some failures only an install can reveal
basics
~20 sTier the checks by cost: lint and schema validation on every supported values file, then rendering plus targeted assertions on a small set of representative combinations, then a server-side dry run. Test the combinations you promise to support, not the exponential cross-product.
solid answer
~50 sStart from the contract: the set of values combinations the chart claims to support is the test matrix, so the honest way to shrink the matrix is to shrink the promise. Then tier by cost. `helm lint --strict` and values-schema validation run per values file and take milliseconds. Rendering runs next, with an explicit longest-supported release name, and produces artefacts you assert on — generic invariants first (every name and label value within 63 characters, the objects the chart promises are present), golden files only where a full diff genuinely helps. A server-side dry run comes last because it needs credentials. Everything that requires a running workload — readiness, in-cluster tests, upgrade from the previous chart version — belongs to a post-merge job against a throwaway namespace. The design questions are what counts as supported, who regenerates golden output, and which failures you accept discovering after merge.
code
bash · 11 linesCHART=./charts/fanout-operator
RELEASE=chat-fanout-prod-eu-central-1-shard-07
helm dependency build "$CHART"
mkdir -p renders
for f in ci/*-values.yaml; do
name=$(basename "$f" -values.yaml)
helm lint "$CHART" --strict --with-subcharts -f "$f"
helm template "$RELEASE" "$CHART" -f "$f" --namespace chat --include-crds > "renders/$name.yaml"
donego deeper
Focus on the individual checks first: know that a chart can be linted and rendered before merge, and that each supported values file needs its own run. The design conversation comes later.
Be ready to describe a concrete pipeline: lint per values file, render each one, assert something about the output, and say which defects each stage catches.
Show the ordering argument — cheap and unambiguous checks first, credential-bearing checks last — and name what a no-cluster suite cannot prove so the later stages are placed deliberately.
Own the tradeoff: the supported matrix is a contract you can shrink, golden output carries a maintenance cost someone must fund, and some failures are cheaper to find after merge. Say where you stop and why.
## Start from the promise, not the tooling A chart that many teams install has a published contract: these values keys, these combinations, this range of release names. That contract *is* the test matrix. Nine independent booleans is 512 renders, and no one maintains 512 expectations — but nobody promised 512 configurations either. The first design act is to write down the combinations you support (default, HA, metrics on, external database, ingress on) and treat everything else as unsupported until someone asks. Shrinking the matrix is a product decision expressed as a testing decision, and it is the lever a lead actually controls. ## Tier by cost, fail early **Tier 0 — free, per values file.** `helm lint --strict`, with `--with-subcharts` when the chart has dependencies, once per supported values file. The schema in `values.schema.json` is enforced as a side effect of both lint and render, so it costs nothing extra. Keep one deliberately invalid values file whose job is to *fail*, and assert that it does; a schema that silently stops rejecting bad input is a regression nobody notices otherwise. **Tier 1 — render and assert.** Render each supported combination with an explicit release name at the longest length Helm permits, and with `--include-crds` if the chart ships CustomResourceDefinitions, since they are otherwise left out of the output entirely. Then assert. Prefer **generic invariants** over per-resource expectations: no `metadata.name` or label value over 63 characters, every promised kind present, no empty documents. One invariant covers every object the chart will ever gain; a golden file covers only what someone remembered to add. **Tier 2 — golden renders, used sparingly.** A committed render of the default values is a real safety net: it makes an accidental change to any object visible in the diff. It also has a maintenance cost that teams consistently underestimate. Every intentional change regenerates it, large diffs get rubber-stamped, and after a few months the golden file is a formality. Keep them small, one per scenario, regenerated by a single documented command, and back them with the targeted assertions that carry the actual meaning. **Tier 3 — server-side dry run.** The first check that consults the API server, and the first that needs credentials. It validates the objects rather than the text: schema, unknown fields, name syntax, admission. Because it needs a cluster it is often the boundary between the pull-request job and the post-merge job. ## What a no-cluster suite structurally cannot prove Be explicit about this in an interview, because it is where the judgement shows. Nothing before Tier 3 can establish that an image exists, that a workload becomes ready, that a probe passes, that an upgrade from the previously published chart version applies cleanly, or that the chart's own in-cluster tests succeed. Those need an install into a throwaway namespace and belong to a job that runs after merge — or before, if your platform can afford ephemeral clusters on every pull request. Deciding which of those you can afford to discover late *is* the design. ## The parts people forget - **Dependencies.** If the chart has subcharts, the pre-merge render is only meaningful once dependencies are resolved; otherwise you are testing a chart that is missing half its templates. - **The upgrade path.** The riskiest change to a widely-installed chart is rarely a bad render; it is a change that renders beautifully and cannot be applied over what the previous version installed. That is a cluster-level check, but the pre-merge suite should at least flag the kinds of change that require it. - **Ownership.** Say who regenerates golden output and who owns the values-file directory. A suite with no owner drifts into a set of checks people re-run until they are green. ## How to choose, out loud The answer an interviewer wants is not a list of commands, it is a cost/benefit ordering with a stated stopping point: what you run on every pull request because it is free and unambiguous, what you run because the failure it catches is expensive enough to justify the maintenance, and what you deliberately leave to a later stage because catching it earlier would cost more than the failure does. A chart with five consumers and a chart with five hundred sit at different points on that curve, and saying so is the substance of the answer.
- Which checks would you deliberately move to after merge rather than before?Anything needing a running cluster and time: installing into a throwaway namespace, waiting for workloads, running the chart's in-cluster tests, and applying the change as an upgrade over the previously published version. They are the checks with real signal, and also the ones whose cost per pull request is highest. Move them post-merge unless the chart's blast radius makes a bad merge unacceptable.
- How do you stop committed golden renders from being rubber-stamped?Keep them per-scenario and small so a diff is readable, regenerate them with one documented command so nobody hand-edits, and put the meaning in targeted assertions rather than the diff. When a golden diff is the only evidence a change is safe, the review is a formality; when it is corroboration for an assertion that already failed loudly, it earns its place.
- A team asks you to support a values combination nobody tests. How do you answer?Either it joins the supported matrix — a values file, a lint run, a render and its assertions — or it stays documented as unsupported. The cost of support is a permanent line in the suite, not a one-off render, and being explicit about that is how the matrix stays finite. Charging that cost visibly is also how you find out whether the combination is actually needed.
saying these in an interview costs you the question
- Proposes rendering every possible values combination
- Treats a committed golden render as free to maintain
- Claims a no-cluster suite proves the workload works
- Runs the credential-bearing check before the free ones
- Leaves the supported values matrix undocumented
- Renders without resolving the chart's dependencies first