How do you decide how wide to make a shared Helm chart's values.yaml surface?
answer
- The two opposite failure modes
- Ask what the chart exists to make true
- A default is an invitation to override
- Serve the long tail generically
- Read the values stored with real releases
basics
~20 sExpose what installers legitimately vary, keep non-configurable what the chart exists to guarantee, and absorb the long tail with generic pass-through keys rather than a named key per request - because every key shipped is one strangers now depend on.
solid answer
~50 sStart from what the chart is *for*. A platform chart exists to make some things true - a security context, a set of labels, a scrape configuration - and those should be rendered unconditionally, not defaulted, because a default is an invitation to change it. Everything genuinely site-specific - replicas, image, scheduling, resource sizing, the application's own configuration - is variance the chart must expose or it will be forked. Between those two sits the long tail, and the answer there is one or two generic hatches rather than a named key per request: a named key is a promise you support forever, and its cost is paid on every future refactor of the templates behind it. Drive the decision with evidence rather than requests - the values actually stored with the existing releases tell you what installers really set. And where a hatch can defeat a guarantee, back the guarantee with a cluster-side control instead of a template.
go deeper
You are not expected to design a shared chart's API, but know the vocabulary: values.yaml is what installers configure, and the keys a chart exposes are a commitment its author has to keep working.
Be able to argue both sides concretely - what a missing key costs the consumer, what an extra key costs the maintainer - and to spot when a request is better served by an existing pass-through key than by a new one.
Show that you would look at what installs actually set before changing the surface, and that you distinguish a default from a guarantee. Be ready to say how you would stop a chart drifting into a hundred keys.
Own the strategy: which properties the platform enforces and where that enforcement really lives, how generic hatches and a cluster-side control divide the work, and what discipline keeps the surface from growing one reasonable request at a time.
### Two ways to get this wrong A chart's values surface fails in opposite directions, and both are expensive. **Too narrow.** A team needs a toleration you did not anticipate. They cannot express it, so they copy the chart into their own repository and add three lines. Your next fix never reaches them. Multiply by a handful of teams and the platform chart is now five slowly diverging charts, and you find out at the worst possible moment - during an incident, when the fix you shipped turns out not to be running anywhere. **Too wide.** Every request became a key. The chart is now a YAML re-encoding of the pod spec with no opinion left in it: hundreds of keys, every install unique, no two consumers exercising the same render path, and no refactor possible because any template change might move a key someone depends on. The chart has stopped being a package and become a very slow templating library. The judgement is where to sit between them, and it is a judgement about *purpose*, not about taste. ### The rule: guarantee, variance, tail Sort every candidate setting into one of three buckets. **Guarantees** are the reason the chart exists. If the platform's job is that every workload runs unprivileged with a read-only root filesystem and carries the labels the fleet's tooling depends on, those are rendered from the template unconditionally. Crucially, a guarantee must not be shipped as a *default*, because a default with a values key behind it is exactly as strong as the weakest team's willingness to override it, and the override happens at 2 a.m. during an incident and is never reverted. If you cannot bring yourself to make it non-configurable, admit that it is not a guarantee. **Variance** is what legitimately differs per install and per environment: replica count, image reference, resource requests, scheduling constraints, ingress hostnames, the application's own configuration. This is the chart's real API, and it should be small, conventionally named, and documented. **The tail** is everything else - the one-off annotation, the sidecar's volume, the extra environment variable one team needs. Serve it with generic pass-through keys. The whole point is that the tail costs you nothing to support: you did not name it, you do not interpret it, and next quarter's request lands in a key that already exists. ### Evidence beats requests Deciding from a queue of feature requests biases the surface toward whoever asks loudest. Deciding from data does not. The values a caller supplied are stored with each release, so `helm get values` across every install of the chart - nine installs of a monitoring-stack chart across three clusters, say - tells you which keys are actually set, which are set to something other than the default, and which have never been touched by anyone. A key nobody has ever set is a maintenance liability with no benefit. A default that every single install overrides is not a default; it is a wrong guess that should be changed, or a required input that should be declared as one. ### The hatch versus the guarantee The uncomfortable corollary of shipping a generic hatch is that it can be used to defeat the very things the chart is meant to enforce - a pass-through block can add a container, a volume, an annotation that changes how the fleet treats the workload. Templates are a poor place to police this: you would be writing a policy engine in Go template, badly, and it only ever inspects what arrives through *your* chart. The right shape is layered: the chart makes the good path easy and the fleet's admission policy engine makes the bad path impossible, applied to everything on the cluster regardless of whether it came from your chart, a hand-applied manifest, or someone else's. Where a chart-owned entry must survive a user's map, merge the user's entries beneath yours rather than trusting them to leave yours alone. ### Keeping the surface honest over time New keys arrive one plausible request at a time, which is how a surface becomes unmaintainable without anyone deciding it should. A workable discipline: a new key needs a second asker or a guarantee behind it; anything satisfiable by an existing hatch goes to the hatch; and the whole surface is reviewed against the collected release values periodically rather than never. The candidate who can articulate this is describing the difference between a chart the platform team owns and a chart the platform team is owned by.
- A team asks for a values key to turn off a security context your chart sets deliberately. What do you do?Treat the request as evidence about the guarantee, not as a key request. Find out what they are actually trying to do - it is usually one capability or one writable path, which can often be expressed narrowly without opening the whole setting. If the underlying need is legitimate and general, change the guarantee for everyone. If it is a genuine one-off exception, it belongs in a cluster-side exemption with an owner and an expiry, not in a values key every installer inherits.
- How do you find out which of your chart's keys anyone actually uses?Read the values recorded with the live releases rather than asking. Helm stores the caller-supplied values with each revision, so iterating over the installs of a chart shows exactly which paths are set and to what. Keys nobody sets are candidates for removal, keys everybody overrides mean the default is wrong, and keys set to wildly different things tell you where the real variance in your estate is.
- One of nine installs needs a setting nobody else wants. Do you add a key for it?Prefer the existing hatch. A named key for a single consumer buys you a permanent constraint on the template behind it in exchange for a convenience one team enjoys. If the setting cannot be expressed through a hatch - because the chart must branch on it, or it appears in several rendered objects - then it earns a key, and you write down why so the next reviewer does not have to re-derive the argument.
A chart's values file is a control panel, not a wiring diagram: the switches are what operators may change, the soldered joints are what the machine guarantees, and adding a switch for every request eventually leaves you with no machine at all.
saying these in an interview costs you the question
- Adds a key for every request until the chart mirrors the pod spec
- Ships a guarantee as an overridable default and calls it enforced
- Exposes nothing, so consumers quietly fork the chart
- Decides the surface from the loudest team rather than usage data
- Tries to police pass-through content inside templates
- Assumes a key nobody sets costs nothing to keep