skip to content

How would you plan a Helm adoption of hand-created objects across many teams?

level: principalimportance: nice to knowfreq 20%

answer

  1. The metadata is the migration primitive
  2. Boundaries first, commands second
  3. Render and diff before claiming
  4. Per-object stamping gives you a burn-down
  5. Decide the uninstall contract in advance

basics

~10 s

Treat the ownership metadata as the migration primitive: inventory what exists, decide release boundaries, make each chart render what is already live, stamp objects per release, and settle uninstall behaviour before anything is claimed.

solid answer

~50 s

The mechanics are easy — a label and two annotations — so the hard parts are boundaries and sequencing. First decide what a release is: usually one per service per namespace, with shared and cluster-scoped resources kept out of application charts entirely, because those are where cross-team collisions come from. Then, per service, get the chart rendering something close to what is already live so adoption is not an unreviewed change. Stamp ownership per object rather than deploying with the blanket claim flag, so the check keeps guarding everything not yet migrated and the burn-down is measurable: objects carrying the metadata are migrated, the rest are not. Before claiming anything stateful, agree what `helm uninstall` should do to it, and mark the keepers in the chart. Finally, ban the blanket flag in pipelines once the migration is done.

go deeper

for a junior

You are unlikely to lead this, but know that bringing existing resources under a chart is a metadata change plus an upgrade, and that it changes what a later uninstall will delete.

for a middle

Be able to script the per-object part reliably and to explain why the blanket claim flag is not the shortcut it looks like on a cluster with more than one owner.

for a senior

Show that you sequence it: inventory, close the render gap, stamp, verify, and settle the fate of stateful objects first. Be specific about how you would enumerate collisions rather than discovering them at apply time.

for a principal

Own the boundaries and the guardrails: what a release is allowed to own, which resources stay outside Helm on purpose, how progress is measured, and what stops the safety check from being permanently disabled once the migration ends.

## Why this is a design question, not a command The technical act of adoption is trivial: satisfy Helm's ownership check by putting `app.kubernetes.io/managed-by: Helm` plus `meta.helm.sh/release-name` and `meta.helm.sh/release-namespace` on a live object, and the next upgrade takes it under management. What makes a fleet migration hard is everything around that: deciding which release should own each object, doing it without a change of behaviour, and not handing a future engineer a deletion they did not sign up for. ## Step one: inventory and boundaries Start from what exists and who touches it. Cluster-scoped objects and anything shared between teams — ClusterRoles, admission configuration, shared ConfigMaps, custom resource definitions — are the recurring source of ownership collisions, because only one release can hold the annotations. The right answer for those is usually that no application chart owns them; they belong to a platform release, or to nothing Helm-managed at all. Getting that decision made first removes most of the errors the migration would otherwise generate. For the rest, the boundary that works is one release per service per namespace, with the release name and namespace fixed by convention. Ownership metadata is written in terms of that pair, so a convention decided late means re-stamping. ## Step two: make the chart match reality first Adoption is not a no-op. Once claimed, an object is reconciled toward whatever the chart renders, so a chart that renders slightly different resource limits or a different Ingress path performs a live change at the moment of adoption. The discipline that keeps this boring is: render, compare against the live object, close the gaps in values, and only then claim. Where the gap cannot be closed — a setting a person applied that the chart has no field for — that is a chart change with its own review, not something to discover during a migration window. ## Step three: stamp, do not blanket-claim There is a flag that skips the ownership check for a whole install or upgrade. It is the wrong default for a migration programme, for two reasons. It is indiscriminate: it will claim objects another release owns, silently, which is exactly the failure a multi-team migration is most exposed to. And it destroys your progress signal — with the check disabled you cannot tell adopted objects from ones that were merely swept up. Stamping per object gives you both safety and a burn-down: the fraction of objects carrying ownership metadata naming their intended release is the migration metric, and every un-stamped collision still fails loudly. ## Step four: settle the uninstall contract This is the part teams skip. Adopted objects join the release's stored manifest, so removing the release removes them. For stateless objects that is what everybody wants. For a PersistentVolumeClaim, a Service other systems resolve by name, or an Ingress holding a provisioned address, it is a latent outage. Each of those needs an explicit decision — kept by the chart's resource policy, or deliberately left outside the release — recorded before the object is claimed, because after adoption the decision is invisible until someone runs uninstall. ## Step five: guardrails after the migration When the programme finishes, the blanket claim flag should not survive in any pipeline. Its presence means a release will absorb any object that ever collides with it, with no log line, which is the same failure mode you spent the migration avoiding. A CI check that rejects the flag in deploy commands is cheap and permanent. ## The tradeoffs to name out loud - **Speed versus safety.** The blanket flag migrates a namespace in an afternoon; per-object stamping takes a sprint and is reversible. On a fleet with multiple owners, reversibility wins. - **Rewrite versus adopt.** Some resources are better deleted and recreated by the chart than adopted — anything cheap and stateless. Adoption earns its cost only where identity matters: allocated addresses, storage, external references. - **Coverage versus blast radius.** Pulling every object into a chart maximises reproducibility and simultaneously maximises what one uninstall can destroy. Deciding that some resources stay outside Helm on purpose is a legitimate answer, not a failure to finish. - **Who runs it.** A platform team can stamp metadata across namespaces far faster than each service team can, but the service team is the one who knows which hand-applied settings still matter. The realistic split is central tooling and enumeration, local review and sign-off.

  • How would you measure progress on a migration like this?
    Count objects, not services: for each namespace, how many live objects carry ownership metadata naming their intended release versus how many the corresponding chart renders. That number moves monotonically, is queryable from the cluster, and — unlike a checklist of migrated services — it exposes the objects everyone quietly skipped.
  • Which resources would you deliberately leave outside Helm ownership?
    Cluster-scoped objects shared by several teams, since only one release can hold the annotations; anything whose deletion is unrecoverable and whose lifecycle is genuinely longer than the release, such as storage claims holding production data; and resources another controller writes, where a chart claiming them just starts a fight.
  • A team wants to adopt a Service whose external address must not change. What do you require before they do?
    That the chart renders the Service with the same identity and no field that would trigger recreation, that the diff against the live object is reviewed, and that the chart marks it with the keep resource policy so a later uninstall cannot take the address away. Then stamp that object specifically rather than deploying with a blanket claim.

saying these in an interview costs you the question

  • Runs the blanket claim flag across a whole estate to finish faster
  • Adopts objects before the chart renders what is already live
  • Never decides what uninstall should do to claimed resources
  • Lets application charts own shared cluster-scoped objects
  • Treats a migrated service list as the metric instead of objects
  • Leaves the claim flag in the deploy pipeline afterwards

context