How do you split responsibility between Helm releases and an operator's custom resources for a stateful platform?
answer
- sort responsibilities by what triggers them
- who acts when nobody is watching
- requests roll back, consequences do not
- one platform release, many tenant releases
basics
~20 sHelm owns what is decided at deploy time — packaging, version pinning, per-tenant configuration, release history. The operator owns everything triggered by cluster events — backup, failover, staged version migration. The test is what has to happen when nobody is at a keyboard.
solid answer
~50 sDraw the line by trigger, not by subject matter. Anything a person or pipeline asks for is chart-shaped: which controller version runs, how many ledger instances a tenant gets, what configuration each environment uses, and a recorded revision history you can roll back. Anything reality triggers is controller-shaped: promoting a replica when a primary dies, taking scheduled backups, walking a major version migration one member at a time. Helm still packages and versions the operator itself, and renders the custom resources that state each tenant's desired shape. What Helm cannot do at that boundary is undo consequences: `helm rollback` restores a previous revision's manifest as a new revision, so rolling a migration request back asks the operator to reconcile toward the old spec against data that has already moved. Design the release layout — one platform-owned operator release, many tenant releases — so those two kinds of change are never the same change.
code
bash · 8 lines# platform team, once per cluster
helm upgrade --install ledger-operator ops/ledger-operator -n platform-system
# product team, many times a week
helm upgrade --install payments ./ledger-umbrella -n payments -f envs/prod.yaml
# what a rollback here restores: the requested spec, not the data
helm rollback payments 41 -n paymentsgo deeper
Know the basic division: charts describe what you want deployed and record what was deployed; a controller is what reacts to failures afterwards. Do not expect a chart to perform a failover.
Explain which settings end up as custom resource fields rendered by a chart and which behaviours belong to the controller, and why the operator itself still arrives as a release.
Show the operating consequences: upgrade ordering between controller and tenants, uninstall blast radius, and that helm rollback restores a requested spec rather than migrated data.
Own the allocation across the estate — release layout, who reviews a migration-triggering change, and the honest cost of running two control planes with different triggers and failure modes.
### The question behind the question Interviewers asking this are not asking you to pick a winner. They want to see whether you can allocate responsibilities across two control planes with different failure semantics and then live with the seams. The organising principle is the trigger: * **Deploy-time decisions** — chosen by a human or a pipeline, reviewable in a diff, recorded as a revision. Which controller version runs. What each tenant's payments ledger should look like. What the staging values differ by. These belong to Helm, because Helm's whole value is turning a decision into a versioned, inspectable, rollback-able artefact. * **Runtime reactions** — triggered by events nobody scheduled. A primary is lost. A backup window arrives. A migration reaches its third member. These belong to a controller, because they must happen while nobody is running commands. Everything else follows. "Should backups be in the chart?" — the *schedule* is a deploy-time decision, so it is a field of a custom resource the chart renders; *taking* the backup is a runtime reaction, so it is the controller's. That single split answers most variants of the question cleanly. ### What Helm keeps owning, and it is a lot Adopting an operator does not shrink Helm's job to nothing: * **Packaging and version pinning of the operator itself.** The controller arrives as a chart, and that release is where its version, resources and watch scope are declared. * **Per-tenant desired state.** The custom resource for a ledger is a rendered manifest like any other, so charts still give you templating, values precedence, a values schema and per-environment differences. * **History and audit.** The release record answers "what did we ask for, when, and with which values" — for the desired state, not for what the controller then did. * **Ordering and grouping.** Which objects ship together and in what order is a chart-level concern. ### Where the seam actually hurts: rollback This is the point worth making unprompted. `helm rollback` reads a stored revision, applies its manifest, and writes it as a *new* revision. For a Deployment that is a genuine undo. For a custom resource that requested a major version migration of a payments ledger, it is not: the manifest goes back, the operator reads the older spec, and now has to reconcile a cluster whose data has already moved forward — which is at best refused and at worst destructive. Helm rolls back requests; it does not roll back consequences. So the design rule is: never let a revision that changes only configuration be the same revision that requests an irreversible operation. Keep migration-triggering fields in their own release or their own change, review them as a data operation rather than a deploy, and be explicit that recovery for them is the operator's restore path, not `helm rollback`. ### Release layout across a fleet A workable shape, and one you should be able to defend: * **One operator release per cluster**, platform-owned, cluster-scoped, upgraded on the platform team's cadence, with the CRDs it defines treated as a shared cluster asset that no tenant release may install. * **Many small tenant releases**, each holding the custom resources for one team's workloads plus the ordinary objects around them, owned and changed by that team. * **Upgrade order:** controller first, then tenants, because a newer controller is expected to understand older specs while the reverse is not guaranteed. Canary the controller on one non-critical tenant before the fleet. * **Blast radius:** because the two are separate releases, `helm uninstall` of a tenant cannot remove the controller, and a controller rollback cannot rewrite a tenant's spec. An umbrella chart that bundles an API subchart and a worker subchart is fine for a product team's own components; do not extend it to swallow the platform's operator, or one team's rollback becomes everyone's outage. ### The costs to name out loud A strong answer volunteers the price. You now run two control planes: one that changes things when asked, one that changes things on its own. Debugging starts with "which of them wrote this field", and that question has to be answerable — which is why field ownership under the default server-side apply path matters at the platform level rather than as trivia. You have taken on an operator's upgrade risk, its RBAC surface, and its bugs, for a component you probably did not write. And you have created a class of change — the migration request — whose review process is not the same as a deploy's, which people will not notice until the first one goes wrong. ### What decides it in the end Write down the day-two operations the payments ledger actually needs, and for each one ask who performs it at 03:00 with nobody watching. Every operation with no acceptable answer other than "a process already running in the cluster" is the operator's. The rest — the majority, in most estates — stays in charts, where a human decided it, a reviewer saw it, and a revision recorded it.
- A tenant's custom resource requested a major version migration and it went badly. What does `helm rollback` give you?The previous revision's manifest, applied as a new revision. The operator then sees the older requested version against data that has already migrated, which it will usually refuse or, worse, attempt. Helm rolls back the request, never the consequence. Recovery is the operator's restore path from a backup, and the deploy pipeline should treat migration-triggering changes as data operations with their own review and their own rollback plan.
- Would you ever put a tenant's custom resources in the same release as the operator?Only for a single-team, single-instance setup where the two genuinely never move independently. On a shared platform it is wrong: every controller bump re-renders tenant specs, an uninstall removes both, and one team's rollback rewinds the shared controller. The default is a platform-owned operator release per cluster and one release per tenant.
- How do you sequence upgrading the operator against the many releases that depend on it?Controller first, tenants after, on the assumption that a newer controller understands older specs while an older controller may not understand newer ones. Canary the controller on a low-risk tenant, confirm reconciliation is healthy, then let tenant releases adopt new spec fields at their own pace. Never bundle the controller upgrade into a tenant's deploy.
saying these in an interview costs you the question
- Treats it as chart versus operator instead of both
- Expects helm rollback to undo a data migration
- Puts failover logic in chart templates or hooks
- Bundles the shared operator into a team's umbrella chart
- Upgrades tenant specs before the controller that reads them
- Ignores the cost of running two control planes