What goes into a Helm chart whose job is installing an operator, and what stays out?
answer
- the controller is just objects
- which lifecycle does each object share
- what an uninstall would take with it
- two releases, two owners, two cadences
basics
~20 sIt ships the controller's own runtime — its CustomResourceDefinitions, Deployment, ServiceAccount and RBAC — with values describing the controller: image, watched namespaces, resources. The custom resources describing real instances normally live in a separate release.
solid answer
~40 sAn operator is ordinary Kubernetes objects, so a chart can install it like anything else: the CustomResourceDefinitions that teach the API server the new kinds, a Deployment running the controller image, a ServiceAccount with the Role or ClusterRole and bindings it needs, and any webhook Service, ConfigMap or PodDisruptionBudget it wants. Its `values.yaml` is a description of the *controller* — image repository and tag, which namespaces to watch, replicas and leader election, resource requests, log level — not of the workload the controller will manage. The custom resources that ask for an actual payments ledger normally belong to a different release, owned by the team that wants the ledger. Keeping them apart means upgrading the controller never re-renders or deletes a tenant's instance, and uninstalling a tenant never removes the shared controller.
code
yaml · 13 lines# operator release values.yaml
image:
repository: registry.example.com/ledger-operator
tag: 1.9.3
watchNamespaces:
- payments
replicaCount: 2
leaderElection:
enabled: true
resources:
requests:
cpu: 100m
memory: 192Migo deeper
Know that an operator is installed like anything else — a Deployment, a service account and permissions, plus the definitions of its new kinds — and that its chart values configure the controller itself.
Explain why the custom resources that request an instance usually sit in a different release from the controller: independent upgrade, uninstall blast radius, and who owns each change.
Talk through the operating consequences — a controller restart pauses reaction but not the workload, stale appVersion hides which controller is running, and shared cluster-scoped definitions are contested by every install.
Own the release layout across a fleet: one platform-owned operator release per cluster against many tenant releases, and the permission and review boundary that split creates.
### The chart installs the controller, not the thing the controller manages An operator has no special packaging format. It is a controller image plus permissions plus the CRDs that define the kinds it watches, and all of those are ordinary Kubernetes objects — which is precisely why the overwhelmingly common way to get an operator into a cluster is `helm install` of a chart the operator's authors publish. "Charts versus operators" is a comparison of *models*, not of rival products; in practice charts are how operators arrive. Such a chart typically contains: * the **CustomResourceDefinitions** for the kinds the controller watches (shipped in the chart's `crds/` directory, which has its own lifecycle rules worth knowing separately); * a **Deployment** running the controller image, usually with leader election so two replicas do not both reconcile; * a **ServiceAccount** plus **Role/ClusterRole** and bindings — typically broad, since the controller creates workloads, Services and Secrets on the user's behalf; * optional plumbing: a Service and certificate Secret for an admission webhook, a ConfigMap of controller settings, a PodDisruptionBudget, metrics annotations; * a `values.yaml` whose keys are all controller-level: `image.repository`, `image.tag`, watched namespaces, `replicaCount`, resource requests, log level, node placement, whether to install the CRDs at all. Notice what is absent: nothing about the payments ledger. There is no `storageSize`, no `backupSchedule`, no `instances`. Those are fields of a *custom resource*, and a custom resource is a request made to the controller after it is running. ### Where the custom resources live Two releases is the normal answer, and the reasons are all lifecycle reasons. **Independent upgrade.** Bumping the controller from 1.9.2 to 1.9.3 should be a change to one release affecting one Deployment. If the tenant ledgers are in the same release, every controller bump re-renders their custom resources, and any template change sweeps them along with it. **Blast radius on uninstall.** `helm uninstall` deletes what the release owns. If the controller and the ledgers are one release, removing the operator removes the ledgers too — and worse, the controller may be torn down in the same operation, so nothing is left running to process the removal of the things it manages. Splitting them makes "remove the operator" and "remove a ledger" different, deliberate acts. **Rollback semantics.** `helm rollback` on the operator release should restore the previous controller image and nothing else. A tenant's ledger spec is not part of that decision and should not be rewound by it. **Ownership and permissions.** The operator release is usually the platform team's: cluster-scoped, one per cluster, installed once. The ledger releases are the product teams': namespaced, many, changed often. Different reviewers, different cadence, different blast radius. **Ordering.** A custom resource cannot be created before its CRD exists. Two releases make that ordering explicit — install the operator, then the workloads — rather than something you hope the chart got right on a fresh cluster. ### When a single release is acceptable A small, single-team chart that installs a modest controller alongside the one custom resource it manages is not automatically wrong; the coupling is only a problem if the two ever need to move independently. Go in with eyes open: an uninstall now tears down both together, and every controller upgrade touches the workload's spec. ### Two details that catch people out **`appVersion` versus `version`.** For an operator chart these diverge constantly. `version` is the chart's own packaging version; `appVersion` should track the controller image tag, so `helm list` tells an on-call engineer which controller is actually running. Charts that leave `appVersion` stale make the release history useless for exactly the question people ask during an incident. **Upgrading the operator release is not an outage of the workloads.** Replacing the controller Deployment stops reconciliation for a few seconds; the ledger's own pods keep serving, because they are not owned by the operator release and nothing deletes them. What you lose during that window is reaction — if a primary fails while the controller is restarting, failover waits. That is a very different risk profile from upgrading the workload itself, and stating the difference clearly is what separates a middle answer from a hand-wave.
- What would you check before letting an operator chart install its CustomResourceDefinitions in a shared cluster?Whether those kinds are already present and who put them there, because CRDs are cluster-scoped and shared by every namespace — two releases of the same operator chart in one cluster are fighting over one set of definitions. Most operator charts expose a value to skip installing them for exactly that reason, so a second installation can consume the definitions the first one owns.
- Should the chart that installs the operator also carry the operator's RBAC, or should that be separate?Keep it in the same release. The ServiceAccount, ClusterRole and bindings exist only to make that controller Deployment work; they share its lifecycle exactly, and separating them just creates a way to have one without the other. The split worth making is between the operator's runtime and the custom resources tenants create, not within the runtime itself.
- How does a values file for an operator chart differ from one for the workload it manages?The operator's values describe a controller process — image, watched namespaces, replicas, log level, resource requests. The workload's values describe a request — instance count, storage size, application version — which end up as fields of a custom resource. Mixing them into one file is the usual sign the two releases were merged when they should not have been.
saying these in an interview costs you the question
- Puts the managed instance's storage and backup settings in operator values
- Says operators replace charts, so no chart is involved
- Bundles tenant custom resources into the operator release
- Expects an operator upgrade to restart the managed workloads
- Leaves appVersion stale so the running controller version is unknown
- Assumes each release gets its own copy of the CRDs