When would you build a Kubernetes Operator for an application instead of shipping a Helm chart with a StatefulSet, and what does choosing the operator cost you?
answer
- Chart = render once; operator = loop forever
- Need app state to decide? operator
- Failover, PITR, resharding, ordered upgrades
- Costs: RBAC blast radius, availability, CRD upgrades
- Charts often install the operator — both, not either
basics
~20 sBuild an operator when the application needs ongoing, application-specific decisions — ordered upgrades, failover, backup and restore, resharding — that no template can make. If install-and-forget is enough, a chart plus StatefulSet is cheaper. The operator costs you a privileged, always-running component to write, secure, upgrade and monitor.
solid answer
~60 sA Helm chart is **templating at apply time**: it renders manifests once. An operator is a **control loop that keeps running**. So the deciding question is: *after install, are there decisions that require knowledge of the application's internal state?* Build an operator when the answer is yes — quorum-aware upgrades, promoting a replica on primary failure, consistent backups and restore, resharding, schema migrations, drift correction. Stay with a chart when the workload is essentially "run these Pods with this config", which the built-in controllers already handle. The costs are real: another Go service to write and test against a moving API; a Pod with broad RBAC over Secrets, PVCs and workloads, so a large blast radius; CRDs that are cluster-scoped and versioned, needing an upgrade story; a new failure mode where the operator is down or wedged; and an on-call surface that now includes the automation itself. Many teams do well with a chart plus a CronJob for backups until the pain justifies more. Note that the two also compose: charts commonly install operators.
code
yaml · 28 lines# Helm-rendered: correct at apply time, no ongoing decisions
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: orders-db
spec:
replicas: 3
serviceName: orders-db
template:
spec:
containers:
- name: postgres
image: postgres:16.3
---
# Operator-managed: intent plus day-2 policy the controller enforces
apiVersion: db.example.com/v1
kind: PostgresCluster
metadata:
name: orders-db
spec:
version: "16.3"
replicas: 3
failover:
automatic: true
maxReplicationLagSeconds: 1
backup:
schedule: "0 2 * * *"
target: s3://backups/orders-dbgo deeper
Say that a chart renders manifests once while an operator keeps running and reacting, and give one example task that needs the loop.
List the concrete triggers — failover, backup/restore, ordered upgrades, drift correction — and note that the built-in controllers already handle generic workload management.
Lead with the decision criteria and volunteer the costs: RBAC scope, operator availability, CRD upgrade path, added debugging layer; prefer adopting a mature operator over writing one.
Frame it as a platform investment decision — build only where domain semantics justify it, standardize on a small set of vetted operators, and define the support model and exit path before adoption.
## The core distinction A Helm chart is a **package manager and template engine**. `helm install` renders YAML from values and posts it to the API server; `helm upgrade` renders again and applies a diff. Between those commands nothing is watching. Reconciliation of what it produced is done by the built-in controllers — the Deployment controller keeps replicas up, the StatefulSet controller maintains ordinal identity and storage. An operator is a **process that never stops**. It watches its custom resources and the objects it created and acts continuously. That difference — one-shot render versus continuous control loop — is what should drive the decision, not fashion. ## Signals that you need an operator 1. **Ordering and state-aware upgrades.** "Upgrade replicas one at a time, wait for replication lag under 1s, then switch over the primary." A StatefulSet rolling update can restart Pods in order, but it cannot check replication lag or decide when a switchover is safe. 2. **Failover and leader promotion.** Deciding *which* replica becomes primary and repointing the Service requires reading the application's internal state. 3. **Backup, restore and point-in-time recovery.** Consistency requires application coordination (quiesce, snapshot, WAL archiving) and restore is a multi-step orchestration. 4. **Topology changes.** Resharding, rebalancing partitions, growing a quorum from 3 to 5 members safely. 5. **Continuous drift correction and self-service.** You want teams to `kubectl apply` a small custom resource and get a governed, validated, RBAC-scoped result — the operator becomes a platform API with defaults and policy baked in. 6. **Lifecycle actions triggered by events**, such as rotating credentials on a schedule or reacting to a node drain. ## Signals that a chart is enough - Stateless services, or stateful ones where the vendor's own upgrade path is "stop, replace image, start". - Configuration is fully expressible as YAML and nothing needs to react to runtime state. - The operational tasks you have can be met by a CronJob (backups) plus alerts and a human runbook, at acceptable frequency and risk. - The team lacks the capacity to own a Go service that follows Kubernetes API changes. A useful test: *list the manual runbook steps you perform per month, and their cost.* If they are frequent, urgent and mechanical, they are worth encoding. If they are rare and require judgment anyway, automation buys little and adds risk. ## The real costs of an operator **Engineering cost.** Controllers are deceptively hard: reconciliation must be idempotent, tolerate partial failure, handle conflicts, and cope with being restarted at any moment. Testing needs an API server (envtest) and end-to-end clusters. **Security cost.** The controller needs a ServiceAccount with write access to workloads, Secrets, PVCs and often across all namespaces. Compromising it is close to compromising the cluster. Scoping to specific namespaces and resources mitigates but rarely eliminates this. **Availability cost.** A wedged operator means no failovers and no backups, silently, unless you monitor the operator itself — reconcile errors, queue depth, last successful reconcile age. If the operator ships an admission webhook, an outage there can block API writes cluster-wide. **Upgrade cost.** CRDs are cluster-scoped: two teams cannot run conflicting versions of the same operator side by side without care. You need CRD versioning, a conversion story, and a plan for upgrading the controller while custom resources exist. **Cognitive cost.** Users now debug through an extra layer; "why is my database not ready" requires reading the custom resource's conditions and the operator's logs, not just Pod events. ## They are not mutually exclusive In practice most platforms use both: a Helm chart (or OLM bundle) installs the operator and its CRDs, and application teams then create custom resources. Charts also remain the right packaging for the many workloads that need no controller. A pragmatic path is to start with a chart, encode the two or three most painful runbook steps as Jobs or CronJobs, and only invest in an operator when those hit their ceiling — or adopt a mature upstream operator instead of writing one. ## Interview framing Say the distinction first (templating versus control loop), give one concrete task a template cannot do (state-aware failover or consistent restore), then volunteer the costs unprompted. Senior signal is refusing to build an operator by default and naming the conditions under which you would.
- Your chart-based deployment already uses a CronJob for backups. What would push you to an operator anyway?A CronJob can run a backup script but cannot decide *when it is safe*, cannot restore in an orchestrated way, and cannot fail over. Once I need consistent restore with a target time, automatic promotion of a replica, or upgrades gated on replication lag, the logic needs live application state and error handling that a scheduled script cannot express — that is the tipping point.
- How do you reduce the blast radius if you do run an operator?Scope it: run it namespaced with a Role rather than cluster-wide where the project supports it, restrict which resource types and which namespaces it may write, avoid granting blanket access to all Secrets, and run it with a restricted Pod security context. Then monitor it as a production service — reconcile error rate, queue depth, last successful reconcile — and know how to pause reconciliation safely during an incident.
- Can you use both Helm and an operator together?Yes, and that is the norm. Helm or an OLM bundle installs the operator's Deployment, RBAC and CRDs, while application teams create custom resources afterwards. Some operators are even implemented as Helm-backed operators, where the controller re-renders and applies a chart on every reconcile, which gives you drift correction without writing Go.
A Helm chart is a recipe printed once; an operator is a chef who stays in the kitchen tasting and adjusting. You only hire the chef if the dish keeps needing decisions.
saying these in an interview costs you the question
- Claiming an operator is 'Helm but better' — they solve different problems and are usually used together
- Building an operator for a stateless service where a Deployment already encodes everything
- Ignoring that the operator is a privileged, always-on component that itself must be monitored and upgraded
- Forgetting that CRDs are cluster-scoped, so operator versions collide across teams in a shared cluster
- Assuming a StatefulSet rolling update is equivalent to a state-aware upgrade with health gates