A team wants automated backups, version upgrades and failover for a stateful database running on Kubernetes, all driven through the Kubernetes API. Explain what the Kubernetes Operator pattern is and how it delivers that.
answer
- CRD + controller = operator
- Encodes day-2 ops, not just install
- Declarative spec, controller closes the gap
- status field reports reality
- Built-in controllers are the same pattern
basics
~20 sAn Operator is a custom Kubernetes resource type plus a controller running in the cluster. Users declare the desired database in a custom object; the controller continuously drives the real system to match that spec, automating install, backup, upgrade and failover.
solid answer
~50 sAn Operator packages operational knowledge as software. It has two halves: 1. A **CustomResourceDefinition** that adds a new kind (say `PostgresCluster`) to the Kubernetes API, so users `kubectl apply` a spec: version, replicas, storage size, backup schedule. 2. A **controller** running in the cluster that watches those objects and reconciles reality toward the declared spec — creating StatefulSets, Services and Secrets, taking scheduled backups, doing rolling version upgrades, promoting a replica when the primary dies — and writes the outcome back into the object's `status`. The payoff is that **day-2 operations** (what a human SRE would do after install) become part of the same declarative API as everything else: same `kubectl`, same RBAC, same GitOps flow, same audit trail, same drift correction. Kubernetes' own Deployment and StatefulSet controllers work exactly this way; an Operator applies the pattern to domain-specific software that the built-in controllers know nothing about.
code
yaml · 18 linesapiVersion: db.example.com/v1
kind: PostgresCluster
metadata:
name: orders-db
namespace: payments
spec:
version: "16.3"
replicas: 3
storage:
size: 200Gi
class: fast-ssd
backup:
schedule: "0 2 * * *"
retention: 14d
status:
phase: Running
primary: orders-db-1
lastBackupTime: "2026-08-12T02:00:11Z"go deeper
Be able to state the two ingredients (custom resource + controller) and give one day-2 example such as scheduled backups or a version upgrade.
Explain that the controller creates and owns standard resources, reports through status and conditions, and needs RBAC via its ServiceAccount; name Kubebuilder or the Operator SDK as the usual tooling.
Frame it as operational knowledge turned into a control loop, and be honest about the costs: another privileged workload to run, upgrade and monitor, and a failure mode where the operator itself is down.
Discuss operators as a platform API strategy — turning services into declarative, RBAC-governed, GitOps-managed resources — and when buying a mature operator beats building one.
## The problem Kubernetes handles stateless workloads well: you declare "5 replicas of this image" and a built-in controller keeps five Pods alive. Clustered, stateful software is harder. Installing PostgreSQL, Kafka or Elasticsearch is the easy part; the hard part is everything after install — *day-2 operations*: taking consistent backups, restoring them, upgrading the engine version without data loss, resharding, rotating credentials, promoting a replica when the primary dies, scaling storage. Traditionally that knowledge lives in runbooks and in the heads of a few engineers, executed by hand or by scripts outside the cluster. ## The pattern An **Operator** is an application-specific controller that extends the Kubernetes API to manage a piece of software on the user's behalf. It is the union of two things: **1. Custom resources.** A CustomResourceDefinition registers a new object kind with the API server, so the cluster now serves `/apis/db.example.com/v1/postgresclusters`. Users write YAML like any other Kubernetes object. This is the *declarative interface*: what the user wants, not how to get it. **2. A controller.** A normal Pod (usually a Deployment) running in the cluster, holding credentials via a ServiceAccount and RBAC. It watches those custom objects and the resources it creates from them, and repeatedly compares desired state to observed state, issuing the API calls and domain actions needed to close the gap. It reports progress and health back on the object's `status` field, so `kubectl get postgrescluster` shows real information. ## Why this is better than scripts - **Declarative and self-healing.** If someone deletes the Service the operator created, the operator recreates it. A one-shot install script cannot notice. - **One API, one auth model.** Access to "create a database" is now an RBAC rule on a resource. Requests are audited by the API server. GitOps tools can manage databases exactly as they manage Deployments. - **Knowledge is executable.** "To upgrade: drain replicas one at a time, check replication lag, then fail over" stops being a wiki page and becomes code that runs the same way at 3am. - **Composable.** The operator itself is just a workload; it can be installed, upgraded and monitored like anything else. ## What the controller actually does for a database Typical responsibilities for a mature database operator: - **Provision**: create StatefulSet, headless Service, PVCs, generate and store credentials in a Secret, run initialization. - **Backup/restore**: run scheduled backup Jobs to object storage; expose restore as another custom resource (`PostgresRestore`) that points at a backup. - **Upgrade**: sequence a version change — pre-flight checks, upgrade replicas, switchover, upgrade the old primary. - **Failover**: watch health, promote a healthy replica, repoint the Service, update `status`. - **Scale**: add replicas, grow volumes, rebalance. ## Where the pattern is *not* worth it A stateless HTTP service needs no operator — a Deployment already encodes everything Kubernetes needs to know. Operators earn their cost where an application has real domain semantics that Kubernetes cannot infer: quorum, leader election, data placement, schema migrations, or ordered upgrade rules. Deploying complexity you do not need means you now maintain a controller *and* the database. ## Ecosystem Most operators are written in Go with **Kubebuilder** or the **Operator SDK** (which also supports Helm- and Ansible-backed operators), and many are distributed through **OLM** and catalogs like OperatorHub. Community maturity is often described using a five-level capability model, from basic install up to full autopilot. When evaluating a third-party operator, look at that maturity, at how it scopes permissions, and at whether its status reporting is good enough to alert on. ## Framing for an interview The crisp definition to say out loud: *"An operator is a CRD plus a controller that encodes day-2 operations — it turns the runbook for a specific application into a control loop, so operating it becomes declarative like the rest of Kubernetes."* Then give one concrete day-2 example (backup or version upgrade) to prove you mean operations, not just installation.
- Which parts of a system are the operator's responsibility versus the built-in Kubernetes controllers'?The operator owns application-specific semantics: quorum, leader promotion, backup consistency, upgrade ordering, credential lifecycle. It does not re-implement Pod scheduling, restarts or rollout of a plain Deployment — it creates StatefulSets/Deployments/Services and lets the built-in controllers do the generic work. A good operator is a thin layer of domain knowledge on top of standard primitives.
- Where do the operator's own permissions come from, and why does that matter?The controller runs as a Pod with a ServiceAccount bound by Role or ClusterRole. It typically needs broad rights over the resources it manages — StatefulSets, Services, Secrets, PVCs — often cluster-wide. That makes the operator a high-value target and a large blast radius, so you scope it to specific namespaces and specific resource types wherever possible rather than granting cluster-admin.
- How does a user find out whether the operator succeeded?Through the custom object's `status`, which the controller writes: a phase, standard `conditions` such as Ready or Degraded, observed generation, and domain fields like current primary or last backup time. Well-behaved operators also emit Kubernetes Events and metrics so failures can be alerted on rather than discovered by reading YAML.
It is like hiring the database's own on-call engineer and shrinking them into a Pod: you write down what the cluster should look like, and they keep watching and fixing it around the clock.
saying these in an interview costs you the question
- Saying an operator is just a Helm chart or an install script — it is a continuously running control loop, not a one-shot templating step
- Claiming an operator replaces Deployments and StatefulSets rather than creating and managing them
- Thinking every application needs an operator, including stateless HTTP services
- Believing the operator runs outside the cluster like a CI job instead of as a Pod using the Kubernetes API
- Forgetting that state and progress must be reported in the custom resource's status, so users have no visibility