Kubernetes operator projects are commonly rated against a five-level capability or maturity model. What do those levels describe, and how would you use them when evaluating a third-party operator before running it in production?
answer
- 1 install, 2 upgrade, 3 lifecycle, 4 insights, 5 autopilot
- Level 3 = backup/restore/failover — where value starts
- Level 4 = domain metrics + conditions, not Pod health
- Self-reported in the catalog — rehearse it yourself
- Higher level, broader RBAC and blast radius
basics
~20 sThe levels run: 1 basic install, 2 seamless upgrades of the managed application, 3 full lifecycle (backup, restore, failover), 4 deep insights (metrics, alerts, logs), 5 autopilot (auto-scaling, auto-tuning, auto-remediation). Use them to check the operator actually covers the operations you would otherwise do by hand.
solid answer
~50 sThe capability model describes **how much operational work the controller takes over**: 1. **Basic install** — provisions the application from a spec. 2. **Seamless upgrades** — upgrades the managed application version, and itself, without manual steps. 3. **Full lifecycle** — backup, restore, failover, scaling: the real day-2 runbook. 4. **Deep insights** — meaningful metrics, alerts, events and conditions about the managed application, not just about Pods. 5. **Autopilot** — automated scaling, tuning, remediation and configuration based on observed behaviour. When evaluating an operator I map the levels against my actual requirements: if I need point-in-time restore and it stops at level 1, I still own the hard part and should say so explicitly. I also check whether the claimed level is real — does it expose restore as a resource, does it publish Prometheus metrics and Conditions, does upgrading the operator upgrade the CRDs safely — and what RBAC it demands, because a level-5 operator usually wants broad, cluster-wide permissions.
code
bash · 5 lineskubectl get crds | grep example.com
kubectl explain postgrescluster.spec --recursive | grep -i -E 'backup|restore|upgrade'
kubectl get postgrescluster orders-db -o jsonpath='{.status.conditions}' | jq
kubectl delete pod orders-db-1 # rehearse failover
kubectl get clusterrole postgres-operator -o yaml # inspect demanded RBACgo deeper
Recall the five levels in order and that they describe how much operational work the controller performs.
Map the levels to concrete features you can check in the CRDs and docs, and note that levels are cumulative and self-declared in catalog metadata.
Turn it into an evaluation procedure: rehearse failover and restore, inspect metrics and conditions, review demanded RBAC, and record the gaps your team still owns.
Use the model to set a platform-wide bar for what may run in shared clusters, weighing automation benefit against permission scope, vendor lock-in and the ability to exit the operator.
## Why a maturity model exists "Operator" is a very broad label. A twenty-line controller that copies a ConfigMap and a mature database operator that performs point-in-time recovery are both operators. Teams adopting one need a common vocabulary for *how much operational responsibility it actually assumes*, because that determines what work remains on the humans. The five-level capability model (popularised by the Operator Framework and used across OperatorHub listings) provides that vocabulary. ## The levels **Level 1 — Basic install.** The operator can provision the application from a custom resource: create workloads, services, storage and initial configuration. Configuration changes may or may not be honoured after creation. This is roughly what a Helm chart already gives you, wrapped in an API. **Level 2 — Seamless upgrades.** The operator can upgrade the *managed application* version — sequencing restarts, running migrations, respecting ordering constraints — and can itself be upgraded without breaking existing custom resources. This level is where CRD versioning discipline starts to matter, because a newer controller must keep reading objects created by the older one. **Level 3 — Full lifecycle.** The classic day-2 set: backups on a schedule, restore from a backup, failover and failback, scaling out and in, credential rotation. This is the level at which an operator meaningfully reduces on-call load, and it is the level most enterprise buyers actually need. **Level 4 — Deep insights.** The operator exposes real observability *about the managed application*: Prometheus metrics (replication lag, backup age, quorum health), Kubernetes Events at meaningful transitions, and rich `status.conditions` so both humans and automation can tell healthy from degraded. Crucially this is domain telemetry, not just "the Pod is running". **Level 5 — Autopilot.** The operator acts on those insights: horizontal or vertical auto-scaling of the managed application, automatic tuning of configuration to the workload, automatic remediation of detected faults, and often automated capacity or cost decisions. Very few operators genuinely reach this level. The levels are cumulative in spirit — an honest level-4 operator does everything at levels 1 to 3. ## Using the model as an evaluation tool The model is most useful as a **gap-finding checklist**, not a scoreboard: 1. **Write down your own day-2 requirements first**: restore RPO/RTO, upgrade cadence, failover expectations, what you must alert on. 2. **Map each requirement to a level** and check the operator's documentation and CRDs for a corresponding capability. If restore is a real feature, there is usually a resource or command for it; if it is "documented as a manual procedure", the operator is level 1–2 regardless of what the marketing says. 3. **Verify, do not trust.** Install it in a test cluster and rehearse the operations you care about: kill the primary, restore yesterday's backup into a new cluster, upgrade the operator itself and confirm existing custom resources still reconcile. 4. **Check observability quality.** Are there metrics for backup age and replication lag? Do conditions flip to Degraded when something is actually wrong? Without that, higher levels are unverifiable in production. 5. **Check the cost side.** Higher capability usually means broader RBAC (cluster-wide access to Secrets and PVCs), more CRDs, sometimes admission webhooks that become a cluster-wide availability dependency. A level-5 operator with cluster-admin is a serious blast radius decision. 6. **Check the project's health.** Release cadence, CRD versioning practice, upgrade notes, whether it is a single-vendor project, and whether an escape hatch exists — can you take your data out if you stop running the operator? ## Where the tooling fits Operators are typically scaffolded with **Kubebuilder** or the **Operator SDK** (Go, and additionally Helm- or Ansible-backed variants), and distributed through **OLM** with a ClusterServiceVersion that declares the capability level shown in catalogs. That declaration is self-reported by the author, which is exactly why hands-on verification matters. ## Interview framing Name the five levels in order, then immediately pivot to judgment: the level is a claim, and your job is to test the two or three operations you would otherwise perform manually at 3am. Interviewers are looking for someone who treats an operator as a piece of production software they now own, not as a black box that came from a catalog.
- An operator claims level 3 but restore is documented as a manual procedure. How do you treat it?As level 1–2 for planning purposes. If restore is not expressed as an API object the controller executes, the operator is not carrying that responsibility and my team still owns the runbook, the rehearsal schedule and the RTO. I would document the gap explicitly and keep the manual procedure tested rather than assume coverage.
- What is the risk of adopting a level-5 autopilot operator?Automated remediation and tuning can amplify a fault: a wrong signal makes the controller scale, restart or reconfigure repeatedly, and it now has the permissions to do so cluster-wide. You need clear guardrails — bounded limits, rate limiting, good telemetry and a documented way to pause reconciliation — before letting software act without a human in the loop.
saying these in an interview costs you the question
- Treating the advertised capability level as verified fact rather than a self-reported claim
- Assuming higher level is always better, ignoring the wider RBAC and larger blast radius that comes with it
- Confusing level-2 application upgrades with upgrading the operator itself — both matter and they are different
- Describing level 4 as 'the Pods have liveness probes' instead of domain telemetry like replication lag or backup age
- Skipping a restore rehearsal because the operator lists backup as a feature