What do Kubernetes data-service operators such as CloudNativePG and Strimzi add on top of plain pods and volumes before a failover, drain or rolling restart is safe?
answer
- facts Kubernetes cannot see
- role label and -rw Service
- primary budget blocks the drain
- switchover when node is cordoned
- restart only if ISR allows
basics
~20 sThey add knowledge of the database's own state. CloudNativePG tracks and labels the primary, moves Services, switches over before a drain and creates PodDisruptionBudgets. Strimzi restarts brokers one at a time and only when in-sync replicas allow it.
solid answer
~50 sAn operator adds **data-aware checks** that the built-in controllers cannot make. **CloudNativePG** manages pods and PVCs directly rather than through a StatefulSet. It tracks which instance is primary, stores that in the `Cluster` status, labels pods with `cnpg.io/instanceRole`, and points the `<name>-rw` Service at the primary and `-ro` at the replicas. It creates a PodDisruptionBudget with `minAvailable: 1` for the primary, so a drain cannot evict the pod while it is primary. When it sees the primary's node cordoned, it switches over to a replica first. **Strimzi** runs brokers through its own `StrimziPodSet` resource. Its rolling restart asks Kafka whether restarting a broker would drop any topic below `min.insync.replicas`, and waits if it would. It also creates a PodDisruptionBudget, with `maxUnavailable` defaulting to 1. The value is the check itself. You still need to read what the operator does during a drain before you trust it.
code
yaml · 11 linesapiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: ocr-db
spec:
instances: 3
primaryUpdateStrategy: unsupervised
primaryUpdateMethod: switchover
enablePDB: true
storage:
size: 412Gigo deeper
Recall that an operator watches a custom resource and knows the database's roles, which the built-in controllers do not.
Explain how role labels drive the -rw and -ro Services, and how a PodDisruptionBudget interacts with the operator during a drain.
Name the concrete checks: switchover on a cordoned node, the primary budget, Strimzi's in-sync replica check before a restart. Say which actions bypass them.
Weigh the operator as a dependency: its upgrade cadence, CRD churn, who is on call for it, and whether its safety model matches your recovery objectives.
## Why plain Kubernetes controllers are not enough The built-in controllers decide things from **pod-level facts**: is the pod running, is it Ready, how many replicas exist. A replicated data service has facts Kubernetes cannot see: which member is the primary, which replicas are in sync, whether a restart right now would lose writes. An **operator** is a controller for a custom resource that reads those facts through the database's own interfaces and acts on them. This question is not about how to build an operator. It is about what two widely used ones actually do, and what that tells you about safety. The running example: a document-OCR pipeline on a 3-control-plane, 27-worker self-managed cluster, with PostgreSQL holding the extracted text and Kafka carrying the page events. ## CloudNativePG (PostgreSQL) CloudNativePG uses a `Cluster` custom resource (`spec.instances: 3`, for example). It does **not** run on a StatefulSet and does not use an external HA manager. It creates the pods and PVCs itself and talks to the Kubernetes API directly. What it adds: - **Role tracking**: the current and target primary are recorded in the `Cluster` status, and each pod gets a `cnpg.io/instanceRole` label. - **Role-aware Services**: `<name>-rw` selects the primary, `<name>-ro` selects replicas, `<name>-r` selects any instance. After a failover the labels move, so clients follow without changes. - **Failover and switchover**: when the primary fails it promotes the most suitable replica. `failoverDelay` and `switchoverDelay` tune the timing. - **Drain awareness**: if the primary's node is marked unschedulable (a `kubectl cordon`, which `kubectl drain` does first), the operator switches over to a replica on a schedulable node. - **PodDisruptionBudgets** (while `enablePDB` is true, the default): one for the primary with `minAvailable: 1`, and, from three instances up, one for the replicas that allows only one replica to be evicted at a time. The primary budget means a drain waits until the switchover has made that pod a replica. - **Rolling updates**: replicas are updated first. Then the primary is handled according to `primaryUpdateStrategy` (`unsupervised` by default, or `supervised` to wait for a human) and `primaryUpdateMethod` (`restart` by default, or `switchover`). - **Fencing**: an annotation, `cnpg.io/fencedInstances`, stops chosen instances on purpose. ## Strimzi (Kafka) Strimzi uses `Kafka` and `KafkaNodePool` custom resources. Brokers run under Strimzi's own `StrimziPodSet` resource, which replaced StatefulSets so the operator can manage each pod individually. What it adds: 1. **Availability-checked rolling restarts.** Before restarting a broker, the operator asks Kafka which partitions that broker holds and whether taking it down would push any topic below its `min.insync.replicas`. If it would, it waits and retries rather than restarting. 2. **Controller quorum checks** before restarting a KRaft controller node. 3. **A PodDisruptionBudget** per cluster, whose `maxUnavailable` defaults to 1 and can be changed through the resource's template. 4. **Manual restart hooks**: the `strimzi.io/manual-rolling-update` annotation asks the operator to perform the restart, so the operator's safety checks still run. ## What this buys, side by side | Safety question | Plain StatefulSet | CloudNativePG | Strimzi | |---|---|---|---| | Which member takes writes? | Unknown | Role label plus `-rw` Service | Kafka's own partition leaders | | Can this member stop now? | Pod count only | Primary PDB plus switchover | In-sync replica check | | Who promotes on failure? | Nobody | Operator | Kafka controller quorum | | Drain of a critical member | Evicted if the budget allows | Blocked until switchover | Blocked by the budget, restart checked | ## What an operator does not remove - **You still run it.** The operator is software with its own upgrades, CRD versions and bugs. When it is down, nothing reconciles. - **Its model has edges.** A PodDisruptionBudget still counts Ready pods, and some checks only run when the operator performs the action itself, not when something else deletes a pod. - **Storage is still storage.** Failover is only as good as the replicas' data, and a zonal volume still limits where a pod can move. - **Backups need configuring.** The operator can run them, but retention and restore testing are still your job. The senior answer names the specific checks, shows how they meet `kubectl drain`, and still treats the operator as a dependency to test before production.
- A CloudNativePG cluster has `spec.instances: 1`. What happens when someone runs `kubectl drain` on its node?There is no replica to switch over to, and the primary PodDisruptionBudget with `minAvailable: 1` blocks the eviction, so the drain waits. CloudNativePG's `nodeMaintenanceWindow` with `inProgress: true` and `reusePVC: true` removes that budget for a single instance, accepting downtime until the node returns. The alternative is `enablePDB: false`, which is meant for non-production clusters.
- Someone deletes a Strimzi broker pod directly with `kubectl delete pod`. Does the in-sync replica check protect the topic?No. The in-sync replica check runs when the operator decides to roll a broker. A direct delete bypasses both that check and the PodDisruptionBudget, because the budget only governs the Eviction API. The pod is recreated, but if another broker is already down, `acks=all` producers can fail meanwhile. Ask the operator for a restart with the `strimzi.io/manual-rolling-update` annotation instead.
- Why might CloudNativePG's default `primaryUpdateMethod` matter for the OCR pipeline's write availability during an upgrade?The default is `restart`, which restarts the primary in place after the replicas are updated, so writes stop until it comes back. With `switchover`, the operator promotes an already-updated replica first, so the write outage is only the switchover itself. For writers that retry, either may be fine. For a tight write window, `switchover` is usually the better choice.
saying these in an interview costs you the question
- CloudNativePG is just a StatefulSet with a nicer YAML front end.
- A PodDisruptionBudget alone knows which pod is the database primary.
- Strimzi restarts brokers on a fixed timer regardless of replica state.
- With an operator installed, deleting pods by hand is always safe.
- An operator removes the need to test backups and restores.
- Operators need no upgrades or monitoring of their own.