skip to content

When PostgreSQL runs in a Kubernetes StatefulSet with per-replica PersistentVolumeClaims, what does Kubernetes actually provide, and what is still left for you to build?

level: middleimportance: must knowfreq 64%

answer

  1. plumbing versus database semantics
  2. identity, claim, ordered rollout
  3. who is primary right now
  4. readiness is not role
  5. failover, backups, upgrades stay yours

basics

~20 s

A StatefulSet gives each PostgreSQL pod a stable name, its own volume and ordered rollout. It knows nothing about which pod is primary, replication, failover, backups or safe upgrades, so all of that is still yours.

solid answer

~40 s

Kubernetes gives the database **identity and storage plumbing**: a stable pod name and DNS entry per replica, a PersistentVolumeClaim per ordinal that follows the pod when it is rescheduled, and ordered, one-at-a-time updates. It gives **no database semantics**. The StatefulSet controller does not know which replica is the primary, does not configure streaming replication, cannot promote a standby, and never re-points the write Service after a failure. A readiness probe says a process answers, not that it is the writable primary or caught up. Backups, point-in-time recovery, major-version upgrades and a switchover before draining the primary's node all stay with you. That gap is why teams either add an operator such as CloudNativePG or use a managed database.

go deeper

for a junior

Remember the split: a StatefulSet gives stable names, a volume per replica and one-at-a-time updates, but it does not understand databases.

for a middle

Explain why a readiness probe and a Service are not enough to route writes, and list the missing pieces: role, replication, failover, backups and upgrades.

for a senior

Show you have seen the failure: writes landing on standbys, a crash-looping major upgrade, or a drain evicting the primary. Explain what an operator adds and what it costs.

for a principal

Treat the gap as an ownership question: who is on call for replication and recovery, and whether an operator or a managed service fits the team's skills and the platform's goals.

## The question behind the question When an interviewer asks this, they are testing whether you confuse **"Kubernetes can keep a pod with a disk running"** with **"Kubernetes can run a database"**. The first is true. The second needs a lot more, and knowing where the line sits is the start of any decision about running PostgreSQL, Kafka or any other data service in a cluster. Take a concrete case: a document-OCR pipeline on a 3-control-plane, 27-worker self-managed cluster. OCR workers (each requesting 0.35 CPU core) write extracted text into a 412 GiB PostgreSQL database. Someone proposes a three-replica StatefulSet with `volumeClaimTemplates`. What have they actually bought? ## What Kubernetes provides A **StatefulSet** is the workload controller for pods that need a durable identity. For a database it contributes: - **Stable identity**: pods are named `ocr-db-0`, `ocr-db-1`, `ocr-db-2`, and with a headless Service each gets a stable DNS name. The name survives a reschedule. - **Per-replica storage**: each ordinal gets its own **PersistentVolumeClaim** (PVC), and a replacement pod with the same ordinal mounts the same claim, so the data comes back with the identity. - **Controlled rollout**: by default pods are updated one at a time, so a bad image does not hit every replica at once. - **Restart on crash**: the kubelet restarts a failed container, and the controller recreates a deleted pod. These are real benefits. Running PostgreSQL in a Deployment with one shared claim would be much worse. ## What Kubernetes does not provide The StatefulSet controller treats all replicas the same. Everything that makes a set of PostgreSQL processes a *database cluster* is missing: | Concern | StatefulSet alone | Who must supply it | |---|---|---| | Which replica is primary | Unknown; all ordinals are equal | Your scripts or an operator | | Streaming replication setup | Not configured | Init logic or an operator | | Failover (promote a standby) | Never happens | An HA manager or operator | | Routing writes to the primary | A Service selects every matching pod | Role labels kept up to date | | Preventing two primaries | No fencing concept | An HA manager or operator | | Backups and point-in-time recovery | None | Backup tooling | | Major-version upgrade | Swapping the image does not migrate data files | A planned procedure | | Draining the primary's node | Evicts it like any pod | Switchover first | Two details trip people up: 1. **Readiness is not role.** A readiness probe that runs a query says the process accepts connections. It does not say the pod is the writable primary or that a standby is caught up. A Service built only on readiness sends writes to read-only standbys. 2. **Ordinal 0 is not "the primary".** After a failover the primary might be `ocr-db-2`. Anything that hard-codes `ocr-db-0` as the writer breaks at the first failover. ## What fills the gap There are three honest options: 1. **Build it yourself** by running an HA manager beside PostgreSQL and writing the glue: role labels, Service updates, backup jobs. It works, and it is a lot of code nobody else maintains. 2. **Use an operator** such as **CloudNativePG**. It manages pods and PVCs directly (it does not use a StatefulSet), labels the current primary, keeps `-rw`, `-ro` and `-r` Services pointing at the right pods, handles failover and switchover, and creates PodDisruptionBudgets for you. 3. **Use a managed database** outside the cluster and let the provider own failover, backups and patching. The pipeline connects to it like any external dependency. ## How to answer in an interview - Credit the StatefulSet with identity, per-replica storage and ordered updates. - Name the missing pieces explicitly: role, replication, failover, write routing, fencing, backups, upgrades. - Point out that readiness does not equal primary. - Say that the gap is why operators exist, and that a managed service is often the cheaper answer when one is available. The weak answer is "StatefulSets are for databases, so it's fine". The strong answer treats the StatefulSet as the storage and identity layer and asks who owns everything above it.

  • A three-replica PostgreSQL StatefulSet sits behind one ClusterIP Service selecting all its pods. What goes wrong for the OCR writers?
    The Service balances connections across all ready pods, so about two thirds of connections land on read-only standbys and every `INSERT` there fails. Writers need a Service that selects only the current primary, which means something must label the primary pod and move that label on failover. An operator such as CloudNativePG does this with a role label and a dedicated `-rw` Service.
  • Does updating the PostgreSQL image in the StatefulSet from one major version to the next upgrade the database?
    No. The rollout replaces containers one at a time, but the new major version cannot start on the old data directory without a migration step such as `pg_upgrade` or a dump and restore. A plain image swap usually leaves the new pod crash-looping. Major upgrades need a planned procedure, which an operator can automate and a StatefulSet cannot.
  • Is a StatefulSet with per-replica claims ever enough on its own for a data service?
    Yes, when the software handles clustering itself and treats every member the same, or when a single instance with restarts is acceptable. Examples are a development database or a cache you can rebuild. For a primary/standby database whose loss matters, the missing role, failover and backup logic has to come from somewhere.

A StatefulSet is a hotel that guarantees each guest the same room and luggage every night. It does not decide which guest runs the meeting, or who takes over when that guest leaves.

saying these in an interview costs you the question

  • A StatefulSet handles PostgreSQL failover automatically.
  • Ordinal 0 is always the primary, so point writes at it.
  • A passing readiness probe proves a pod is the writable primary.
  • Changing the image tag performs a major-version database upgrade.
  • Per-replica volumes mean backups are no longer needed.