skip to content

Databases & Data Services

Whether Postgres or Kafka belongs in a cluster at all: what a StatefulSet with per-replica PVCs really buys, what an operator adds before failover is safe, and where a node drain meets quorum. The honest answer is often a managed service.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

5

When PostgreSQL runs in a Kubernetes StatefulSet with per-replica PersistentVolumeClaims, what does Kubernetes actually provide, and what is still left for you to build?

level: middleimportance: must knowfreq 64%

answer

  1. plumbing versus database semantics
  2. identity, claim, ordered rollout
  3. who is primary right now
  4. readiness is not role
  5. failover, backups, upgrades stay yours

basics

~20 s

A StatefulSet gives each PostgreSQL pod a stable name, its own volume and ordered rollout. It knows nothing about which pod is primary, replication, failover, backups or safe upgrades, so all of that is still yours.

solid answer

~40 s

Kubernetes gives the database **identity and storage plumbing**: a stable pod name and DNS entry per replica, a PersistentVolumeClaim per ordinal that follows the pod when it is rescheduled, and ordered, one-at-a-time updates. It gives **no database semantics**. The StatefulSet controller does not know which replica is the primary, does not configure streaming replication, cannot promote a standby, and never re-points the write Service after a failure. A readiness probe says a process answers, not that it is the writable primary or caught up. Backups, point-in-time recovery, major-version upgrades and a switchover before draining the primary's node all stay with you. That gap is why teams either add an operator such as CloudNativePG or use a managed database.

go deeper

for a junior

Remember the split: a StatefulSet gives stable names, a volume per replica and one-at-a-time updates, but it does not understand databases.

for a middle

Explain why a readiness probe and a Service are not enough to route writes, and list the missing pieces: role, replication, failover, backups and upgrades.

for a senior

Show you have seen the failure: writes landing on standbys, a crash-looping major upgrade, or a drain evicting the primary. Explain what an operator adds and what it costs.

for a principal

Treat the gap as an ownership question: who is on call for replication and recovery, and whether an operator or a managed service fits the team's skills and the platform's goals.

## The question behind the question When an interviewer asks this, they are testing whether you confuse **"Kubernetes can keep a pod with a disk running"** with **"Kubernetes can run a database"**. The first is true. The second needs a lot more, and knowing where the line sits is the start of any decision about running PostgreSQL, Kafka or any other data service in a cluster. Take a concrete case: a document-OCR pipeline on a 3-control-plane, 27-worker self-managed cluster. OCR workers (each requesting 0.35 CPU core) write extracted text into a 412 GiB PostgreSQL database. Someone proposes a three-replica StatefulSet with `volumeClaimTemplates`. What have they actually bought? ## What Kubernetes provides A **StatefulSet** is the workload controller for pods that need a durable identity. For a database it contributes: - **Stable identity**: pods are named `ocr-db-0`, `ocr-db-1`, `ocr-db-2`, and with a headless Service each gets a stable DNS name. The name survives a reschedule. - **Per-replica storage**: each ordinal gets its own **PersistentVolumeClaim** (PVC), and a replacement pod with the same ordinal mounts the same claim, so the data comes back with the identity. - **Controlled rollout**: by default pods are updated one at a time, so a bad image does not hit every replica at once. - **Restart on crash**: the kubelet restarts a failed container, and the controller recreates a deleted pod. These are real benefits. Running PostgreSQL in a Deployment with one shared claim would be much worse. ## What Kubernetes does not provide The StatefulSet controller treats all replicas the same. Everything that makes a set of PostgreSQL processes a *database cluster* is missing: | Concern | StatefulSet alone | Who must supply it | |---|---|---| | Which replica is primary | Unknown; all ordinals are equal | Your scripts or an operator | | Streaming replication setup | Not configured | Init logic or an operator | | Failover (promote a standby) | Never happens | An HA manager or operator | | Routing writes to the primary | A Service selects every matching pod | Role labels kept up to date | | Preventing two primaries | No fencing concept | An HA manager or operator | | Backups and point-in-time recovery | None | Backup tooling | | Major-version upgrade | Swapping the image does not migrate data files | A planned procedure | | Draining the primary's node | Evicts it like any pod | Switchover first | Two details trip people up: 1. **Readiness is not role.** A readiness probe that runs a query says the process accepts connections. It does not say the pod is the writable primary or that a standby is caught up. A Service built only on readiness sends writes to read-only standbys. 2. **Ordinal 0 is not "the primary".** After a failover the primary might be `ocr-db-2`. Anything that hard-codes `ocr-db-0` as the writer breaks at the first failover. ## What fills the gap There are three honest options: 1. **Build it yourself** by running an HA manager beside PostgreSQL and writing the glue: role labels, Service updates, backup jobs. It works, and it is a lot of code nobody else maintains. 2. **Use an operator** such as **CloudNativePG**. It manages pods and PVCs directly (it does not use a StatefulSet), labels the current primary, keeps `-rw`, `-ro` and `-r` Services pointing at the right pods, handles failover and switchover, and creates PodDisruptionBudgets for you. 3. **Use a managed database** outside the cluster and let the provider own failover, backups and patching. The pipeline connects to it like any external dependency. ## How to answer in an interview - Credit the StatefulSet with identity, per-replica storage and ordered updates. - Name the missing pieces explicitly: role, replication, failover, write routing, fencing, backups, upgrades. - Point out that readiness does not equal primary. - Say that the gap is why operators exist, and that a managed service is often the cheaper answer when one is available. The weak answer is "StatefulSets are for databases, so it's fine". The strong answer treats the StatefulSet as the storage and identity layer and asks who owns everything above it.

  • A three-replica PostgreSQL StatefulSet sits behind one ClusterIP Service selecting all its pods. What goes wrong for the OCR writers?
    The Service balances connections across all ready pods, so about two thirds of connections land on read-only standbys and every `INSERT` there fails. Writers need a Service that selects only the current primary, which means something must label the primary pod and move that label on failover. An operator such as CloudNativePG does this with a role label and a dedicated `-rw` Service.
  • Does updating the PostgreSQL image in the StatefulSet from one major version to the next upgrade the database?
    No. The rollout replaces containers one at a time, but the new major version cannot start on the old data directory without a migration step such as `pg_upgrade` or a dump and restore. A plain image swap usually leaves the new pod crash-looping. Major upgrades need a planned procedure, which an operator can automate and a StatefulSet cannot.
  • Is a StatefulSet with per-replica claims ever enough on its own for a data service?
    Yes, when the software handles clustering itself and treats every member the same, or when a single instance with restarts is acceptable. Examples are a development database or a cache you can rebuild. For a primary/standby database whose loss matters, the missing role, failover and backup logic has to come from somewhere.

A StatefulSet is a hotel that guarantees each guest the same room and luggage every night. It does not decide which guest runs the meeting, or who takes over when that guest leaves.

saying these in an interview costs you the question

  • A StatefulSet handles PostgreSQL failover automatically.
  • Ordinal 0 is always the primary, so point writes at it.
  • A passing readiness probe proves a pod is the writable primary.
  • Changing the image tag performs a major-version database upgrade.
  • Per-replica volumes mean backups are no longer needed.
open as a page

A team wants to run PostgreSQL and Kafka inside its Kubernetes cluster rather than use managed services. How do you decide, and what would make you say no?

level: principalimportance: must knowfreq 52%

basics

~20 s

Decide by ownership and risk, not by preference. In-cluster operators suit teams that can run databases, need portability or have no managed option. Otherwise a managed service is usually cheaper overall. Say no when nobody can own recovery.

open as a page

A three-broker Kafka cluster on Kubernetes rejects acks=all writes during routine node drains, although a PodDisruptionBudget allows only one unavailable broker. How can that happen, and what do you change?

level: seniorimportance: should knowfreq 41%

basics

~20 s

A PodDisruptionBudget counts Ready pods, not in-sync replicas. A restarted broker can be Ready while still catching up, so a second eviction is allowed. Slow volume reattachment stretches that window. Gate restarts on replica health and serialise drains.

open as a page

What do Kubernetes data-service operators such as CloudNativePG and Strimzi add on top of plain pods and volumes before a failover, drain or rolling restart is safe?

level: seniorimportance: should knowfreq 44%

basics

~20 s

They add knowledge of the database's own state. CloudNativePG tracks and labels the primary, moves Services, switches over before a drain and creates PodDisruptionBudgets. Strimzi restarts brokers one at a time and only when in-sync replicas allow it.

open as a page

For a database on Kubernetes, what do you give up and gain by using local NVMe PersistentVolumes instead of network-attached block storage?

level: middleimportance: nice to knowfreq 31%

basics

~20 s

Local NVMe gives much lower, steadier I/O latency, but the volume is tied to one node. If that node is lost, the replica's data goes with it, and the pod cannot move. Database replication has to provide the durability.

open as a page