skip to content

A team wants to run PostgreSQL and Kafka inside its Kubernetes cluster rather than use managed services. How do you decide, and what would make you say no?

level: principalimportance: must knowfreq 52%

answer

  1. who owns recovery at 3 a.m.
  2. is a managed option available
  3. operator is a dependency too
  4. dedicated pool, off-cluster backups
  5. no tested restore, no go

basics

~20 s

Decide by ownership and risk, not by preference. In-cluster operators suit teams that can run databases, need portability or have no managed option. Otherwise a managed service is usually cheaper overall. Say no when nobody can own recovery.

solid answer

~50 s

I start with **who owns recovery**. Running PostgreSQL or Kafka in the cluster means the team owns failover, backups, restore testing, upgrades and 3 a.m. pages, both for the data service and for its operator. A **managed service** moves most of that to the provider for a price, with less control and some lock-in. In-cluster wins when there is **no managed option** (for example on premises), when **portability or residency** is a hard requirement, when **scale** makes managed pricing unreasonable, or when the team already has database skills. I would say no if nobody can own on-call for the database, if restores have never been tested, if storage is not good enough, or if cluster upgrades and node churn would keep disrupting the data. If we go ahead: a mature operator such as CloudNativePG or Strimzi, a dedicated tainted node pool, tested backups, and cluster maintenance that waits on database health.

go deeper

for a junior

Recall that running a database in the cluster means your team is responsible for failover, backups and upgrades.

for a middle

Explain what an operator automates and what it leaves: backups off-cluster, restore tests and maintenance that respects database health.

for a senior

Bring operating evidence: drain incidents, restore timings, storage behaviour, and what dedicated node pools changed.

for a principal

Make the decision explicit: criteria, ownership split, the conditions for saying no, and the signals that would make you reverse it.

## Frame it as ownership, not technology Kubernetes can run PostgreSQL and Kafka well. Operators such as **CloudNativePG** and **Strimzi** are mature, and StatefulSet-style identity plus per-replica volumes solve the plumbing. The real question is **who will own the data when something goes wrong**, and whether that is a better use of the team than paying a provider. The case: a document-OCR pipeline on a **3-control-plane, 27-worker self-managed cluster**. OCR workers each request 0.35 CPU core. They publish page events to Kafka and store extracted text in a PostgreSQL database of about 412 GiB. The team asks to run both data services in the cluster. ## The decision criteria | Criterion | Favours in-cluster | Favours managed service | |---|---|---| | A managed option exists | None (on premises, air-gapped) | Available where the cluster runs | | Team skills | Engineers who have run the database in production | No database on-call experience | | Portability and residency | Must run the same way everywhere | One provider is acceptable | | Cost at scale | Large, steady footprint | Small or bursty footprint | | Storage quality | Fast, replicated or local disks with automated rebuilds | Weak or shared storage | | Cluster churn | Stable nodes, planned upgrades | Frequent node replacement | | Recovery objectives | Met by tested operator backups | Tight, and cheaper to buy | A self-managed cluster is a strong hint. If it runs in a data centre, a cloud-managed database may not be available at all, or reaching it may add latency and data-transfer cost. That changes the question from "whether" to "how safely". ## What "in-cluster" really commits you to - **Operating the operator**: its upgrades, CRD version changes and bugs, and a plan for when its controller is down. - **Backups and restores**: backups to object storage outside the cluster, retention and regular restore tests. The cluster is not a backup location. - **Maintenance discipline**: node drains and Kubernetes upgrades must wait on database health, not only on PodDisruptionBudgets. - **Isolation**: a dedicated node pool with taints, so a burst of OCR workers cannot starve the database and a routine worker rollout does not drain database nodes. - **Storage decisions**: network block storage (mobile, slower) or local NVMe (fast, tied to nodes). - **Observability and on-call**: replication lag, under-replicated partitions, disk growth and operator health, with people who can act on them. - **Blast radius**: a cluster-wide incident (a control-plane outage, a bad admission webhook, an etcd problem) now affects the data too. Some teams give data services their own cluster for this reason. ## When to say no 1. Nobody will be on call for the database, or the plan is "the platform team will figure it out". 2. Restores have never been tested, and there is no time budgeted to test them. 3. The only storage is slow or shared, with no plan for replica rebuilds. 4. Nodes are replaced often by automation that does not know about database health. 5. A managed service exists, meets the requirements, and costs less than the engineering time. ## How to present the decision Write the decision down as a short record the team can challenge: the options considered, the criteria above with an honest rating for each, the chosen option, who is on call, and the recovery objectives it must meet. Include a cost comparison that counts engineering time, not only infrastructure spend, because the time spent running an operator, testing restores and handling drains is usually the largest hidden cost of the in-cluster option. ## A reasoned recommendation For the OCR pipeline, a reasonable answer is: - **Kafka in-cluster with Strimzi** if there is no managed option and the team can run it, because event throughput is steady and the data is replayable for a limited time. - **PostgreSQL** in-cluster with CloudNativePG only if backups go off-cluster and restores are tested. Otherwise a managed database if one is reachable, because the extracted text is the system of record. - Either way, write down the ownership split, the recovery objectives, and a review date. There is no universal right answer. A principal is expected to make the costs visible, say what would change the decision, and refuse the option nobody can own.

  • Would you put the OCR pipeline's databases in the same cluster as the OCR workers, or in a separate cluster?
    A separate cluster isolates the data from application-driven churn: frequent upgrades, admission changes and noisy neighbours. It costs another control plane to run. The same cluster with a tainted, dedicated node pool is simpler and often enough when upgrades are planned carefully. I would choose a separate cluster when the application cluster changes often or has had cluster-wide incidents.
  • What evidence would make you reverse an in-cluster decision a year later?
    Repeated incidents where maintenance disrupted the database, missed recovery objectives in restore tests, operator upgrades that keep slipping, or database on-call load that crowds out platform work. Also a managed option becoming available at an acceptable cost. Deciding on these signals in advance makes the reversal a planned review, not a crisis.

saying these in an interview costs you the question

  • Kubernetes is never suitable for running databases.
  • Kubernetes makes databases highly available automatically.
  • Installing an operator removes the need for database skills on the team.
  • Volume snapshots inside the cluster count as off-site backups.
  • Managed services are always cheaper regardless of scale.
  • A PodDisruptionBudget is enough to make cluster upgrades safe for data.