skip to content

Why do organizations running Kubernetes split workloads across many clusters instead of one large cluster, and what does each driver buy them?

level: middleimportance: must knowfreq 62%

answer

  1. one cluster, one failure domain
  2. cluster-scoped objects are shared
  3. blast radius, region, compliance, tenancy
  4. etcd latency keeps clusters regional
  5. fixed overhead multiplies per cluster

basics

~10 s

A Kubernetes cluster is one failure and control domain, so teams run many clusters to limit blast radius, place workloads in specific regions, draw compliance boundaries and give tenants isolation that namespaces cannot provide.

solid answer

~50 s

Everything in one cluster shares one API server, one etcd, one set of cluster-scoped objects (CRDs, admission webhooks, ClusterRoles) and one upgrade schedule, so a fault at that layer hits every workload in the cluster. Teams add clusters for four main reasons. **Blast radius:** a bad webhook, a failed upgrade or an overloaded API server only affects the cluster it happens in. **Region:** a cluster is normally kept within one region because etcd and the control plane are sensitive to latency, so serving users close to them, or surviving the loss of a region, needs another cluster. **Compliance:** a PCI or data-residency scope is much easier to audit when it covers a whole cluster. **Hard tenancy:** some tenants need their own CRD versions, their own cluster admins or no shared kernel. Upstream scale limits and different upgrade schedules are secondary reasons. The cost is overhead that repeats for every cluster: N control planes, N add-on stacks, N upgrades, plus cross-cluster service discovery.

go deeper

for a junior

Remember that everything in a cluster shares one API server and one etcd, so a problem there affects every workload in the cluster.

for a middle

Name the four drivers (blast radius, region, compliance, hard tenancy) and explain why cluster-scoped objects and etcd latency are what push teams toward separate clusters.

for a senior

Give a real incident where one cluster's shared layer, such as a webhook, a CRD or an upgrade, took down unrelated teams, and put the per-cluster overhead against it.

for a principal

Frame the split as matching cluster boundaries to real failure and compliance domains, and justify each extra cluster against the fleet tooling and cost it adds.

## What one Kubernetes cluster shares A **cluster** is one Kubernetes control plane (kube-apiserver, etcd, kube-scheduler, kube-controller-manager) plus the nodes registered with it. Namespaces divide the *namespaced* API objects, but a large part of the cluster is shared by everything in it: - **The control plane.** Every request from every team goes to the same API server and is stored in the same etcd. - **Cluster-scoped objects.** `CustomResourceDefinition`, `ClusterRole`, `ValidatingWebhookConfiguration`, `MutatingWebhookConfiguration`, `StorageClass`, `PriorityClass` and `Node` have no namespace. The cluster holds one copy of each. - **The release.** The whole cluster runs one Kubernetes minor version, and an upgrade affects everyone. - **Nodes and kernels,** unless you deliberately partition them. This makes the cluster the natural **failure domain** and **control domain**. Whatever breaks at the cluster level breaks for every workload in it. ## The four main drivers | Driver | What goes wrong with one cluster | What a separate cluster buys | |---|---|---| | **Blast radius** | A webhook with `failurePolicy: Fail` goes down, a CRD change breaks controllers, or the API server is overloaded, and every team is affected | The fault stays inside one cluster | | **Region / latency** | etcd needs a majority of members to acknowledge each write, so stretching it across regions slows every write and makes partitions dangerous | Each region has its own control plane and can fail independently | | **Compliance / residency** | An auditor must reason about shared nodes, shared admins and shared cluster-scoped policy | The whole cluster is in scope, which is easy to explain and prove | | **Hard tenancy** | Tenants need different CRD versions, their own cluster-admin, or no shared kernel | Each tenant gets its own cluster-scoped objects and its own admins | ## Secondary drivers 1. **Scale limits.** Upstream large-cluster guidance is tested up to roughly 5,000 nodes and 150,000 pods, and many operators hit API-server or etcd pressure well before that. 2. **Upgrade independence.** A team with a strict freeze window can stay on its own schedule without holding back the whole platform. 3. **Environment separation.** Keeping development, staging and production in separate clusters stops a test run from exhausting production's control plane. 4. **Hardware or network specialisation.** Examples are a GPU fleet or an isolated network segment. ## What splitting costs Every extra cluster multiplies fixed costs: - **Another control plane.** That means three or five etcd members plus API servers to run, or another managed control-plane fee. - **Another add-on stack.** Each cluster needs its own cluster DNS, ingress or Gateway, metrics, log shipping, a policy engine and a GitOps agent. - **Another round of upgrades and certificate renewals** every release cycle. - **Stranded capacity.** Spare room in one cluster cannot absorb a spike in another. - **Cross-cluster plumbing.** Callers in one cluster need a way to find services in another, for example the Multi-Cluster Services API with `ServiceExport` and `ServiceImport`. This is why fleet tooling such as **Cluster API** (declarative cluster lifecycle) matters: without it, N clusters means N times the manual work. ## A worked example A loyalty-points accrual service runs in one 38-node cluster that has a mixed spot and on-demand node pool, and its namespace holds 1,180 pods. A platform team rolls out a mutating admission webhook with `failurePolicy: Fail`, and its backing pods land on spot nodes that get reclaimed together. Until the webhook recovers, pod creation fails for every namespace the webhook matches, including accrual scale-ups. With the accrual service in its own cluster, or with the webhook rolled out to one cluster first, the outage would have been limited. That is the blast-radius argument in concrete form. ## When one cluster is still the right answer - A small team with no residency or compliance split. - Workloads that need **tight bin-packing** across shared spot capacity. - An organisation that has no automation to run a fleet yet. The usual end state is a **small number of clusters cut along real drivers** (region × environment, plus any compliance scope), not one cluster per team by default.

  • Why does a cluster-scoped object like a CRD push two tenants into separate clusters?
    `CustomResourceDefinition` has no namespace, so a cluster serves one definition per resource name, with one set of versions and one schema. If two tenants need incompatible versions of the same operator's CRDs, or one tenant's admission webhook must not affect another, namespaces cannot separate them. Only a separate control plane gives each tenant its own copy of the cluster-scoped objects.
  • Why not stretch one Kubernetes cluster across two regions for disaster recovery?
    etcd is a Raft quorum, so every write waits for a majority of members. Latency between regions slows every API write, and a network partition leaves the minority side unable to accept writes. The stretched cluster is also still one blast radius: a bad upgrade or webhook hits both regions at once. Two regional clusters with independent control planes are the usual design.

Clusters are like separate buildings on a campus rather than rooms in one building: rooms share one fire alarm, one electrical panel and one renovation schedule, while a second building costs its own utilities but survives the first one's fire.

saying these in an interview costs you the question

  • More clusters always means higher availability at no extra cost.
  • Namespaces give the same isolation as separate clusters.
  • Stretch one cluster across regions for disaster recovery.
  • A cluster has no practical upper size limit.
  • Splitting clusters costs nothing beyond buying extra nodes.