skip to content

Your Kubernetes platform is one 38-node cluster with mixed spot and on-demand pools, and its loyalty-points accrual namespace alone runs 1,180 pods. How would you decide whether to split it into several clusters?

level: principalimportance: should knowfreq 46%

answer

  1. drivers before size
  2. 31 pods per node vs 110
  3. fixed cost per cluster
  4. stranded spot headroom
  5. tooling before the split

basics

~20 s

Split only when a concrete driver (blast radius, region, compliance scope or hard tenancy) outweighs the extra control planes, add-on stacks, upgrades and stranded capacity, and only once fleet tooling can run many clusters as consistently as one.

solid answer

~50 s

I would start from failure and compliance domains, not size. The accrual namespace averages about 31 pods per node (1,180 / 38), well under the kubelet's default `maxPods` of 110, and 38 nodes is far from the roughly 5,000-node envelope in upstream large-cluster guidance. Scale alone does not justify a split. Next I would check the real drivers: past cross-team outages from shared webhooks, CRDs or upgrades; residency or audit scope for accrual data; and regional survivability. Against that I would count the costs: each cluster adds an HA control plane or a managed fee, a full add-on stack, an upgrade and certificate cycle every release, and stranded headroom, because spare spot capacity in one cluster cannot absorb another cluster's peak. I would split only after declarative lifecycle (Cluster API or equivalent), templated configuration, fleet observability and cross-cluster discovery exist. The likely result is a few clusters cut by region and environment, with accrual moved only if it has its own driver.

go deeper

for a junior

Remember that a bigger cluster is not automatically a problem; the real reasons to split are about failures, regions and compliance.

for a middle

Check the numbers: pods per node against kubelet maxPods, cluster size against upstream guidance, and which control-plane signals would show real strain.

for a senior

Show the operating costs of a split in concrete terms: add-on consistency, upgrade cycles and on-demand floors for spot-heavy workloads.

for a principal

Make a decision tied to named drivers, state which tooling must exist first, and propose an end state that will not grow with headcount.

## Start from drivers, not size The weakest reason to split a Kubernetes cluster is "it feels big". The strong reasons are **failure domains** and **control domains**: what can break everyone at once, what an auditor needs fenced off, and what must survive a regional outage. A principal-level answer names the drivers that actually apply, prices the split, and states what has to be true before the split is safe. ## Checking the scale argument For the cluster in question: - 1,180 accrual pods across 38 nodes is **about 31 pods per node** for that namespace alone (1,180 / 38 = 31.05). - The kubelet's default `maxPods` is **110**, so per-node density leaves plenty of room even after other namespaces. - Upstream large-cluster guidance is tested to roughly **5,000 nodes and 150,000 pods**. A 38-node cluster is far from that. So scale is not the driver. The signals that *would* show control-plane pressure are elevated kube-apiserver request latency, requests rejected with 429 by API Priority and Fairness, growing etcd database size and commit latency, and a steady backlog of pending pods waiting for scheduling. Measure those before accepting a size argument. ## The cost side | Dimension | One larger cluster | Several smaller clusters | |---|---|---| | Control planes | One HA control plane to run or pay for | One per cluster | | Add-ons | One stack of DNS, Gateway/ingress, metrics, logging, policy engine, GitOps agent | One stack per cluster, kept consistent | | Upgrades and certificates | One cycle per release | N cycles, often left at different versions | | Capacity | Spot and on-demand headroom is pooled; bin-packing is efficient | Each cluster needs its own on-demand floor and spare room | | Blast radius | A cluster-level fault hits every team | Contained to one cluster | | Cross-cluster calls | None | Needs discovery (for example `ServiceExport`/`ServiceImport`) and routable networking | | Access and audit | One RBAC surface | N surfaces, simpler for each compliance scope | The mixed spot and on-demand pool matters. In one cluster, a spot reclaim can be absorbed by on-demand headroom that every team shares. Split the cluster, and each piece needs its own on-demand floor for pods that cannot tolerate reclaims, which raises total cost. ## Tooling that must exist before splitting 1. **Declarative cluster lifecycle.** Cluster API or an equivalent, so creating and upgrading N clusters is one reviewed change, not N runbooks. 2. **Templated cluster configuration.** Add-ons, admission policy and RBAC delivered identically to every cluster, usually by a GitOps controller. 3. **Fleet observability.** One place to see every cluster's version, health and capacity. 4. **Cross-cluster discovery and networking** for any service that callers in other clusters need. 5. **An upgrade cadence** that stops clusters drifting several minor versions apart. Without these, a split turns one well-run cluster into several neglected ones. ## A decision sketch for this cluster - **Region.** If accrual must survive a regional outage, add a second regional cluster and put both behind cluster-set discovery. - **Compliance.** If accrual ledgers fall under an audit or residency scope, give that scope its own cluster so the audit boundary is the cluster boundary. - **Blast radius.** If platform-wide webhooks or CRD upgrades have already caused incidents across teams, separate production from the platform team's experiments first. That usually means clusters per environment, and staged rollouts of platform components across clusters. - **None of the above.** Keep one cluster and improve isolation inside it, which is the tenancy leaf's subject, not this one. ## Common end states and anti-patterns - **Good:** a small matrix of region × environment, plus a cluster for any hard compliance scope. - **Anti-pattern:** one cluster per team by default. Overhead grows with headcount, and versions drift. - **Anti-pattern:** splitting first and building fleet tooling later. - **Anti-pattern:** stretching one cluster across regions to avoid running two.

  • How would the decision change if accrual data for EU customers had to stay in the EU?
    Residency becomes a hard driver. I would run an EU-region cluster for the EU accrual workload and its data stores, keep other regions' data out of it, and let only non-residency data or aggregated results cross regions. The split follows the residency boundary, so the audit scope is a whole cluster rather than a set of namespaces and node pools.
  • Which signals would convince you the single cluster is hitting control-plane limits?
    Sustained rises in kube-apiserver request latency, requests rejected with 429 by API Priority and Fairness, a growing etcd database with rising commit latency, and a persistent backlog of pods pending scheduling. If those appear under real load and tuning does not fix them, scale becomes a genuine driver. Pod count alone does not.
  • Why is one cluster per team rarely the right default?
    Fixed per-cluster overhead then grows with headcount: another control plane, another add-on stack and another upgrade cycle per team. Capacity fragments, and clusters drift across versions. Most teams need isolation of permissions and resources, not their own control plane. Team-level clusters make sense only when a team has a real driver such as a compliance scope or conflicting cluster-scoped dependencies.

saying these in an interview costs you the question

  • Split because 1,180 pods in one namespace is too many for Kubernetes.
  • One cluster per team is the default best practice.
  • Many small clusters are always cheaper because each is smaller.
  • Split first and build the fleet tooling afterwards.
  • A large cluster has no blast-radius risk if its nodes are redundant.