skip to content

As a Kubernetes platform owner, how do you decide between namespace-per-tenant in a shared cluster and a cluster per tenant, and how would you defend that choice?

level: principalimportance: should knowfreq 38%

answer

  1. trust before tenant count
  2. three tiers of isolation
  3. cluster-scoped needs force clusters
  4. pooled capacity versus fleet toil
  5. publish a promotion path

basics

~20 s

Choose by trust and by what tenants need to control. Trusted teams that need only namespaced objects fit namespace-per-tenant. Hostile code, compliance walls, or tenants needing their own CRDs, webhooks or cluster-admin justify separate clusters, at the cost of fleet operations and unpooled capacity.

solid answer

~50 s

I start from the threat model, not the tenant count. If tenants are our own teams and need only namespaced objects, **namespace-per-tenant** with the standard bundle is the cheapest answer: one control plane, pooled capacity and one upgrade. I move a tenant to **dedicated nodes** when it runs untrusted code or input, and to **its own cluster** when any of these holds: tenants may be hostile to each other; a contract or regulator demands no shared control plane; the tenant needs cluster-scoped control such as its own CRD versions, admission webhooks or cluster-admin; or its API load or upgrade cadence would harm others. A shared cluster whose tenants' operators installed 612 CRDs is a good example: every tenant is pinned to the same CRD versions. I defend the choice with a written matrix of what each model isolates, the cost per tenant, and an explicit path to promote a tenant up a tier.

go deeper

for a junior

Know that teams can share one cluster through namespaces or each get their own cluster, and that the second is stronger but costlier.

for a middle

Explain which surfaces each model isolates: kernel, nodes, control plane and cluster-scoped objects.

for a senior

Identify the concrete triggers for a separate cluster: hostile code, compliance, cluster-scoped needs, and API load or upgrade divergence.

for a principal

Build a tiered policy with priced tiers and a promotion path, and defend it with the threat model instead of blanket rules.

## The decision in one line **Isolation strength should follow trust and control needs. Headcount should not drive it.** Namespace-per-tenant is the cheapest model and a soft boundary. A cluster per tenant is the strongest model and the most expensive. Dedicated node pools sit between them. A principal is expected to state the criteria, show the costs, and design a path between tiers rather than pick one model for everyone. ## What each model isolates | Model | Kernel | Nodes | Control plane | Cluster-scoped objects | Relative cost | |---|---|---|---|---|---| | Namespace-per-tenant | Shared | Shared | Shared | Shared | Lowest | | Namespace + dedicated node pool | Separate | Separate | Shared | Shared | Medium | | Cluster per tenant | Separate | Separate | Separate | Separate | Highest | A fourth option, **a virtual control plane per tenant** hosted inside a shared cluster, gives tenants their own API view and cluster-scoped objects while sharing the physical nodes. It is worth naming as a middle path, but it adds its own moving parts. ## Criteria that push toward separate clusters - **Hostile or unknown tenants.** Customer code, or workloads such as a PDF-invoice renderer that parse untrusted uploads, where a kernel escape must not reach another customer. - **Compliance walls.** A contract or regulator requires that no control plane, etcd or admin identity is shared. - **Cluster-scoped control.** CRDs, admission webhooks, StorageClasses and ClusterRoles cannot be scoped to a namespace. In a cluster whose tenants' operators have installed **612 CRDs**, every tenant is pinned to one version of each. One team's operator upgrade becomes a cross-tenant change, and discovery and API server memory grow for everyone. - **Control-plane noise.** One tenant's controllers issuing heavy LIST or WATCH traffic. API Priority and Fairness separates flows but not priority-level capacity. - **Divergent lifecycle.** A tenant that must stay on an older Kubernetes release, or wants a faster upgrade cadence, cannot share a control plane with the others. ## Criteria that keep tenants in a shared cluster - Tenants are **internal teams** under one security organisation. - Workloads need only **namespaced** objects plus platform-provided CRDs. - **Utilisation matters.** An 80-node cluster that autoscales between 20 and 80 nodes pools headroom across tenants. Splitting it into a dozen small clusters leaves each with its own idle minimum and its own control-plane overhead. - The team is **small**. Every extra cluster is another upgrade, another set of add-ons, another certificate rotation and another place for drift. ## A decision process you can defend 1. **Classify tenants** by trust (internal, partner, customer) and by the data they handle. 2. **Record control needs**: will they install CRDs, webhooks or cluster-wide RBAC? 3. **Map each class to a tier.** For example, internal teams get namespaces, untrusted workloads get dedicated nodes with a sandboxed runtime, and customers or regulated data get their own cluster. 4. **Price each tier**: control-plane cost, minimum node count, operational hours per cluster. 5. **Publish the promotion path.** A tenant moving up a tier is a planned migration, not a crisis. Tooling that stamps out both namespaces and clusters from templates keeps that cheap. 6. **Review regularly.** Growth in CRD count, API load or trust changes should trigger a review. ## How to defend it in the room - Lead with the **threat model** and show the isolation table. Stakeholders accept a cost once they see which row it buys. - Admit what soft tenancy leaves shared: the kernel, the control plane and cluster-scoped objects. Do not oversell a namespace. - Show the **cost curve**. Clusters rarely fail on day one. They fail when a small team is running thirty of them. - Keep **multi-cluster operations** separate. Once a tier needs many clusters, provisioning and connecting that fleet is its own programme.

  • A tenant in your shared Kubernetes cluster wants to install an operator that ships its own CRDs and a validating webhook. What do you do?
    Both are cluster-scoped. The CRD version would bind every tenant, and a failing webhook with a broad scope could block other tenants' writes. I either take the operator into the platform, versioned and owned centrally for everyone, or promote that tenant to its own cluster. I do not grant a tenant cluster-scoped install rights in a shared cluster.
  • Leadership proposes one cluster per team for all 14 internal teams to 'be safe'. How do you respond?
    I ask what threat it addresses. For trusted internal teams, the namespace bundle plus dedicated nodes for the few untrusted workloads covers the realistic risks. Fourteen clusters mean fourteen control planes, upgrades and add-on stacks, and unpooled headroom. I would reserve separate clusters for teams with compliance walls or cluster-scoped needs.

saying these in an interview costs you the question

  • Every tenant should get its own cluster because namespaces are insecure.
  • Tenant count alone decides when to split clusters.
  • A tenant can safely install its own CRDs in a shared cluster.
  • Dedicated node pools remove the shared control plane.
  • Separate clusters cost roughly the same as namespaces to operate.