Six teams share one cluster and a seventh wants its own — how do you decide between sharing and one cluster per team?
answer
- cost, blast radius, operational load
- every cluster has a fixed floor
- one cluster, one upgrade calendar
- N clusters is N of everything
- soft separation is not a boundary
basics
~20 sDecide on three axes: the fixed floor each extra cluster carries, the blast radius teams are willing to share, and who absorbs each cluster's operational load. Separation inside one cluster is soft, so a security argument is a different question with a different answer.
solid answer
~50 sI would refuse to answer it as a preference and put it on three axes. Cost: every cluster carries a fixed floor — a deciding half to run and keep available, a full set of add-ons on each one, and per-host capacity reserved for the agent and those add-ons — so seven clusters pay that floor seven times and lose the bin-packing that lets uncorrelated peaks share one pool. Blast radius: one cluster means one upgrade calendar, one cluster-wide change surface and one deciding half whose loss stops every team's new placements at once. Operational load: N clusters is N of everything to patch, rotate and look at during an incident, unless someone builds fleet machinery, which is itself a platform investment. And if the real driver is that one team's data must not be reachable by another even after a compromise, that is a separation argument, not a capacity one, and no budget answers it.
go deeper
Recall the two shapes on offer: many teams inside one cluster, separated by named scopes and budgets, or a cluster each — and that the second one is not simply the same machines rearranged.
Explain the fixed floor a cluster carries before any workload runs: a deciding half to keep available, a set of add-ons on every cluster, and per-host capacity reserved for the agent and those add-ons.
Bring the operating reality: N clusters means N upgrades, N sets of credentials to rotate and N places to look during an incident, and say what fleet machinery you would need before that is sustainable.
Own the call and its price. Name the axis you are optimising, say who pays the floor, be explicit that soft separation is not a security boundary, and propose the split line rather than a cluster per team.
## Three axes, and the one that is not really on the list The shared-versus-dedicated question gets argued on taste far more often than on numbers. There are three axes that actually decide it — **cost**, **blast radius** and **operational load** — plus a fourth thing that is frequently the real driver and is a different question altogether: how strong the separation between teams has to be. | Axis | One shared cluster | One cluster per team | |---|---|---| | Fixed cost | Paid once, amortised across teams | Paid once per team | | Utilisation | Uncorrelated peaks share one pool | Each pool needs its own headroom | | Blast radius | One upgrade, one change surface, everyone | Bounded to one team | | Upgrade cadence | One calendar everyone must accept | Each team on its own | | Operational load | Concentrated on one skilled team | Multiplied, and often pushed onto teams that do not want it | | Separation strength | Soft: names, accounting, access | A real machine and control-plane boundary | ## Cost: the per-cluster floor A cluster is not free before a single workload runs on it: - the **deciding half** — the API surface and the state store behind it — has to run, and to run redundantly if anyone depends on it; - a **full set of add-ons** goes on every cluster: log shipping, metric collection, an external entry point, policy evaluation, certificate handling; - **per-host capacity** is reserved on every machine for the node agent and those add-ons, so a fleet of many small clusters loses a slice of every host; - **headroom** is per pool. One pool absorbs seven teams' peaks when those peaks do not coincide; seven pools each need enough spare capacity to survive their own peak, and the spare is not shareable. That floor is the number most often left out of the argument, because it does not appear on anybody's workload bill. ## Blast radius: what one cluster genuinely shares Against that, the shared cluster concentrates risk. Everything cluster-wide is shared by definition: the deciding half, the upgrade, the installed extension types, cluster-wide configuration, the address space the workload network draws from. Concretely: 1. **One upgrade calendar.** Every team must accept the same window and the same version, and one team's blocker holds up everyone. 2. **One change surface.** A cluster-wide change lands on all seven teams simultaneously; there is no rolling it out to one of them first. 3. **One deciding half.** While it is unavailable, running workloads keep serving — but no new placements, replacements, rollouts or scaling happen for anyone. The failure is shared even though the traffic is not. 4. **Shared exhaustion.** A finite cluster-wide resource — the workload address range is the classic one — is consumed by whoever gets there first. ## Operational load: N of everything Every cluster needs upgrading, its add-ons keeping current, its credentials and certificates rotating, its capacity watching, and a person who knows where to look during an incident. Seven clusters is seven of each, and the honest options are only two: build the machinery to operate them as a fleet — which is a real platform investment with its own staffing — or push the work onto each team, most of which would rather ship their service. A shared cluster concentrates this load on people who chose it; a cluster per team distributes it to people who mostly did not. ## What actually forces a split Some constraints end the debate regardless of cost: 1. **Incompatible cluster-wide requirements.** Two teams needing conflicting installed extension types, or versions that cannot coexist, cannot share — the cluster-wide surface admits only one answer. 2. **An upgrade cadence one team cannot accept**, typically for regulatory or certification reasons. 3. **A separation requirement that soft separation does not meet.** A named scope and a budget stop collisions, accidental consumption and accidental access. They are not a wall around a workload that has broken out of its boundary. When the requirement is that one team's data stays unreachable by another even in that case, the answer is a stronger boundary, and the question has moved from capacity to separation — where the budget has nothing to contribute. ## Where most estates land The binary framing is usually the mistake. What survives contact with reality is a **small number of shared clusters, split along a line that matters** — environment, or risk class, or regulatory scope — with the one or two workloads that genuinely need their own boundary moved out of the shared pool and given one. That pays the per-cluster floor a few times rather than once per team, bounds the blast radius along a line someone can defend, and keeps the operational load on a fleet small enough to actually operate. The decision to state out loud to the seventh team: which of the three axes are you actually buying, and are you willing to pay the floor to get it?
- Which single technical fact most often forces a team onto its own cluster?A cluster-wide requirement that conflicts with another team's. The deciding half, the installed extension types and the upgrade calendar are cluster-wide by construction, so two teams needing incompatible ones cannot share however generous the budget is. Cost arguments lose to that one immediately, which is why it is worth checking first.
- What makes this decision easier than the binary framing suggests?It is rarely all-or-nothing. Most estates settle on a small number of shared clusters split along a line that means something — environment, risk class, regulatory scope — with the one or two workloads that genuinely need their own boundary lifted out. The floor gets paid a few times instead of once per team, and the blast radius follows a line you can defend.
saying these in an interview costs you the question
- Argues extra clusters are free because the hosts are the same
- Treats a named scope as a security boundary between teams
- Forgets each cluster needs its own add-ons, upgrades and credentials
- Assumes one shared cluster is cheapest at any team count
- Decides on team preference rather than blast radius and cost