When is a separate cluster, not a namespace, the only honest trust boundary between two tenants?
answer
- Risk and cost, not isolation strength
- Ask who wrote the workload code
- Blast radius beats likelihood here
- Node pools before second control plane
- A boundary you cannot operate is theatre
basics
~20 sSplit when the tenants' threat models differ by more than the cluster can enforce: one tenant runs code nobody reviewed, the assets sit at very different values, or a shared control plane compromise is unacceptable. Otherwise separate node pools usually buy more than a second cluster.
solid answer
~50 sI ask three questions. Who writes the workload code — if a tenant submits arbitrary jobs nobody reviews, the host kernel is my only boundary and a single escape ends the story, so that tenant belongs on separate hardware or a separate cluster. How far apart are the assets — regulated payment flow beside a marketing site means one tenant's ordinary compromise is the other's incident, and the blast radius, not the likelihood, drives the split. And can we accept a shared control plane, since one API server and one platform team are common-mode to every namespace in the cluster. If none of those bite, I prefer the cheaper rungs: dedicated node pools, sandboxed runtimes for untrusted work, tighter scoping. A fleet of five badly patched clusters is worse than one well-run cluster, so I count the operational cost of the boundary as part of the decision and write the accepted risk down with an owner and a review date.
go deeper
Be ready to say that a namespace and a separate cluster are not the same strength of separation, and that running code nobody reviewed next to valuable workloads is the risky case.
Explain the intermediate options between one namespace and two clusters — dedicated node pools, stronger workload isolation for untrusted jobs — and what each one actually removes from the threat list.
Demonstrate the judgment on a live design: pick a rung, justify it by attacker position and asset value, and state plainly what the chosen arrangement still leaves exposed.
Own the decision end to end: argue it in risk and operating-cost terms to funders, refuse the reflex split, and close it as a written accepted risk with an owner and a trigger for review.
## The question behind the question `Namespace or separate cluster` is usually presented as an isolation argument and is really a **risk and operating-cost argument**. Everyone in the room already knows a namespace is weaker than a cluster. The lead's job is to say when that weakness matters enough to pay for the alternative, and to say it in terms the people funding the platform can act on. ## Criteria that genuinely force a split **Who supplies the workload code.** This is the strongest signal. If a tenant runs code you reviewed and built, your boundary is defence in depth. If a tenant submits arbitrary code — a research GPU cluster taking student- and lab-submitted training jobs is the clean example — then the host kernel is not a backstop, it is *the* boundary, and it is being probed by the workload itself. Add host mounts for a shared dataset volume and the boundary has already been removed by configuration before any attacker arrives. Here the attacker is an authenticated, legitimate tenant, and the assets are other labs' checkpoints and models plus the compute itself. That combination — untrusted code plus valuable neighbours — is where a shared cluster stops being defensible. **Asymmetric asset value.** When one tenant's worst day is a defaced page and another's is money moving incorrectly, an ordinary compromise on the low side becomes an incident on the high side. The split is justified by blast radius rather than by probability, and it is worth saying that explicitly, because otherwise the argument gets answered with `the marketing site is not a likely target`, which is not the point. **Regulatory or contractual separation.** When an auditor or a customer contract requires demonstrable separation, a namespace is hard to evidence and a distinct cluster is easy. That is a legitimate reason even where the technical risk is modest. **Common-mode control plane.** One API server, one datastore, one set of platform components installed for everyone, one platform team. Every namespace boundary in the cluster depends on all of it. If your model cannot accept that dependency for a given tenant, no arrangement inside the cluster fixes it. ## The ladder, cheapest rung first Splitting is not binary, and jumping straight to `two clusters` skips options that often carry most of the benefit: 1. Namespace separation alone — fine for tenants with similar assets and reviewed code. 2. Dedicated node pools — moves the boundary onto the kernel line without a second control plane, and removes the shared node credential. 3. Stronger workload isolation for the untrusted subset — sandboxed or virtualised runtimes, so that only the jobs that need it pay the cost. 4. Separate clusters — a separate control plane and separate platform state. 5. Separate cloud accounts or subscriptions around those clusters — where the identity and network blast radius must be split too. Most real decisions land on rung 2 or 3. Recommending rung 4 by reflex is how platform teams acquire a fleet. ## What a split does not buy Be explicit about the residue, or the org will believe it bought more than it did. After two clusters exist, the tenants typically still share: the identity provider that both clusters trust; the platform team and their credentials; the images and platform components installed in both; the network and cloud account around them; and the same people on call. The boundary has **moved up a level**, not disappeared, and the highest-value credential in the estate is now the one that administers both clusters. ## Operating cost is part of the security argument Every cluster is a recurring bill: upgrades, node lifecycle, cluster-level configuration, drift between clusters, and one more place for a misconfiguration to hide. A boundary the organisation cannot operate is theatre — five stale clusters are a worse security outcome than one well-run cluster with dedicated node pools. A lead who cannot make that argument tends to get overruled later on cost, which is the worst of both outcomes. ## How to close the decision Don't answer a request for one shared cluster with a flat no. Name the assumption the shared arrangement rests on, offer the ladder with the cost of each rung, propose detection where the boundary stays soft, and record the outcome as an accepted risk with a named owner and a review date tied to a trigger — a new tenant type, untrusted workloads arriving, or an asset changing value. That converts an argument into a decision the organisation owns, which is the actual deliverable at this level.
- A team wants everything on one cluster for cost. What do you offer instead of a flat refusal?Name the assumption their proposal rests on, price the intermediate rungs — dedicated node pools, a sandboxed runtime for the untrusted subset — and propose detection where the boundary stays soft. Then record the residual risk as an accepted decision with a named owner and a review trigger. That way the org keeps the saving and also owns the exposure explicitly.
- Which risks survive a split into two clusters?The shared identity provider both clusters trust, the platform team and their administrative credentials, the shared platform components and images installed in both, the surrounding network and cloud account, and the same operators on call. The boundary moves up a level rather than away, so the model should be redrawn around the new highest-value credential.
- How do you avoid the fleet-of-clusters failure mode?Split only where a threat model genuinely differs, and count the recurring cost — upgrades, cluster-level configuration, drift, one more place to hide a misconfiguration — as part of the decision. If the org cannot commit to operating the extra cluster to the same standard, the split lowers real security while raising the diagram's quality.
Two locked offices in one building versus two separate buildings: the second stops anyone crawling through the shared ceiling void, but both still share the street, the power supply and the same security company.
saying these in an interview costs you the question
- Always recommends a separate cluster to be safe
- Ignores the operating cost of extra clusters
- Claims a cluster split removes all shared risk
- Treats tenant-submitted code the same as reviewed code
- Argues likelihood when blast radius is the issue
- Leaves the decision unwritten with no owner