skip to content

Two teams share one cluster that is always up and move to a cluster per job. Which failures does that isolate, and which stay shared?

level: seniorimportance: should knowfreq 48%

answer

  1. what was the scope of the failure
  2. machines separate, dependencies do not
  3. versions and restarts become per team
  4. storage, metadata and quota stay common
  5. bursty lookups can get worse

basics

~20 s

It isolates everything that lived on the machines: memory exhaustion, a crashed coordinating process, runtime and library versions, restarts and upgrades. It does not isolate shared storage, a shared metadata service, the quota with whoever grants machines, or the credentials both teams use.

solid answer

~40 s

Moving to a cluster per job separates the things whose scope was the machines. Each job gets its own worker processes — the processes that run pieces of a job and own their memory — its own coordinating process, its own runtime and library versions, and its own upgrade and restart schedule, so one team exhausting memory or crashing its job now harms only itself. What does not move is everything outside the machines: the storage layer both teams read and write, any shared metadata service consulted at planning time, the quota held with whatever grants machines, the network, and shared credentials. Those remain common failure points, and the second one is the one candidates forget: a thousand short-lived clusters can hammer a shared metadata service far harder than one long-lived cluster did.

go deeper

for a junior

Know the basic claim: separate machines per job means one job's crash or memory exhaustion cannot touch another team's job.

for a middle

Be able to list what actually moves — memory, local disk, runtime versions, restarts, blast radius — and explain why each of those had the machines as its scope.

for a senior

Name the couplings that survive: shared storage and its limits, a shared metadata service, the machine quota, credentials and images — and explain why bursty start-ups can make one of them worse.

for a principal

Decide which failures the organisation is actually paying to isolate, and whether a per-job model buys enough of that to justify operating it alongside whatever else you run.

## What isolation means at this level Isolation here means one job's trouble not reaching another job belonging to someone else. It is a question about **shares, separate processes and separate lifecycles**, not about kernel-level fencing or per-process limits on a single machine — that is a different subject with a different owner. The useful way to answer is to ask, for each failure you can name: what was its scope? If the scope was the machines, a cluster per job isolates it. If the scope was something outside the machines, nothing about the supply model changed it. ## What a cluster per job genuinely isolates - **Memory exhaustion.** A worker process — the process that runs pieces of the job and owns their memory — can be driven out of memory by one job's data. On shared machines, the other team's worker processes were on those machines too. Now they are not. - **Local disk saturation.** One job filling the scratch space it writes to when memory runs short used to fill it for everyone on that machine. - **Losing the coordinating process.** The single process that plans the pieces, hands them out and tracks what finished takes its job down with it when it dies. Per job, that blast radius is one job. On shared machines it was one job too — but a crash that took the machine down with it reached everything on that machine. - **Runtime and library versions.** This is the big one in practice. A shared cluster has one set, and upgrading it is a negotiation between every team on it. Per job, each job brings its own, and one team can move without asking. - **Restarts and maintenance.** Restarting a cluster per job is invisible; restarting a shared one is an outage for everybody submitted to it. - **Leftovers.** Files, caches and half-cleaned directories from a previous job cannot confuse a later one, because there is no later one on those machines. ## What stays shared | shared component | still coupled? | how it bites | |---|---|---| | the storage layer both teams read and write | yes | request rate limits, throughput ceilings, one team's scan starving another's | | a metadata or catalog service consulted before running | yes | more clusters means more short bursts of lookups, often worse than before | | the quota with whatever grants machines | yes | one team's fleet of clusters can exhaust what the other team needs to start | | network paths and egress | yes | bandwidth is shared regardless of who owns the machines | | credentials and access paths | yes | a permission held by both teams is still held by both teams | | the pipeline that builds machine images | yes | a bad image reaches every new cluster at once | The third row is the surprise in the interview. Moving off a shared cluster does not remove contention; it moves it up a level, from competing for worker processes inside a pool to competing for the right to raise machines at all. And the second row frequently gets **worse**: a long-lived cluster consults a metadata service steadily, whereas hundreds of short-lived clusters consult it in bursts at start-up, which is a harsher shape for that service. ## Where the boundary sits There are two boundaries that sound alike and are not: 1. **Between jobs sharing a pool** — this leaf's boundary, drawn with separate processes, separate machines and separate lifecycles. 2. **Between processes on one machine** — requests, limits and kernel-level fencing, which belong to the container and orchestration subject, not here. Even with a cluster per job, if two jobs' machines are co-tenants on the same underlying hardware, that second boundary is the only one protecting them, and it is not yours. How a finite pool divides its capacity between competing jobs while they run — shares, priorities, taking capacity back — is its own subject and not what "isolation" means here. ## The honest summary to give A cluster per job converts a **shared-machine problem into a shared-dependency problem**. It is a real improvement for memory, versions, restarts and blast radius, and it is not the total isolation teams often believe they are buying. Ask what the two teams still touch in common; whatever that list contains is what you did not fix. If the answer includes the storage layer, the metadata service and a machine quota, then you have changed which failures are possible, not removed the possibility of shared failure.

  • Why can a move to a cluster per job make a shared metadata service less stable rather than more?
    Because the load shape changes. One long-lived cluster consults it steadily; many short-lived ones consult it in a burst every time they start. The total may be similar while the peak is far higher, and it is the peak that trips rate limits and timeouts.
  • Two teams now have separate clusters but jobs still fail to start during the morning peak. What is the likely coupling?
    The quota or capacity held with whatever grants machines. Separate clusters still draw from one entitlement, so one team raising many clusters can leave nothing for the other. The contention moved from inside a pool to the act of obtaining machines.
  • Does a cluster per job protect a team from another team's bad machine image?
    No. If both teams' clusters are built from the same image or base configuration, a broken one reaches every newly raised cluster at once — and arguably faster than on a long-lived cluster, which is not rebuilt on every job.

saying these in an interview costs you the question

  • Claims a cluster per job gives complete isolation between teams
  • Forgets that shared storage and its rate limits remain common
  • Ignores the quota held with whatever grants the machines
  • Assumes more, shorter clusters are always gentler on shared services
  • Confuses isolation between jobs with per-process limits on one machine