A job needs machines before it can run. What are the three ways a platform supplies them, and who sizes them under each?
answer
- three ways to be given machines
- already up, raised, or invisible
- who sized it, and when
- lifetime of the machines versus the job
- sizing shifts submitter to provider
basics
~20 sA standing pool that is already up and shared by many jobs; a cluster raised for one job and torn down with it; or a managed compute service that shows no machine at all. Sizing moves from submitter to provider across the three.
solid answer
~50 sThere are three supply models. A **standing pool** is a cluster that is up before any particular job arrives and stays up after it ends, with many jobs submitted to it; whoever owns the pool sized it once, for aggregate demand, and keeps it running. A **per-job cluster** is machines raised for one job and torn down when that job finishes, so its lifetime is the job's lifetime and nothing it held locally outlives it. A **managed compute service** is one you hand work to and are never shown a machine by, with sizing, starting and stopping owned entirely by the provider. Sizing follows that order: the pool owner sizes the first, the submission sizes the second — sometimes as an exact count, sometimes as a hint the platform turns into one — and the third often takes no machine sizing at all.
go deeper
Be able to name the three and say one sentence about each: already up and shared, raised for one job and torn down, or no machine shown at all. That alone answers the screening version of this question.
Explain who sizes the machines under each model and when that decision is made, and note that a cluster raised per job may be sized from a hint rather than an exact count.
Talk about lifecycle rather than sizing: what a long-lived pool accumulates across jobs, what a per-job cluster gives up by keeping nothing, and which decisions disappear when no machine is exposed.
Frame it as who the organisation wants owning capacity and upgrades. Each model you operate adds an access path, a failure mode and a sizing conversation, so the number of models in use is itself a decision.
## What is settled before the program matters A distributed job is a program plus a request for machines. Nothing in the program runs until something has granted it some, and the arrangement that grants them decides several things the program itself cannot: how long you wait to start, who chose how much you got, who may take it away, and what is left behind when you are done. Two terms first, because the rest depends on them. - A **worker process** is the process on a cluster machine that actually runs pieces of the job and owns the memory those pieces use. Several worker processes can sit on one machine. - The **coordinating process** is the single process that plans the pieces, hands them out, tracks what finished, and receives anything the program asks to bring back to one place. Every supply model produces both. They differ only in where those processes come from, who decided how many there are, and how long the machines under them live. ## The three models 1. **A standing pool.** A cluster that is already up before a job arrives and stays up after it ends. Jobs are submitted to it all day. Someone — usually a platform team — chose its size once, against the aggregate demand of everyone who submits to it, and owns keeping it alive, patched and reachable. 2. **A per-job cluster.** Machines raised for one job and torn down when it finishes. Its lifetime is exactly the job's lifetime. Nothing it cached in memory or wrote to a local disk outlives it, because the disks go too. 3. **A managed compute service.** You hand work to a service and are never shown a machine. There are machines, of course, but none of them is yours to name, size, log into or keep. Starting, stopping and sizing are the provider's, and so is the decision about what to do when one of them dies. ## Who owns sizing This is the axis interviewers actually probe, because it is where candidates over-generalise from the one platform they have used. Under a standing pool, sizing is a **capacity decision made in advance** by the pool's owner, about everybody's work at once. An individual job asks for a share of what is already there. Under a per-job cluster, sizing is a **per-submission decision**, and platforms differ in how literal it is. Some take an exact worker-process count and machine size from the submission. Others take a hint — a workload size, a family of machine, a maximum spend — and derive the count themselves. Either way the decision is made once, at submission, and by default lasts as long as the job. Under a managed compute service, there is frequently **no machine sizing to give**. Where the service accepts anything at all it is coarse: a capacity tier, a concurrency limit, a ceiling on spend. "You size the cluster" is exactly the sort of sentence that is true of one product and false of the next, so say which model you mean before you say who sizes it. ## Side by side | | standing pool | per-job cluster | managed compute service | |---|---|---|---| | exists before the job | yes | no | no machine is exposed | | sized by | the pool owner, once, for everyone | the submission, exactly or from a hint | the provider | | torn down by | nobody, until the pool is retired | job completion | the provider | | caches and local files reusable later | possibly, by a later job | no | no machine to hold them | | one job's memory pressure reaches | other jobs on the pool | only itself | the provider's concern | | runtime and library versions | one set, shared | per job | the provider's, within limits | ## Lifecycle is the real difference Sizing is the visible difference; lifecycle is the consequential one. A standing pool is a long-lived thing with an owner, an upgrade schedule and state that accumulates across jobs. A per-job cluster has no life outside the job, which is why it gives each job its own runtime versions and its own blast radius and why it can reuse nothing. A managed compute service takes the whole lifecycle off your books, and with it the ability to make any decision that depends on a machine existing. ## What the choice does not decide It does not decide what your program computes, nor how the work is divided into pieces. It also does not, by itself, settle what happens when several jobs want the same finite pool at the same moment — that contention is its own subject. Keep the answer on supply: who owns the machines, who sized them, and how long they live.
- Under a cluster raised for one job, who actually decides how many worker processes it gets?It depends on the platform. Some take an exact count and machine size from the submission; others take a hint — a data volume, a machine family, a spend ceiling — and derive the count themselves. The constant is that the decision is made once, at submission, not by the program while it runs.
- Does a managed compute service remove capacity decisions entirely?It removes machine-level ones. What usually remains is coarse: a capacity tier, a limit on how much work runs at once, or a spend ceiling. Those are the only dials, so capacity planning becomes a question about concurrency and budget rather than about worker processes and machines.
- Can a job outlive the machines it ran on?No — the job ends when its processes end. The relationship runs the other way: a standing pool outlives every job submitted to it, a cluster raised for one job dies with it, and under a managed service the question is not answerable because no machine is exposed to attach a lifetime to.
A standing pool is an office the company leases all year: it is there whether you use it today or not, and someone else picked its size. A per-job cluster is a van hired for one move — exactly as big as you asked for, and empty of your things the moment you return it. A managed compute service is posting a parcel: you never see the vehicle, and nobody asks you how large it should be.
saying these in an interview costs you the question
- Assumes every supply model lets you pick a machine size
- Thinks a cluster is always up somewhere waiting for work
- Believes a cluster raised per job keeps its cache for the next run
- Cannot say who starts and stops the machines in each model
- Treats the three models as billing labels with no lifecycle difference