skip to content

Machines for the Job

Where the machines a job runs on come from and how large they are: a pool that is always up, a cluster raised for one job, or a service that shows you none - and what each costs.

on this pageshow

explore

questions

23

A job needs machines before it can run. What are the three ways a platform supplies them, and who sizes them under each?

level: juniorimportance: must knowfreq 68%

answer

  1. three ways to be given machines
  2. already up, raised, or invisible
  3. who sized it, and when
  4. lifetime of the machines versus the job
  5. sizing shifts submitter to provider

basics

~20 s

A standing pool that is already up and shared by many jobs; a cluster raised for one job and torn down with it; or a managed compute service that shows no machine at all. Sizing moves from submitter to provider across the three.

solid answer

~50 s

There are three supply models. A **standing pool** is a cluster that is up before any particular job arrives and stays up after it ends, with many jobs submitted to it; whoever owns the pool sized it once, for aggregate demand, and keeps it running. A **per-job cluster** is machines raised for one job and torn down when that job finishes, so its lifetime is the job's lifetime and nothing it held locally outlives it. A **managed compute service** is one you hand work to and are never shown a machine by, with sizing, starting and stopping owned entirely by the provider. Sizing follows that order: the pool owner sizes the first, the submission sizes the second — sometimes as an exact count, sometimes as a hint the platform turns into one — and the third often takes no machine sizing at all.

go deeper

for a junior

Be able to name the three and say one sentence about each: already up and shared, raised for one job and torn down, or no machine shown at all. That alone answers the screening version of this question.

for a middle

Explain who sizes the machines under each model and when that decision is made, and note that a cluster raised per job may be sized from a hint rather than an exact count.

for a senior

Talk about lifecycle rather than sizing: what a long-lived pool accumulates across jobs, what a per-job cluster gives up by keeping nothing, and which decisions disappear when no machine is exposed.

for a principal

Frame it as who the organisation wants owning capacity and upgrades. Each model you operate adds an access path, a failure mode and a sizing conversation, so the number of models in use is itself a decision.

## What is settled before the program matters A distributed job is a program plus a request for machines. Nothing in the program runs until something has granted it some, and the arrangement that grants them decides several things the program itself cannot: how long you wait to start, who chose how much you got, who may take it away, and what is left behind when you are done. Two terms first, because the rest depends on them. - A **worker process** is the process on a cluster machine that actually runs pieces of the job and owns the memory those pieces use. Several worker processes can sit on one machine. - The **coordinating process** is the single process that plans the pieces, hands them out, tracks what finished, and receives anything the program asks to bring back to one place. Every supply model produces both. They differ only in where those processes come from, who decided how many there are, and how long the machines under them live. ## The three models 1. **A standing pool.** A cluster that is already up before a job arrives and stays up after it ends. Jobs are submitted to it all day. Someone — usually a platform team — chose its size once, against the aggregate demand of everyone who submits to it, and owns keeping it alive, patched and reachable. 2. **A per-job cluster.** Machines raised for one job and torn down when it finishes. Its lifetime is exactly the job's lifetime. Nothing it cached in memory or wrote to a local disk outlives it, because the disks go too. 3. **A managed compute service.** You hand work to a service and are never shown a machine. There are machines, of course, but none of them is yours to name, size, log into or keep. Starting, stopping and sizing are the provider's, and so is the decision about what to do when one of them dies. ## Who owns sizing This is the axis interviewers actually probe, because it is where candidates over-generalise from the one platform they have used. Under a standing pool, sizing is a **capacity decision made in advance** by the pool's owner, about everybody's work at once. An individual job asks for a share of what is already there. Under a per-job cluster, sizing is a **per-submission decision**, and platforms differ in how literal it is. Some take an exact worker-process count and machine size from the submission. Others take a hint — a workload size, a family of machine, a maximum spend — and derive the count themselves. Either way the decision is made once, at submission, and by default lasts as long as the job. Under a managed compute service, there is frequently **no machine sizing to give**. Where the service accepts anything at all it is coarse: a capacity tier, a concurrency limit, a ceiling on spend. "You size the cluster" is exactly the sort of sentence that is true of one product and false of the next, so say which model you mean before you say who sizes it. ## Side by side | | standing pool | per-job cluster | managed compute service | |---|---|---|---| | exists before the job | yes | no | no machine is exposed | | sized by | the pool owner, once, for everyone | the submission, exactly or from a hint | the provider | | torn down by | nobody, until the pool is retired | job completion | the provider | | caches and local files reusable later | possibly, by a later job | no | no machine to hold them | | one job's memory pressure reaches | other jobs on the pool | only itself | the provider's concern | | runtime and library versions | one set, shared | per job | the provider's, within limits | ## Lifecycle is the real difference Sizing is the visible difference; lifecycle is the consequential one. A standing pool is a long-lived thing with an owner, an upgrade schedule and state that accumulates across jobs. A per-job cluster has no life outside the job, which is why it gives each job its own runtime versions and its own blast radius and why it can reuse nothing. A managed compute service takes the whole lifecycle off your books, and with it the ability to make any decision that depends on a machine existing. ## What the choice does not decide It does not decide what your program computes, nor how the work is divided into pieces. It also does not, by itself, settle what happens when several jobs want the same finite pool at the same moment — that contention is its own subject. Keep the answer on supply: who owns the machines, who sized them, and how long they live.

  • Under a cluster raised for one job, who actually decides how many worker processes it gets?
    It depends on the platform. Some take an exact count and machine size from the submission; others take a hint — a data volume, a machine family, a spend ceiling — and derive the count themselves. The constant is that the decision is made once, at submission, not by the program while it runs.
  • Does a managed compute service remove capacity decisions entirely?
    It removes machine-level ones. What usually remains is coarse: a capacity tier, a limit on how much work runs at once, or a spend ceiling. Those are the only dials, so capacity planning becomes a question about concurrency and budget rather than about worker processes and machines.
  • Can a job outlive the machines it ran on?
    No — the job ends when its processes end. The relationship runs the other way: a standing pool outlives every job submitted to it, a cluster raised for one job dies with it, and under a managed service the question is not answerable because no machine is exposed to attach a lifetime to.

A standing pool is an office the company leases all year: it is there whether you use it today or not, and someone else picked its size. A per-job cluster is a van hired for one move — exactly as big as you asked for, and empty of your things the moment you return it. A managed compute service is posting a parcel: you never see the vehicle, and nobody asks you how large it should be.

saying these in an interview costs you the question

  • Assumes every supply model lets you pick a machine size
  • Thinks a cluster is always up somewhere waiting for work
  • Believes a cluster raised per job keeps its cache for the next run
  • Cannot say who starts and stops the machines in each model
  • Treats the three models as billing labels with no lifecycle difference
open as a page

A job over input that ends gains extra worker processes halfway through its run. Where can that new capacity first do useful work?

level: juniorimportance: must knowfreq 58%

basics

~20 s

On pieces of work that have not started yet, normally from the next round of work onward. A piece already running on another worker process is not cut in half and shared, and finished work is not redone to spread it more evenly.

open as a page

The provider takes back a discounted machine in the middle of a run. Besides the piece it was computing, what does the job lose?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Everything that machine held goes at once: every worker process on it, the pieces they were running, anything cached there, and intermediate output produced there that other workers had not yet fetched. Finished work can therefore have to be produced again.

open as a page

A job's worker processes each run several concurrent work slots — what do the slots inside one process share?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Work slots in one worker process share that process's single memory pool, its local scratch space, its fixed start-up overhead and its fate: they compete for the same memory, and all of them die together when the process does.

open as a page

Twelve jobs are submitted to one shared pool that is already fully occupied — what decides which starts next?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Nothing starts until machines free up. Submitted jobs wait in an admission queue, and a share policy — strict order of arrival, weighted shares, or a guaranteed minimum per group — decides which waiting job the freed capacity goes to.

open as a page

In a distributed job, what does the single process that is not a worker do, and what does it never do?

level: juniorimportance: must knowfreq 72%

basics

~20 s

One process, the coordinating process, turns the program into pieces of work, hands them to worker processes, tracks which finished, and receives anything the program asks to bring back. It runs no piece of the data itself.

open as a page

Why does changing the worker count of a job over endless input cost a pause, when the same change costs a job with an end almost nothing?

level: middleimportance: must knowfreq 55%

basics

~20 s

A job over endless input keeps per-key state, and which worker process owns a given key follows how many worker processes there are. Change that count and the retained state has to move to its new owners before any record can be processed correctly.

open as a page

A job's worker processes can be few and large or many and small at one fixed total — what changes?

level: middleimportance: must knowfreq 72%

basics

~20 s

Fewer, larger worker processes pay the fixed per-process overhead fewer times, keep more handoffs inside one address space and give any single unit of work a bigger pool to draw on. They also lose more when one dies and are harder to place.

open as a page

A job filters a billion rows down to four million and returns them all to the submitting program — what fails first?

level: middleimportance: must knowfreq 66%

basics

~20 s

The coordinating process runs out of memory. Every returned record leaves the workers and lands in that one process's heap on one machine, so the cost follows the result size, not the cluster size — adding worker processes does not help.

open as a page

A job ran on machines raised only for it and cached a working set on them. What survives after those machines are torn down?

level: middleimportance: should knowfreq 52%

basics

~20 s

Only what was written to storage outside those machines. Caches held in worker-process memory, intermediate output written to local disk, temporary directories and logs left on the machines all disappear with them, so the next run starts cold.

open as a page

Two worker processes of one job sit on the same machine — does moving records between them still cross a network path?

level: middleimportance: should knowfreq 44%

basics

~20 s

Mostly yes. Two worker processes have separate address spaces, so records are encoded into bytes at one end and decoded at the other even on one machine. Sharing a machine skips the wire, not the encoding — only slots inside one process avoid it.

open as a page

A five-second query waits six hours behind one overnight job on a shared pool — why, and what fixes it?

level: middleimportance: should knowfreq 58%

basics

~20 s

Under strict order-of-arrival admission the overnight job holds the pool and the queue is blocked at its head, so short work inherits the long job's runtime. Weighted shares, a per-group guaranteed minimum, a ceiling on one job's share, or reclaiming capacity all break it.

open as a page

Should the coordinating process run inside the cluster beside the workers or on the machine that submitted the job?

level: middleimportance: should knowfreq 48%

basics

~20 s

Inside the cluster, the coordinating process sits close to the workers and outlives the submitting machine. Outside it, the session stays interactive but every scheduling message crosses a slower link, and closing the submitting machine ends the job.

open as a page

Two teams share one cluster that is always up and move to a cluster per job. Which failures does that isolate, and which stay shared?

level: seniorimportance: should knowfreq 48%

basics

~20 s

It isolates everything that lived on the machines: memory exhaustion, a crashed coordinating process, runtime and library versions, restarts and upgrades. It does not isolate shared storage, a shared metadata service, the quota with whoever grants machines, or the credentials both teams use.

open as a page

A running job hands worker processes back to free capacity. Why is releasing one riskier than acquiring one?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A worker process with no piece in flight can still be holding output that other worker processes have not fetched yet. Release it and that output goes with it, so the pieces that produced it may have to run again — often costing more than the release saved.

open as a page

Fifty minutes into a sixty-minute run the provider takes back one machine, and work that finished forty minutes ago is done again - why?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Because that machine held intermediate output from earlier rounds that later pieces were still fetching. Losing it turns completed pieces back into work to be done, and the later the withdrawal, the more such output has accumulated.

open as a page

Late in a long job, one worker process holding eight work slots dies — what is lost with it, and how does the process's size set that?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Everything that lived in that process goes: all eight units of work in flight, from the beginning rather than from where they stopped, plus any intermediate output and cached data held only there. The process's size is the multiplier on all of it.

open as a page

On a shared pool, the share policy takes capacity back from a job already running — what does that job lose?

level: seniorimportance: should knowfreq 48%

basics

~20 s

It loses the worker processes taken, every piece of work they were part-way through, and anything they still held that other workers had not yet consumed. The job itself survives; only what lived on the reclaimed capacity has to be produced again.

open as a page

A long-running job loses the coordinating process on one machine and a worker process on another — how do the two losses differ?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Losing a worker process costs the pieces it was running and the intermediate output it still held; the rest of the job continues. Losing the coordinating process leaves nothing planning, handing out or tracking work, so the job stops.

open as a page

Eight teams submit a hundred jobs a day, from seconds long to hours long. How would you place those workloads across the three supply models?

level: principalimportance: should knowfreq 42%

basics

~20 s

Place by wait tolerance and lifecycle ownership, not by team. Short interactive work needs machines already up; long scheduled jobs earn their own cluster and their own runtime versions; work with no operator behind it fits a service that shows no machine. Then limit how many models you operate.

open as a page

Your platform can change the worker count of a job over endless input on demand. What makes an aggressive change policy cost more than it saves?

level: principalimportance: should knowfreq 32%

basics

~20 s

Every change stops the job while its retained state is relocated, and that cost follows how much the job holds rather than how much capacity moves. A policy that reacts faster than the pause is long spends its savings on pauses and provokes further changes.

open as a page

A machine the provider may take back costs far less per hour. For which shapes of distributed job does that saving fail to arrive?

level: principalimportance: should knowfreq 35%

basics

~20 s

For jobs that redo more than they save: long runs whose later rounds depend on output held only on those machines, and latency-bound continuous jobs, where each withdrawal buys a pause, a rewind and a backlog that a job with no spare throughput cannot clear.

open as a page

A job that must have all 200 worker processes at once faces a pool ceiling of 150 — what happens?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

It waits forever. Under all-or-nothing placement a job starts only when every worker process it asked for is available together, so a request above the ceiling is not slow but impossible, and no amount of waiting changes it.

open as a page