skip to content

A job's worker processes each run several concurrent work slots — what do the slots inside one process share?

level: juniorimportance: must knowfreq 68%

answer

  1. two counts, not one total
  2. what is paid once per process
  3. one pool, one fate
  4. slots are co-resident, not isolated

basics

~20 s

Work slots in one worker process share that process's single memory pool, its local scratch space, its fixed start-up overhead and its fate: they compete for the same memory, and all of them die together when the process does.

solid answer

~50 s

A job's total capacity is granted as **worker processes** — the processes on cluster machines that run pieces of the job and own the memory those pieces use — and each process runs some number of **work slots**, the concurrent units of work it may run at once. The slots in one process are co-resident, not just co-counted. They draw from one memory pool, so two hungry units can exhaust it while the rest of the job's capacity sits idle elsewhere. They share the process's fixed costs — start-up, bookkeeping, connections, one copy of any data sent to every worker — which is why those costs are paid per process rather than per slot. They share one address space, so two of them can hand a record over in memory. And they share one lifetime: if the process ends, every slot in it ends too.

go deeper

for a junior

Recall that capacity arrives as processes, that several concurrent work slots can live inside one process, and that those slots share the process's memory and its lifetime.

for a middle

Explain which costs are paid once per process rather than once per slot — start-up, connections, one copy of data sent to every worker — and why that makes the two counts a genuine decision.

for a senior

Show that you reason about shared fate and shared memory: a single greedy unit can starve its co-resident slots, and one dying process takes every slot in it down at once.

for a principal

Treat the slots-per-process default as a platform decision the teams inherit: say who owns it, what it costs the fleet, and on what evidence a team is allowed to override it.

## Two numbers, not one A job asks whatever grants it machines for a total: so much processing capacity, so much memory. That total never arrives as one undivided pool. It arrives as **worker processes** — the processes on cluster machines that actually run pieces of the job and own the memory those pieces use — and each worker process runs some number of **work slots**, a work slot being one of the concurrent units of work a single process may run at the same time, all of them sharing that process's memory. Sixty-four-way concurrency is therefore two decisions and not one: how many processes, and how many slots inside each. Eight processes of eight slots, sixteen processes of four, and sixty-four processes of one all add up to sixty-four, and they do not behave alike. What an interviewer is checking is whether you know the slots are **co-resident**, not merely co-counted. ## What the slots in one process share - **One memory pool.** The process owns the memory its request was granted, and every slot draws from that one pool. Nothing is reserved per slot by default. Two greedy units of work running at the same moment can exhaust it while capacity elsewhere in the job goes unused. How that pool is then divided between the operators running inside it is a separate subject. - **One fate.** The slots are threads of work inside one operating-system process. If the process ends — a fault, a machine going away, the runtime giving up — every slot ends with it, and the units of work they held are lost rather than paused. - **One set of fixed per-process costs.** Starting the runtime, its internal bookkeeping, the connections it keeps open to the coordinating process and to peer workers, and any read-only data the job sent to every worker are paid once for the process, not once per slot. - **One local scratch area**, and one share of that machine's disk and network. Slots all writing locally at the same moment contend for the same device. - **One address space.** Two slots in one process can hand a record to each other as a reference in memory. Two slots in different processes cannot, even when the processes sit on the same machine. ## What they do not share A slot still carries its own unit of work and its own progress. One unit failing normally fails that unit alone; one unit finishing frees the slot for the next piece. What the slots do **not** get from each other is isolation: no private pool, no separate lifetime, no protection from a neighbour that takes the memory first. | Property | Per work slot | Per worker process | |---|---|---| | Unit of work in progress | yes | — | | Memory | drawn from the shared pool | granted once, owned here | | Start-up and runtime overhead | — | paid once | | Copy of data sent to all workers | — | held once | | Survives its process dying | no | — | | In-memory record handoff | possible between slots here | not across processes | ## Where designs differ This is a family whose members disagree, so state the variation rather than one engine's arrangement: - In the oldest model in this class, **each unit of work is its own operating-system process**. The slot count per process is effectively one, the fixed overhead is paid once per unit of work, and a read-only copy is held per unit rather than per process — much of why that model was expensive for short tasks. - Engines built around long-lived multi-threaded worker processes make the slot count a first-class sizing decision, and some let one process hold slots belonging to different parts of the same job at the same time. - A **managed compute service** — one you hand work to and are never shown a machine by — may expose neither number. You buy a capacity unit and the provider picks the shape. - How strictly memory is accounted also varies: some runtimes track usage per unit of work and fail the greediest one, others simply let the pool run out and fail whichever unit asked next. ## Why this is first-screen material A candidate who hears "sixty-four slots" as a single number will size a job by concurrency alone, be surprised when the same total behaves differently after a reshaping, and will not be able to explain why one process dying cost eight units of work instead of one. Everything else about worker shape — the overhead, the blast radius, the in-memory handoff — follows from knowing that a process is the thing that owns memory and a slot is a thing that borrows it.

  • Two slots in one process each need a large working set at the same moment — what is the failure mode?
    They draw on the same pool, so the process can run short even though the job's total is ample and other processes are idle. Which of the two is penalised depends on the runtime: some account memory per unit of work and fail the greediest, others let the pool run out and fail whichever unit asked next. Dividing that pool between the operators inside is a separate subject.
  • Does a lookup table sent to every worker cost memory once per slot or once per process?
    Once per process, wherever slots are threads sharing one address space — that is one of the main reasons for putting more slots in fewer processes. Where each unit of work is its own operating-system process, the same data is held once per unit, so the same job holds many more copies for the same total capacity.
  • If one unit of work throws an error, does its worker process die?
    Normally no: the unit fails, the slot is freed and the failure is reported for retry, while the other slots carry on. A process dies when something outside one unit's control ends it — the machine going away, the runtime being unable to continue, or memory exhaustion that no single unit can recover from. Only then do the co-resident slots go with it.

saying these in an interview costs you the question

  • Thinks each work slot gets its own private, fenced memory pool
  • Says adding slots to a process adds memory to that process
  • Believes the other slots keep running when their process dies
  • Treats total capacity as one number with no process count behind it
  • Assumes data sent to all workers is held once per slot everywhere