skip to content

Many Jobs, One Cluster

Several jobs competing for one pool: what is admitted and what waits, how shares and priorities split the machines, and the request so large the pool can never grant it.

on this pageshow

questions

4

Twelve jobs are submitted to one shared pool that is already fully occupied — what decides which starts next?

level: juniorimportance: must knowfreq 70%

answer

  1. the pool is finite, not elastic
  2. jobs wait, they do not shrink
  3. an admission queue of submitted jobs
  4. a share policy picks the next grant
  5. arrival order, weights, guaranteed minimum

basics

~20 s

Nothing starts until machines free up. Submitted jobs wait in an admission queue, and a share policy — strict order of arrival, weighted shares, or a guaranteed minimum per group — decides which waiting job the freed capacity goes to.

solid answer

~50 s

A shared pool holds a fixed number of machines, so a job asking for more capacity than is currently free is normally held rather than started small. Jobs that cannot be granted yet sit in an `admission queue of submitted jobs`, and whatever owns the pool hands freed machines out of that queue according to a **share policy**: strictly in order of arrival, in proportion to weights given to each team or workload class, or against a guaranteed minimum per group with the spare capacity lent out meanwhile. Which policy is in force is the real answer to 'when does my job run' — the size of the pool only sets the ceiling. Behaviour differs at the edges: some systems will admit a job with fewer worker processes than it asked for and let it pick up more later, others hold the whole request until it can be granted at once.

go deeper

for a junior

Recall that a shared pool is finite: a job that does not fit waits in a queue of submitted jobs rather than starting slowly. Be able to say that some rule decides who is next.

for a middle

Explain the three share-policy families and what each does to a freed machine, and say why a guarantee does not mean idle machines are reserved for a group with no work.

for a senior

Show that you know which policy your platform runs and what it does under contention, and that you separate waiting for a share from waiting for upstream data or for room on a single machine.

for a principal

Frame it as an allocation contract between teams: what floor each workload class is owed, whether continuous jobs may share a pool with finite ones at all, and who is allowed to change the weights.

## The pool is a fixed budget A **standing pool** is a cluster that is already up before any job arrives and stays up after each one ends, with many jobs submitted to it over a day. Its total capacity is whatever machines it holds, and that number does not change because someone pressed submit. **Submitting a job** means handing a packaged program together with its resource request to whatever owns those machines; the component that owns a pool of machines and grants a submitted job some share of them is the **cluster resource manager**. The request is a shape, not a vague amount. It names a number of **worker processes** — the processes on cluster machines that actually run pieces of the job and own the memory those pieces use — each of a stated size, plus the single **coordinating process** that plans the pieces, hands them out, tracks what finished and receives anything the program asks to bring back to one place. The share of the pool granted to this job is that shape multiplied out. When the pool's free capacity is smaller than the shape, most systems hold the job rather than start it thin, because starting it thin silently changes the job's own parallelism and is not a decision the platform usually makes on the author's behalf. Some do admit a job with fewer worker processes and let it acquire more later; that is a property of the platform, not a law. ## The admission queue of submitted jobs Jobs that cannot be granted yet sit in an **admission queue of submitted jobs**. Two things about it are routinely confused: - It queues **whole jobs**. Records waiting in a broker in front of the job, and pieces of work waiting for a free **work slot** — one of the concurrent units of work a single worker process may run at once, all of them sharing that process's memory — are two other queues entirely, with different owners and different remedies. - Position in it is not always the whole story. Only under a strict order-of-arrival rule does the earliest submission always win; under the other policies a later job belonging to an under-served group can be granted first. ## The share policy is the actual answer A **share policy** is the rule that decides how a finite pool is divided between competing jobs, and whether capacity is taken back from a job already running. Three families cover most of what you will meet: | Policy family | Who gets the freed capacity | Characteristic failure | |---|---|---| | Strict order of arrival | the job that has waited longest | one large job blocks everything behind it | | Weighted proportional shares | every group with work waiting, in proportion to its weight | nobody has a floor; all groups slow together under load | | Guaranteed minimum per group, spare lent out | a group currently below its guarantee | the guarantee is only as fast as the lending is undone | A guarantee is an entitlement, not a reservation of idle machines: a group with no work waiting is not holding anything, and its share is lent to whoever does have work. That is why the interesting question about this family is always *how* the entitlement is honoured when the group comes back — by waiting for borrowers to finish, or by reclaiming capacity from them. ## What varies between systems 1. **Partial admission.** Some platforms start a job below its requested width; others hold the entire shape until it fits at once. 2. **The coordinating process.** Where it runs is **coordinator placement** — inside the pool beside the workers, or outside it on the machine that submitted the job. When it runs inside, it is itself part of the grant and may be admitted before the workers it then waits for. 3. **Giving capacity back early.** Some jobs release worker processes they have stopped using while still running; a **continuous job** — one over input that never ends and therefore never reaches a last round of work — normally holds its share for its whole life, so a pool mixing continuous and **finite** jobs needs separate guaranteed shares or the finite work starves. ## What the admission queue is not - It is **not** the ordering of jobs relative to one another. Waiting because an upstream dataset is not ready yet is a dependency between jobs, decided before anything is submitted; waiting here means the program is ready and the machines are not. - It is **not** the packing of processes onto individual machines, nor per-process limits and the neighbour that crowds you out — that is a layer below, and a job that has been granted its share can still be slow for those reasons. - It is **not** a bill. Who pays for the pool, and what an idle pool costs, are separate questions from who is allowed to use it next.

  • What is the coordinating process doing while the job sits in the admission queue?
    It depends on coordinator placement. When the coordinating process runs outside the cluster on the submitting machine, it is already up and holding the submission open while it waits. When it runs inside the pool, it is part of the grant itself: on some platforms it is admitted first and then waits for its worker processes, on others nothing of the job exists until the whole shape can be granted.
  • Does a job ever give part of its share back to the pool before it finishes?
    Sometimes. A job that has stopped using some of its worker processes can hand them back mid-run on platforms that support a mid-run capacity change, which is what keeps a pool of finite jobs fluid. A continuous job never reaches a last round of work, so it normally holds its share until it is stopped — which is why mixing continuous and finite work on one pool needs guaranteed shares rather than order of arrival.

saying these in an interview costs you the question

  • Thinks a job always starts immediately, just more slowly, when the pool is full.
  • Assumes machines are added to the pool automatically whenever a job has to wait.
  • Believes the largest waiting job is always granted the freed capacity first.
  • Cannot name anything that decides the order beyond 'it queues'.
  • Confuses waiting for a share of the pool with waiting for upstream data to be ready.
  • Thinks a guaranteed minimum means machines are held idle for that group.
open as a page

A five-second query waits six hours behind one overnight job on a shared pool — why, and what fixes it?

level: middleimportance: should knowfreq 58%

basics

~20 s

Under strict order-of-arrival admission the overnight job holds the pool and the queue is blocked at its head, so short work inherits the long job's runtime. Weighted shares, a per-group guaranteed minimum, a ceiling on one job's share, or reclaiming capacity all break it.

open as a page

On a shared pool, the share policy takes capacity back from a job already running — what does that job lose?

level: seniorimportance: should knowfreq 48%

basics

~20 s

It loses the worker processes taken, every piece of work they were part-way through, and anything they still held that other workers had not yet consumed. The job itself survives; only what lived on the reclaimed capacity has to be produced again.

open as a page

A job that must have all 200 worker processes at once faces a pool ceiling of 150 — what happens?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

It waits forever. Under all-or-nothing placement a job starts only when every worker process it asked for is available together, so a request above the ceiling is not slow but impossible, and no amount of waiting changes it.

open as a page