A job that must have all 200 worker processes at once faces a pool ceiling of 150 — what happens?
answer
- unsatisfiable, not merely queued
- all of them at once, or none
- the wait never shrinks
- two partial holds can deadlock a pool
- shrink the request or raise the ceiling
basics
~20 sIt waits forever. Under all-or-nothing placement a job starts only when every worker process it asked for is available together, so a request above the ceiling is not slow but impossible, and no amount of waiting changes it.
solid answer
~50 sThis is not queueing, it is an unsatisfiable request. **All-or-nothing placement** means a job can only start once every worker process it asked for is available at the same moment; if the most the pool will ever grant one job is 150 and the job insists on 200, the condition can never be met while both numbers stand. The tell is that the job is still waiting while other jobs are admitted around it and capacity is visibly being handed out — ordinary queueing shortens as jobs ahead finish, this does not. A worse variant appears when two such jobs each accumulate a partial hold and neither can reach its full width, so the pool is fully allocated and nothing is running; policies avoid it by reserving atomically or by timing out a partial hold and releasing it. The fixes are to lower the request, raise the ceiling, or make the job able to start narrower.
go deeper
Recall that some jobs need every worker process at once, and that a request bigger than the pool will ever grant is refused by arithmetic, not delayed by traffic.
Explain how to distinguish an unsatisfiable request from ordinary queueing, and name the different ceilings a request can hit, including a process size no machine can host.
Diagnose the partial-hold stall — a pool fully allocated with nothing running — and pick between atomic reservation, a hold timeout and reshaping the job's request.
Decide whether workloads that demand simultaneity belong on a shared pool at all, and what a reserved window or a raised guarantee commits the organisation to.
## Why some jobs demand everything at once Most work does not need all its worker processes simultaneously. A job whose pieces are independent can run 200 pieces on 50 worker processes across four rounds of work and produce the same answer more slowly. Some jobs cannot. **All-or-nothing placement** describes a job that can only start once every worker process it asked for is available at the same moment, and otherwise waits. The usual reasons: - Every part of the job exchanges records with every other part while running, so a missing participant means the others block forever rather than run slowly. - The job is continuous — its input never ends — and its whole graph must be resident and connected to make any progress at all. - The work is one indivisible computation whose participants must be alive together, so a partial start produces nothing. When the platform knows this, it admits the job as a unit. When it does not, the job is admitted piecemeal and hangs at the first exchange, which is the same failure wearing a worse costume. ## Impossible is not slow A request above the largest share the pool will ever grant a single job is **unsatisfiable**, and the distinction from ordinary queueing is the whole of the diagnosis: | Symptom | Ordinary queueing | Unsatisfiable request | |---|---|---| | Wait time | shrinks as jobs ahead finish | does not shrink, ever | | During an idle pool | job starts | job still waits | | Jobs submitted later | wait behind it | are admitted around it | | Fix | wait, or change the policy | change the request or the ceiling | The ceiling in question can come from several places and it is worth finding out which: the pool's total capacity; a cap on the fraction of the pool one job may hold; a group's guaranteed share when the policy will not lend beyond it; or a per-machine constraint, where a worker process of the requested size does not fit any single machine so no count of them can ever be placed. The last one is the cruellest, because the requested total may be modest. ## The partial-hold stall Two all-or-nothing jobs on the same pool can fail in a way neither would alone. Each is granted part of what it asked for, each waits for the rest, and neither will release what it holds because releasing means abandoning its place. The pool reads as fully allocated, and nothing at all is running. Policies defeat this in one of three ways: 1. **Reserve atomically.** Nothing is handed to the job until the whole shape can be granted, so a partial hold never exists. 2. **Time out a partial hold.** A job that has not reached its full width within some interval releases everything and requeues, which breaks the deadlock at the price of restarting the wait. 3. **Order the claimants.** Only one all-or-nothing job at a time may accumulate a hold, so the second never starts collecting. If you are diagnosing a pool where nothing is running and everything is allocated, this is the first shape to check for. ## What to actually do - **Lower the request.** Most often the 200 was never a requirement — it was a throughput target. If the job can run in more rounds of work on fewer processes, it should. - **Make the shape fit.** Fewer, larger worker processes or more, smaller ones can be the same total capacity while one of the two fits inside the ceiling and the other does not. - **Raise the ceiling for this workload.** Give the job a group whose guarantee is large enough, and accept what that commits. - **Reserve a window.** Run the job when the pool is not contended, with the whole pool promised to it — the standard answer for one large periodic job among many small ones. - **Do not simply resubmit.** Repeated submission of an unsatisfiable request fills the queue and, on policies that grant partial holds, actively makes the stall above more likely. ## What varies Whether a job needs all-or-nothing treatment at all is a property of the job's structure, not of the engine, but engines differ in whether they can express the requirement and in what they do without it. Some platforms accept an explicit minimum width and start once it is met; some accept only an exact shape; some have no notion of it, so a job that truly needs simultaneity has to have that enforced by the way its resources are requested. Before assuming a job must be placed all at once, check whether it genuinely blocks without every participant or whether it would merely be slower — the second case is not this problem.
- How do you tell an unsatisfiable request from a job that is merely far back in the queue?Watch what happens around it. A queued job's wait shortens as the jobs ahead of it finish, and it starts when the pool empties. An unsatisfiable one keeps waiting while later submissions are admitted past it and while capacity is visibly free, because the condition it needs is never met rather than not yet met.
- Why can repeatedly resubmitting the job make things worse?On a policy that grants partial holds, each submission starts accumulating worker processes it will never complete, so the pool ends up allocated to jobs that cannot run. Even where holds are atomic, the queue fills with requests that will never be granted, which delays decisions about everything behind them.
- Is there a fix that does not involve changing the request or the ceiling?Yes: reserve a window. Promise the whole pool to the job at a time when nothing else is contending, so the ceiling on a single job's share does not apply or is temporarily raised. It works well for one large periodic job among many small ones, and badly for anything that must run on demand.
A party of eight arrives at a restaurant whose largest table seats six and which will not join tables together. They can wait all evening: pairs and fours keep being seated around them, and no amount of patience produces a table for eight. The host is not being unfair and the queue is not stuck — the request cannot be satisfied by that room. Either the party splits across two tables and accepts eating separately, or somebody changes the furniture.
saying these in an interview costs you the question
- Says the job will start eventually if the pool is quiet enough.
- Advises resubmitting repeatedly until the request happens to fit.
- Assumes every job can be started narrower and simply run more rounds.
- Overlooks a worker-process size that no single machine can host, whatever the count.
- Treats a fully allocated pool with nothing running as normal contention.
- Thinks the ceiling is always the pool's total, never a per-group or per-job cap.