A five-second query waits six hours behind one overnight job on a shared pool — why, and what fixes it?
answer
- latency of the front, not your own
- head-of-line blocking on admission
- it holds its share until it ends
- shares or a floor, not arrival order
- reclaiming is the only in-flight remedy
basics
~20 sUnder strict order-of-arrival admission the overnight job holds the pool and the queue is blocked at its head, so short work inherits the long job's runtime. Weighted shares, a per-group guaranteed minimum, a ceiling on one job's share, or reclaiming capacity all break it.
solid answer
~50 sThe waiting time of a short job under strict order of arrival is the runtime of whatever is in front of it, not its own size — the classic head-of-line block, where the item at the front of a queue holds up everything behind it regardless of cost. Two things make it bite on a shared pool: the long job was granted a large share, and it may not release any of that share until it ends. The remedies all work by removing one of those two conditions: divide freed capacity by weighted share rather than by arrival so a newly submitted job gets some immediately; give interactive work a guaranteed minimum of its own so it never queues behind the batch group at all; cap how much of the pool any single job may hold, leaving a permanent gap; or let the policy reclaim capacity from the running job. Which of these your platform supports is what decides the real fix.
go deeper
Recall that under order-of-arrival admission your waiting time is set by the job in front of you, not by how small your own job is.
Explain head-of-line blocking and both conditions behind it — the ordering, and the long job holding its share — and name a remedy for each.
Demonstrate the diagnosis first: prove the job was queued rather than slow, then choose between a floor for interactive work and reclaiming capacity, and say what each costs the batch group.
Decide whether interactive and long-running work belong on one pool at all, what latency the organisation is buying with a permanent floor, and who owns the weights when two teams disagree.
## What is actually blocking Two separate facts combine here, and candidates who name only one miss the fix. The first is **head-of-line blocking**: when a queue is served strictly in order, the item at the front holds up everything behind it however cheap those items are. A queue of submitted jobs served in strict order of arrival gives the five-second query a waiting time equal to the remaining runtime of whatever is in front of it. Its own cost is irrelevant to its latency. The second is that the overnight job holds a large **share of the pool granted to it** — the worker processes it was admitted with, where a worker process is the process on a cluster machine that runs pieces of the job and owns the memory those pieces use. Even a policy that would happily grant the short query a slice has nothing to grant while every machine is inside another job's share. ## Why the long job may not give anything back It is tempting to say the long job releases its machines when it finishes, and that is the first sentence to make precise, because engines in this class differ: - Some jobs release worker processes they have stopped using while still running, so a pool of these stays fluid and a short job only waits for the next release rather than for the end of the run. - Others hold the shape they were admitted with from start to finish, whether or not every worker process is busy — so the wait really is the job's full runtime. - A **continuous job** — a job over input that never ends, and which therefore never reaches a last **round of work** — does not finish at all. Under order-of-arrival admission alone, anything behind one waits indefinitely, which is the version of this problem that actually pages people. ## Four ways a policy unblocks the short job 1. **Serve by share rather than by arrival.** Under weighted proportional shares, freed capacity is divided among every group that has work waiting, so a newly submitted query starts collecting machines as soon as any are released rather than at the back of a line. 2. **Give interactive work its own guaranteed minimum.** Put short work in a group with a floor of its own; the batch group may borrow that floor while it is unused, but the short job is then owed it back rather than queued behind the borrower. 3. **Cap the share any one job may hold.** If no single job may hold more than, say, two thirds of the pool, a gap always exists for something small — at the price of a long job that can never use an otherwise idle cluster fully. 4. **Reclaim capacity from the running job.** Let the share policy take worker processes back from the overnight job to satisfy the owed minimum. This is the only remedy that acts on a job already at full width, and it is the only one that destroys work — everything those worker processes were doing has to be redone. A fifth trick sits between admission and reclaiming: allow a small job to be started ahead of a large waiting one whenever it fits in the space available and will finish before the large one's turn comes. It keeps the pool busy without breaking the promise made to the large job, and it only works when the running time of the small job is known or bounded. | Remedy | Acts on | Cost it introduces | |---|---|---| | Weighted shares | future releases | no firm floor for anyone; all groups slow together | | Guaranteed minimum per group | future releases | capacity is committed to a group even when it is quiet | | Ceiling on one job's share | admission | a single large job can never use the whole pool | | Reclaiming from a running job | work already in flight | destroys and redoes the work on the reclaimed capacity | ## Getting the diagnosis right Before reaching for a policy change, confirm the query is actually queued rather than slow: a job that has been granted its machines and is running badly is a different problem with different owners, and so is a query whose input was not ready. The signature of this one is that the short job has no worker processes at all for most of its wall-clock time, and starts promptly the moment the long job's share is released. The other thing to keep separate: this is contention for machines between jobs that are all ready to run. A short job that waits because the dataset it reads is produced by the overnight job is not queued at all — it is ordered behind it by a dependency, and no share policy will help.
- Why does adding machines to the pool often fail to fix this?Because the block is in the ordering, not the total. Under strict order of arrival the next submitted large job takes the enlarged pool too, and the short query is behind it again. Extra machines only help if the policy hands some of them to the short job — which is a policy change, not a capacity change.
- What breaks if you simply cap every job at a small share of the pool?Long jobs lose the ability to use an idle cluster. A cap guarantees a gap for short work at all times, including nights when nothing else is submitted, so the overnight job runs at a fraction of the width it could have had and its runtime grows. Caps that apply only under contention avoid that, at the cost of needing capacity to be reclaimed when contention appears.
- How would you tell this apart from the short query simply being slow?Look at when the job held worker processes. Queued work holds none for most of its wall-clock time and then runs in seconds; slow work holds its machines throughout. The two have completely different owners, and the second is not a pool-sharing question at all.
saying these in an interview costs you the question
- Says the short query must be badly written, without checking whether it ever got machines.
- Assumes the long job releases capacity steadily as it progresses, on every platform.
- Proposes only 'add more machines', which delays the same block rather than removing it.
- Treats a continuous job as something that will eventually finish and free the pool.
- Confuses being queued for machines with being ordered behind an upstream dataset.
- Thinks reclaiming capacity from the running job is free because the work resumes where it stopped.