On a shared pool, the share policy takes capacity back from a job already running — what does that job lose?
answer
- policy takes it back, not the provider
- pieces in flight plus unconsumed output
- the job survives, the workers do not
- cost depends where intermediates live
- warn-then-take, or wait for a boundary
basics
~20 sIt loses the worker processes taken, every piece of work they were part-way through, and anything they still held that other workers had not yet consumed. The job itself survives; only what lived on the reclaimed capacity has to be produced again.
solid answer
~50 sReclaiming is an allocation decision with an execution bill. The worker processes taken away stop, so the job loses the pieces they were running and any intermediate results they were still holding for other workers to pick up — the second of those is the expensive part, because it can force work that had already been marked finished to be done again. What that costs differs sharply by engine class: a runtime that materialises intermediate output locally and lets consumers fetch it later loses that output with the process unless something separate keeps serving it, while a record-at-a-time runtime with per-key state pays instead by rolling the affected work back to its last saved recovery point. Policies soften it in different ways — no reclaiming at all, reclaiming only at the boundary between rounds of work, or a warning interval before the capacity is taken regardless.
go deeper
Recall that losing worker processes does not end the job: the pieces they were running are attempted again elsewhere, and the coordinating process carries on.
Explain why unconsumed intermediate output is the expensive loss rather than the pieces in flight, and describe the boundary between rounds as the cheap moment to take capacity.
Show that you know which design your engine uses and therefore what a reclaim actually costs it, and that you would measure destroyed machine-hours against the latency it buys another team.
Set the policy: which workload classes may be reclaimed from at all, whether the organisation buys a warning interval or lossless waiting, and what the destroyed work is worth against the guarantee it funds.
## Two different actors, one English word Capacity can be taken away from a running job by two completely different parties, and conflating them produces the wrong remedy every time. - **The share policy reclaims capacity from a running job.** This happens inside your organisation, to honour a share someone else is owed. It is a policy you chose and can change, and it is the subject here. - **The provider reclaims a discounted machine.** That is an infrastructure bargain — a cheap machine the provider may take back at short notice for reasons that have nothing to do with your job — and it is a different decision with a different warning and a different remedy. ## What actually leaves with the capacity When worker processes are taken from a running job, the job loses exactly three things: - **The pieces of work in flight on them.** Anything part-way through is abandoned; a piece is the unit that has to be attempted again. - **Whatever those processes were still holding for others.** Intermediate results produced earlier and not yet consumed are the costly loss, because they represent work the job had already finished and counted. - **Anything cached or written locally there.** A result deliberately kept in a worker's memory to be reused later is gone, and the next use of it pays full price. The job as a whole does **not** end. The single **coordinating process** — the process that plans the pieces, hands them out and tracks what finished — is a different process on a different footing, and losing worker processes is not losing it. Whether the job then reattempts just the lost pieces or restarts more broadly is a property of the engine's recovery model rather than of the share policy, and the policy has no say in it. ## The bill depends on what the engine keeps where This is where an answer built from one engine goes wrong, because the three common designs in this class pay in different currencies. | Engine design | What the reclaimed capacity was holding | What the job pays | |---|---|---| | Finite job that materialises intermediate output locally for later fetch | completed output other workers had not fetched | redoing finished work, sometimes a whole earlier round of it | | Record-at-a-time runtime with per-key state | live state and records in flight | rolling affected work back to its last saved recovery point and replaying from there | | Two-phase disk-to-disk model writing each phase out durably | only the current attempt | reattempting that unit, with earlier phases still readable | One mitigation is worth naming because it changes the arithmetic of the first row: a separate process that keeps serving a finished worker's intermediate output after that worker is gone. Where such a thing exists, losing the worker process no longer by itself loses the output, and reclaiming becomes far cheaper. Where it does not, taking capacity back late in a long job can cost more than the machines were worth. ## How policies soften it Policies sit on a spectrum, and knowing where yours sits is most of the operational answer: 1. **Never reclaim.** An owed share is honoured only as running jobs release capacity. Simple and lossless, but the guarantee may take hours to materialise — or never, behind a continuous job. 2. **Reclaim only at a boundary between rounds of work**, when no piece is half-finished. The cheapest point to take capacity, at the price of waiting for that moment to arrive. 3. **Warn, then take.** The job is told capacity is going and given an interval to stop cleanly, save what it must and hand the processes back; anything still running when the interval expires is killed anyway. 4. **Take immediately.** Lowest latency for the owed share, highest destroyed work. ## Making a job cheap to reclaim from The properties that make reclaiming survivable are the same ones that make any capacity loss survivable, and a team that runs on a reclaiming pool should design for them deliberately: - Keep the unit of work small, so an abandoned piece is a small loss. - Prefer work whose intermediate output survives the process that made it, or accept that a late reclaim can cascade. - Save recovery points often enough that a rollback is cheap, and understand that this costs throughput continuously in exchange for a cheaper worst case. - Know which of your jobs must never be reclaimed from, and put those in a group whose share is not lent out. The judgment an interviewer is listening for is that reclaiming trades a **latency** improvement for somebody else against **wasted machine-hours** for you, and that the trade is only sound when the work destroyed is smaller than the wait it saves.
- Why can reclaiming capacity late in a long job cost far more than the machines are worth?Because the loss is not limited to work in progress. Where a runtime keeps completed intermediate output on the worker that produced it, taking that process away removes results other workers were still going to fetch, so work already counted as finished must be produced again — possibly a whole earlier round of it. The later in the job this happens, the more finished work is sitting there to lose.
- What makes the boundary between rounds of work the cheapest moment to reclaim?At that moment no piece is half-finished: one wave has completed and the next has not started, so taking worker processes away abandons nothing in progress. The job simply plans the next wave for a smaller set of processes. The catch is that the owed share has to wait for that moment, and a job with long rounds makes the wait long.
- Does a job get told before its capacity is taken?It depends on the policy. Some take capacity immediately with no notice; others send a warning and an interval in which to stop cleanly, save a recovery point and hand the processes back. A warning is useful only if the job can act within the interval, and anything still running when it expires is stopped regardless.
saying these in an interview costs you the question
- Says the whole job dies when some of its worker processes are taken away.
- Thinks reclaiming is free because the work simply resumes where it stopped.
- Assumes every engine keeps intermediate output on local disk to be fetched later.
- Cannot distinguish the share policy taking capacity back from the provider reclaiming a cheap machine.
- Believes a warning interval guarantees the job can always finish what it started.
- Treats the cost as proportional to the pieces in flight, ignoring already-finished output that is lost.