The provider takes back a discounted machine in the middle of a run. Besides the piece it was computing, what does the job lose?
answer
- the discount has a catch
- the unit taken back is a machine
- several worker processes go at once
- output lived only on that machine
- consumers were still fetching it
basics
~20 sEverything that machine held goes at once: every worker process on it, the pieces they were running, anything cached there, and intermediate output produced there that other workers had not yet fetched. Finished work can therefore have to be produced again.
solid answer
~50 sThe unit the provider takes back is a **machine**, and a machine usually hosts several worker processes - the processes that actually run pieces of the job and own the memory those pieces use. So the loss is not one piece. Whatever those processes held goes with them: pieces in flight, anything kept there for reuse, and - on engines that materialise intermediate output for a later round of work to fetch - output produced minutes ago and already marked complete. Consumers that still needed that output fail their fetch, so the missing shares have to be produced again by re-running the steps that made them. Engines that push records to the next step as they are produced hold no such files; there the loss is the workers' retained state plus records in flight, and the affected part of the job resumes from its last saved recovery point instead.
go deeper
Recall that a discounted machine can be withdrawn at any moment and that everything running on it disappears with it, not only the piece it was working on.
Explain why output produced earlier can vanish: it was materialised on that machine for a later round to fetch, so the fetch fails and the work is done a second time.
Show that you size the blast radius: which worker processes sat there, what they had produced that is still needed, and what the surviving machines must now redo while carrying their own work.
Frame it as a purchase. The hourly discount buys the risk that a machine's entire footprint, produced output included, is discarded at a moment nobody chooses.
## The purchase and the catch A **reclaimable machine** is a discounted machine the provider may take back at short notice for reasons that have nothing to do with the job running on it - capacity sold cheaply is wanted by someone paying full price, or the pool it came from shrank. The discount is genuine and so is the condition attached to it. What makes this a distributed-processing question rather than a general infrastructure one is that a job's workers feed each other. For a single-machine task, losing the machine costs the task. For a job whose later work consumes what earlier work produced, losing a machine costs the work in progress **plus everything that machine held which someone else still needed**. ## Count the machine, not the process Two units are easy to conflate, and the answer depends on keeping them apart. - A **worker process** is the process on a cluster machine that actually runs pieces of the job and owns the memory those pieces use. A large job holds hundreds of them. - A **machine** is the host. It normally carries several worker processes, and the machine is what the provider withdraws. A withdrawal therefore removes a set of worker processes simultaneously, together with the local disks underneath them. Anything whose only copy sat on those disks is gone at the same instant. Start an interview answer here: the blast radius is the machine's whole footprint, and sizing the damage means listing what that footprint contained. ## What the footprint contained 1. **Pieces in flight.** Work handed out and not finished. Everybody names this one, and it is usually the smallest of the three. 2. **Anything kept for reuse.** A working set the job asked to hold on the cluster so a later step would not re-read or recompute it, and, in a long-running job, the per-key state those workers retained. Treat that retained state as an opaque quantity: it has to exist again before those keys can be processed, and making it exist again is not free. 3. **Intermediate output already produced.** On engines that materialise it, the product of an earlier **round of work** - one wave of pieces that all run before the next wave can start - is written where it was produced and pulled across the network by the next wave. A piece is marked complete when it has written its share, not when anyone has read it. The third item is the one that surprises candidates. Completed work is not safe work; it is work whose product may still be sitting on a machine that can disappear. ## How the job finds out The job rarely experiences the withdrawal as an announcement. The **coordinating process** - the single process in a distributed job that plans the pieces, hands them out, tracks what finished, and receives anything the program asks to bring back to one place - usually learns what it learns about any dead machine: heartbeats stop, and the workers that were fetching report that the data promised to them is not there. From that point the missing product has to exist again, which means running again the steps that originally produced it. That is **lost-work recomputation**, and it runs on the machines the job still holds, which are themselves withdrawable. ## Where engines differ This is where a confident answer built on one runtime goes wrong. | how the runtime moves data between rounds | what a withdrawal destroys | how the job gets it back | |---|---|---| | materialises each round's output for the next round to fetch | produced output on local disk, plus pieces in flight and anything cached | produces the missing pieces again from their own inputs | | pushes records to the next step as they are produced | records in flight and the retained state of those workers; there are no files to lose | the affected part resumes from its last saved recovery point and re-reads input from there | | runs continuous work as a rapid succession of small finite jobs | whatever the small job currently in flight had produced | the current small job's lost pieces are produced again | A **detached intermediate-output server** - a separate process that keeps serving a finished worker's intermediate output after that worker is gone - changes the first row only for the case where the worker process ends while the machine lives. When the machine itself is taken back, the server and the disks go with it. ## What this question does not decide How much is then re-run - the single lost piece, its whole round, or the job from its start - is a question about retry granularity, and engines answer it differently. So is what a restart resumes from. The general practice of designing for withdrawable capacity - the warning signal, draining work in progress, spreading risk across pools, mixing guaranteed with discounted machines - belongs to the provider's capacity model. What belongs here is the one thing that model cannot state: in a job whose workers feed each other, losing a machine can cost work that was already finished.
- What can keep a finished worker's intermediate output usable after that worker process has exited?A separate process that goes on serving it - a detached intermediate-output server - so a worker process ending does not by itself lose what it produced. It helps when a process exits or is handed back voluntarily. It does not help when the whole machine is taken back, because the disks holding the output go too. Where the runtime materialised nothing, there is nothing for such a process to serve.
- Is the loss smaller if nothing on that machine had been fetched yet?Usually yes. What must be produced again is what was already produced and is still needed. A machine taken back before anything downstream consumed its output costs the pieces in flight on it plus whatever it cached. The same withdrawal later in the run can cost the product of several completed rounds.
- Does the job run narrower for the rest of the run after a machine is withdrawn?It depends on whether it can get a replacement. A job may ask whatever grants it machines for another and be given one, or run on fewer machines and take longer. Either way the redone work is carried by the machines it still holds, and those are on the same discounted terms.
A shared kitchen where one prep station is wheeled out mid-service. The dish that station was plating is the obvious loss; the expensive loss is the trays of prepared stock the other stations were about to collect, which now have to be made again.
saying these in an interview costs you the question
- Thinks only the in-progress piece is lost and everything already finished is safe.
- Treats the withdrawal as losing one worker process, when a machine hosts several.
- Assumes produced intermediate output is always readable from shared storage afterwards.
- Says every engine keeps produced output on local disk for a later round to fetch.
- Calls a machine that has been taken back a slow worker and waits for it to return.