skip to content

Stopping on Purpose

Taking a job down on purpose for a deploy or a settings change without losing what it knows: draining work in flight, writing a restart point, and resuming from it.

on this pageshow

questions

4

When a continuous job is asked to stop cleanly for a deploy, what must happen before the process exits?

level: middleimportance: must knowfreq 60%

answer

  1. four acts, and the order matters
  2. close the tap before you empty the pipe
  3. position never runs ahead of durable output
  4. the restart point is written last
  5. a hard kill repeats work, not input

basics

~20 s

Stop reading new input, let the records already inside reach the destination durably, advance the recorded read position to match, then write a restart point and exit. A hard kill loses no input but guarantees repeated work.

solid answer

~50 s

A clean stop is four acts in order. **One**, stop pulling from the source, while everything already inside keeps moving. **Two**, drain: let the in-flight records finish their journey and become durable output. **Three**, settle the bookkeeping — the recorded read position, the marker of how far into the input the job has got, is advanced only once the output it covers is safely written, so the two never disagree. **Four**, with nothing in flight, write a restart point taken on purpose: a durable copy of what the job was holding, on storage outside any one machine, so a later build can resume from it. Kill the job instead and you lose no input — a restart re-reads it — but the work done since the last automatic save is repeated, and any output already written for those records can appear twice.

go deeper

for a junior

Remember the order: stop reading, let what is inside finish, save, exit. The point of a deliberate stop is that nothing is halfway through when the job goes away.

for a middle

Explain the ordering rule between output and the recorded read position, and be able to say exactly what a hard kill costs — repeated work and possibly duplicated output, not lost input.

for a senior

Demonstrate that you know your drain can hang and what your engine does when it does. Say what you check after a deploy to confirm nothing was written twice.

for a principal

The interesting call is how long a drain may block a deploy before you stop hard, and who owns the consequence when you do. That is a policy, not a setting nobody chose.

## Why stopping is not one act A running job holds work in three places at once: records it has read but not finished, results it has computed but not yet made visible at the destination, and whatever it has accumulated across records. A crash abandons all three. A deliberate stop is valuable precisely because it can empty the first two and preserve the third — but only if it happens in the right order. ## The order, and why each step is where it is 1. **Stop reading.** The source tap closes first. Everything downstream keeps running. Close it last and you are chasing a moving target: every record you finish, another arrives behind it. 2. **Drain.** The records already inside the job flow through the remaining steps and reach the destination. This is the step people skip, and the one that costs the least to do properly. 3. **Reconcile output with position.** The destination's write is made durable, and the **recorded read position** — how far into the input the job has got — is advanced to cover exactly that output and no further. The ordering rule is absolute in one direction: the position never runs ahead of what is durable at the destination. Some engines go further and arrange **committing output and the read position as one unit**, so that no crash can leave one done and the other not. 4. **Write the restart point, then exit.** With nothing in flight, the job's accumulated contents are internally consistent without any cleverness, and the copy written now is the one the new build will be pointed at. ## What a hard kill actually costs It is worth being precise, because the usual answer overstates the damage in one direction and understates it in another. - **No input is lost.** Given **a retained, re-readable input** — a source that keeps its records for a while and lets a reader start again from a named position — everything read since the last durable save is still there, and the restart reads it again. - **Work is repeated.** Everything computed since the last automatic save is done again. For a job that saves every few seconds that is trivial; for one that saves every few minutes and holds a large accumulated set, it is not. - **Output can duplicate.** Records whose output was written but not covered by an advanced read position are processed and written a second time on restart. Whether anyone notices depends entirely on whether the destination makes a repeat harmless — which is a separate subject with its own owner. - **Partial output may be left behind.** A step that was midway through writing leaves a fragment. Engines that stage each unit's output privately and promote the whole result in one final act make those fragments invisible; ones that write in place do not. ## Where the quiet moment comes from — and here the engines differ The entire drain exists to manufacture a moment when nothing is in flight. Some runtimes get it for free. | Runtime model | Quiet moment | What a clean stop must do | |---|---|---| | **Repeated small finite jobs** — an endless input cut into short bounded pieces, each run as a complete little job | Exists naturally at the end of each piece | Finish the current piece, stop starting new ones | | **Record-at-a-time processing** — each record traverses the whole job on arrival, workers hold results between records | None naturally | Close the source, then wait for the last records to reach the destination | | **Two-phase disk-to-disk model** over a bounded input | Every phase boundary is one | Usually nothing: the job is re-run rather than resumed | This is why the same words describe very different amounts of engineering in two engines, and why a candidate who has used only one of them will describe the drain as either trivial or elaborate and be right about their own engine only. ## The drain that will not finish A drain waits on the slowest thing downstream. If the destination is unreachable, or one worker is stuck on an enormous group, the drain does not complete. Engines therefore bound it: wait this long, then stop hard. Designs differ on what happens next — whether a restart point is written anyway from a less convenient moment, or the stop degrades into the crash case and the last automatic save is what you resume from. Knowing which your engine does is the difference between a deploy that takes an extra minute and one that silently costs you five minutes of reprocessing you did not budget for. And one thing that is not a drain: the platform terminating your process. That is the crash case, and any shutdown work the job manages to do in the seconds it is given is a race it may lose.

  • What happens to the records in flight if the stop is not allowed to drain?
    They are not lost: the recorded read position still names a point before them, so a restart reads them again. What is lost is the work done on them, and the guarantee that the destination sees each effect once — output already written for those records is written a second time unless the destination is arranged so a repeat leaves no extra trace.
  • Must the drain succeed before the stop can complete?
    No, and this is worth checking for your engine. A drain waits on the destination and on the slowest in-flight work, so it can hang; engines bound it with a timeout and then stop hard. What happens to the restart point in that case differs by design — some write one anyway, some leave you resuming from the last automatic save, which is the crash case.
  • Does the coordinating process take part in the drain?
    Yes, and it must outlive it. The coordinating process — the one holding the job's plan and tracking which work is done — is what decides the source has closed, observes that the last records have landed, triggers the save and declares the job stopped. Losing it mid-drain leaves nobody who knows whether the drain finished.

Closing a shop for the night: you lock the door to new customers first, then serve the ones already inside, then count the till and write the total down. Sweeping everyone out mid-purchase loses no customers — they come back tomorrow — but some are charged twice and the till never balanced.

saying these in an interview costs you the question

  • Thinks killing the job loses the records that were in flight.
  • Writes the restart point before the in-flight work has finished.
  • Advances the recorded read position before the output it covers is durable.
  • Assumes every runtime has a natural moment with nothing in flight.
  • Believes a drain always completes and needs no time bound.
  • Treats the platform stopping the process as a clean drain.
open as a page

A job is being taken down for a deploy: what decides whether you can stop it and submit the new version, or must save first?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Whether the job carries anything between records. A job that keeps nothing is stopped and resubmitted, provided its recorded read position is durable. A job holding accumulated results must write a restart point before it exits, or those results are gone.

open as a page

How does a restart point an operator asks for before stopping a job differ from the ones written automatically?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Different moment, different lifetime, different purpose. Automatic saves happen on a timer while the job runs and may be discarded when it stops; a restart point taken on purpose is written once after the drain, for a build not yet deployed.

open as a page

You are budgeting the outage window for a planned stop and resume of a continuous job: what goes into it, and what must it fit inside?

level: principalimportance: should knowfreq 36%

basics

~20 s

Five spans: draining, writing the restart point, the deploy and getting machines, reading the saved copy back, and catching up the input that piled up. It must fit inside the source's retention and whatever freshness the consumers were promised.

open as a page