A job is being taken down for a deploy: what decides whether you can stop it and submit the new version, or must save first?
answer
- ask what the job carries between records
- accumulated results, or nothing at all
- a durable read position matters either way
- carries nothing: stop and resubmit
- holds totals: write a restart point first
basics
~20 sWhether the job carries anything between records. A job that keeps nothing is stopped and resubmitted, provided its recorded read position is durable. A job holding accumulated results must write a restart point before it exits, or those results are gone.
solid answer
~50 sAsk what the job would be unable to obtain again if every one of its processes vanished. A job that reads a record, filters or reshapes it, writes it and forgets holds nothing irreplaceable: the deploy is stop, submit the new version, resume from the **recorded read position** — the durable marker of how far into the input the job has got, advanced only once the work it covers is safely written. A job that has *accumulated* something — running counts per key, half-finished matches, the last value seen per customer — holds it in worker memory or on the workers' own disks, and those die with the processes. That job needs a **restart point taken on purpose**: a durable copy of what it was holding, written before it exits. Treating every deploy as the second case invents ceremony most jobs do not need.
go deeper
Recall the one question that decides everything: does this job remember anything between records? If not, stopping it is stopping it, and the only thing that must survive is how far into the input it had got.
Explain why worker memory is not durable and why a copy written to the machine that is about to go away is worthless. Be able to name what a job accumulates: totals per key, unmatched records, last-seen values.
Show that you check before you deploy rather than after. Say how you would establish what a given job holds, and what you would look at the next morning to prove nothing was silently lost.
Frame it as a property teams should know about their own jobs in advance. A fleet where nobody can say which jobs accumulate has no deploy procedure, only a habit that has not failed yet.
## Two stops that look identical from outside When a job goes away, only two things can be true of the work it was in the middle of: it wrote down enough to resume, or it did not and something will have to be done again. A crash gives you no opportunity to write anything down. A deliberate stop — for a new build, a changed setting, a machine rotation — hands you that opportunity, and this whole subject is about spending it well. The first decision is whether you need to spend it at all. ## The test: what does this job carry between records? Ask what the job could not obtain again if every one of its processes disappeared this second. - **It carries nothing between records.** It reads a record, maps it, filters it, enriches it from an external lookup, writes it, and retains no trace. Whatever sits in a worker's memory at any instant is derived from the record currently in hand. Nothing here is irreplaceable, because reading the input again reproduces it. - **It has accumulated something.** Running totals per key, one side of a match waiting for its partner, the last value seen per customer, the set of identifiers already seen. That content exists *only* because this job has been running — sometimes for weeks — and no future record can reconstruct it. The second kind is a real storage system that happens to live inside a job, and stopping it carelessly destroys a database nobody remembered was there. ## Case one: stop and resubmit — but not for free Even a job carrying nothing has one durable artifact worth protecting: the recorded read position. If it does not survive the stop, the new version starts wherever its defaults say. Starting from the earliest retained record reprocesses everything and writes the output again; starting from the newest leaves a hole exactly the size of the outage. So the cheap case is not *nothing* — it is *one* thing, and it is usually already durable because the running job was maintaining it anyway. A **finite** job — one over a fixed, bounded input — is a further special case. Stopped halfway, it is normally not resumed at all; it is re-run, and the cost is the compute already spent rather than any knowledge. How much of that spend survives depends on the engine: in **the two-phase disk-to-disk model**, where each phase materialises its output to durable files before the next phase reads them, completed work may still be on disk and reusable, while an engine that streams results between steps in memory generally has to redo them. ## Case two: write a restart point before exiting For the accumulating job the sequence is: stop reading new input, let the work already inside the job finish, write a **restart point taken on purpose** — a recovery point the operator asks for before stopping, where a *recovery point* is a durable copy of everything the job would otherwise lose, held on **durable shared storage**, meaning storage outside any one machine and reachable from whichever machine resumes, because a copy on the dead machine's disk is no recovery point — and only then exit. The new version is submitted pointing at that artifact instead of at the beginning of the input. ## Where engines differ, and they do | | Carries nothing between records | Has accumulated something | |---|---|---| | What must survive the stop | The recorded read position | That position **and** the accumulated contents | | Deploy procedure | Stop, submit new version, resume | Drain, write a restart point, submit against it | | Cost of getting it wrong | Reprocessing or a gap in output | Silent loss of weeks of arithmetic | | Outage window dominated by | Acquiring machines, resubmission | Writing, then reading back, what was held | And across the engines themselves: 1. A runtime that serves an endless input as **repeated small finite jobs** — cutting it into short bounded pieces and running a complete little job over each — already reaches moments when nothing is in flight, so the quiet point a clean stop needs exists for free. 2. A runtime doing **record-at-a-time processing**, where each record traverses the whole job as it arrives and workers hold results between records, has no natural quiet moment and must manufacture one. 3. Whether an operator-requested restart point is a distinct kind of artifact or merely a retained ordinary one varies by engine, as does whether the engine promises a later build can read it. One thing does not vary: the platform restarting your process, or rescheduling it onto another machine, restores no work at all. It gives you a running process with an empty memory, which is the crash case wearing a deploy's clothes.
- Does a job that carries nothing between records still have an outage window when you deploy it?Yes. The window is resubmission, acquiring machines, and the input that piled up meanwhile — output stops for all of it. What the cheap case avoids is the two expensive extras: writing out what the job was holding before it exits, and reading all of it back before the new version can process its first record.
- For a job you did not write, how do you tell which case you are in?Look for logic that spans records: grouping by key with an aggregate, matching one input against another, deduplication, anything comparing a record to a previous one, anything with an expiry. Those accumulate. Per-record mapping, filtering, reshaping and external lookups do not. If in doubt, assume it accumulates and write a restart point; the wrong guess that way costs minutes rather than history.
saying these in an interview costs you the question
- Assumes every job needs a restart point written before a deploy.
- Says a job holding nothing can restart anywhere, ignoring its read position.
- Calls the platform restarting the process a recovery of the job's work.
- Treats a deliberate stop as indistinguishable from a crash.
- Assumes a half-finished finite job can always be resumed rather than re-run.