How does a restart point an operator asks for before stopping a job differ from the ones written automatically?
answer
- same bytes, different contract
- one is the engine's, one is yours
- taken after the drain, not during
- may or may not outlive a clean stop
- written for a build not yet deployed
basics
~20 sDifferent moment, different lifetime, different purpose. Automatic saves happen on a timer while the job runs and may be discarded when it stops; a restart point taken on purpose is written once after the drain, for a build not yet deployed.
solid answer
~40 sAn automatic recovery point is the engine's own insurance: written on a timer *while the job keeps running*, optimised for being cheap and frequent, and owned by the running job — several engines discard them on a clean termination, others retain them, and many make it a choice. A restart point taken on purpose is written *because you asked*, once, after the drain, when nothing is in flight; it is meant to outlive the job and to be readable by the build you are about to deploy. Two practical consequences. First, resuming from yesterday's automatic save is not the same deploy: it repeats everything since. Second, the purposeful write is often a whole, self-contained copy rather than only what changed, so it costs more and that cost lands inside the outage window.
go deeper
Know that a job saves copies of what it holds on a timer, and that before a deliberate stop an operator asks for one more — taken after the job has finished the work already inside it.
Explain the two mechanical differences: the purposeful one is written when nothing is in flight, and it is written to be read by code that does not exist yet. Automatic ones are written mid-flight and belong to the run.
Show you know your engine's cleanup behaviour on a clean stop, and that you would not plan a deploy around an automatic save that predates the drain. Say how you verify the artifact exists before killing anything.
The judgment is how much the organisation depends on a promise the engine makes about reading older artifacts, and what the fallback is when an upgrade breaks it: rebuild from the input, or accept a cold start.
## Same mechanism, different contract Both are recovery points: a durable copy of everything the job would otherwise lose, written to storage outside any one machine so that whichever machine resumes can read it. The difference is not the bytes, it is the contract around them — who asks for it, when it is taken, how long it lives, and who is expected to be able to read it. ## The four axes | | Automatic recovery point | Restart point taken on purpose | |---|---|---| | **Who asks** | The engine, on a timer | An operator, once, as part of a stop | | **When** | While the job runs, records still in flight | After the drain, nothing in flight | | **Lifetime** | Belongs to the run; may be discarded when it ends | Kept deliberately, outlives the job | | **Read by** | The same build, resuming itself | A build you have not deployed yet | ### Who asks, and how often Automatic saves exist to bound how much work a crash costs, so they are tuned to be frequent and cheap. Frequency is a trade: more often means less repeated work after a failure and more overhead during normal running. The purposeful one happens once, so cheapness matters much less than completeness. ### When it is taken This is the substantive difference. An automatic save is taken *while records are moving*, so engines that do this need a mechanism to make the parts saved by workers that never paused at the same instant add up to one coherent picture — typically **a marker travelling with the records**, a special element injected into the stream that each worker passes along after saving its own part. A restart point taken on purpose is written after the drain, when nothing is in flight, so coherence is free. It is also *later*: it covers every record the job ever read, which is exactly what you want before a deploy and exactly what the last automatic save cannot give you. ### Lifetime Here is where teams are caught out, and here the honest answer is that engines disagree. Some treat automatic saves as scratch belonging to the run and remove them when the job terminates cleanly, on the reasoning that a job which stopped on purpose has no crash to recover from. Others retain them. Several make it a deliberate choice. The consequence of guessing wrong is discovering, after a clean stop, that the only durable copy you have is the one you did or did not ask for. ### Who is expected to read it An automatic save only has to be readable by the process that wrote it, continuing as itself. That permits formats that are fast to write and closely tied to the running build — several engines write them **incrementally**, recording only what changed since the previous one, so any single file depends on the ones before it. A restart point taken on purpose is aimed at a build that does not exist yet, so it is normally written whole and self-contained. Be careful about how far that promise goes. Whether a *later version of the engine* can read an older purposeful restart point is a property the engine promises or does not — some do across a stated range, some promise nothing between major versions. And whether your *new code* can interpret the contents once the shape of what the job stores has changed is a different problem again, with its own owner; the purposeful artifact makes resumption possible, it does not make an altered layout readable. ## What this means on the day 1. **Do not plan a deploy around the last automatic save.** It predates the drain. Resuming from it means every record since is processed again, and any output already written for those records repeats. 2. **Budget the write.** A whole, self-contained copy of a large accumulated set is not instant, and unlike an automatic save it is not overlapped with useful work — it sits inside the span in which the job produces nothing. 3. **Know your engine's cleanup behaviour before you need to.** Find out on a Tuesday, not during an incident. 4. **A restart point is not a backup of your data.** It is a resumable position for one job, tied to that job's plan and contents. It is not a stored table you can query, and it is not a substitute for being able to re-read the input. ## The case where none of this applies A job carrying nothing between records has no accumulated contents to copy, so its purposeful restart point is, at most, its recorded read position — and that is usually durable already. For those jobs the distinction collapses, and the honest advice is to stop, submit and resume.
- If the automatic saves do survive a clean stop, why ask for a purposeful one at all?Because the last automatic one predates the drain, so resuming from it repeats everything the job did since — the whole point of the drain was to have nothing left to repeat. It may also be incremental, depending on files whose lifetime you do not control, and it carries no promise about being readable by the build you are about to start.
- What can stop a new build from resuming from a restart point you wrote correctly?Compatibility is a promise the engine makes or does not. Some guarantee that a later version reads an older purposeful restart point within a stated range; some promise nothing across major versions. Separately, if the shape of what your job stores has changed, resumption may succeed mechanically and still fail to interpret the contents — a different problem with a different owner.
- Is it worth writing a purposeful restart point even when you do not intend to stop?Occasionally, yes: it gives you a known-good position to fall back to before a risky change, independent of whatever the engine is doing on its timer. The cost is a full write and a pause while it happens, so it is a deliberate act before something you expect might go wrong, not a habit.
saying these in an interview costs you the question
- Says the last automatic save is as good as one taken before the stop.
- Assumes automatic recovery points always survive a clean termination.
- Believes any saved copy can be read by any future build.
- Calls a restart point a backup of the job's output data.
- Forgets the purposeful write happens while output is stopped.
- Thinks an incremental save is self-contained on its own.