A six-hour fit step in a nightly forecasting graph dies near hour five — what must its checkpoints record to resume?
answer
- continue, do not restart
- cursor and generator state, not just weights
- bound to input identity and parameters
- never overwrite the only good checkpoint
- keyed by run and step, not attempt
basics
~20 sEnough to continue rather than restart: the model and optimiser state, the exact position in the input, the random-number state, and the identity of the inputs and parameters being fitted — each published atomically and retained alongside the previous good checkpoint.
solid answer
~40 sModel weights alone are not resumable state. A resumed fit has to land where the failed one stood: the **position in the input** (which shard, epoch and row cursor), the **random-number generator state** so sampling continues rather than repeats, the accumulated training metrics, and the **identity of the declared inputs and hyperparameters** the checkpoint belongs to — resuming a different dataset or a different parameter set silently produces a model nobody asked for. Each checkpoint is written to a new location and published atomically, and the previous complete one is kept, because overwriting the only good checkpoint is how a crash mid-write costs the whole five hours. The checkpoint must also be addressed by run and step rather than by attempt, or the retry will not find what the first attempt wrote.
code
json · 12 lines{
"run_id": "nightly-2026-09-18",
"step": "fit_demand_model",
"input_dataset": "store_sales_features@2026-09-17",
"hyperparameters": { "objective": "tweedie", "learning_rate": 0.05, "max_rounds": 4000 },
"cursor": { "epoch": 3, "shard": 41, "rows_consumed": 812000000 },
"rng_state": "opaque-generator-position",
"learner_state_location": "ckpt/nightly-2026-09-18/fit/00013",
"metrics_so_far": { "train_wape": 0.118 },
"written_at": "2026-09-18T03:41:02Z",
"complete": true
}go deeper
Know that resuming means continuing from where the work stopped, which needs more than the model's parameters — at minimum the place in the input it had reached and a record it can trust.
List what a checkpoint has to hold and why each item matters, and explain why it is written to a new location and published atomically instead of being overwritten in place.
Reason about spacing from the write cost and the failure rate, keep the checkpoint addressed by run and step so a retry finds it, and validate input identity and hyperparameters before accepting a resume.
Decide where checkpointing earns its keep at all: a step comfortably shorter than its optimal interval should simply rerun, and the platform should make the atomic-publish-and-retain behaviour a default rather than each team's private discipline.
## Why weights alone are not resumable state The fit step reads ninety days of per-store, per-item features and takes about six hours. It dies near hour five. Restarting from zero costs a nightly refresh; resuming costs an hour — but only if what was written is genuinely enough to continue from. A checkpoint that holds only model parameters produces a resumed run that is *not* the interrupted one: it restarts the input from the beginning, re-draws the same random choices, and double-counts whatever it had already consumed. To match the uninterrupted run up to what the checkpoint records, it needs: - **Learner state** — model parameters, plus whatever the optimisation carries between updates (momentum, accumulated gradients, the trees built so far and the current residuals). - **Input position** — the shard, epoch and row cursor already consumed, so the resume continues rather than repeats. - **Random-number state** — the generator's position, not merely the seed, so sampling, shuffling and initialisation continue the same stream. - **Accumulated metrics** — running training loss or error, so the resumed run's curve is continuous and early-stopping logic is not reset. - **Binding identity** — the run, the step, the declared input the fit is reading and the hyperparameters in force. A checkpoint is only valid for the thing it was taken from. - **A completeness marker** — the record is either fully published or not present. ## Writing a checkpoint so it survives the crash that caused it The checkpoint is written by the process most likely to die. 1. Write to a **new location** each time; never modify the previous one in place. 2. Flush and confirm the bytes are durable before anything points at them. 3. Publish by an **atomic pointer swap** to the new location, so a reader sees the previous complete checkpoint or the new complete one. 4. **Retain at least the previous complete checkpoint.** Keeping exactly one means a crash mid-publish leaves nothing usable. 5. **Validate on resume** — read the record, confirm its input identity and hyperparameters match the step about to run, and fall back to the previous checkpoint if they do not. A fit that overwrites one file in place breaks exactly this: the crash that truncates it also removes the only thing that would have saved the five hours. ## Where the checkpoint lives relative to the graph The fit is one node. When it fails, the graph retries **that node**, and the retry must find the checkpoint the previous attempt wrote — so the checkpoint location is keyed by **run and step**, not by attempt. Key it by attempt and every retry starts from zero while cheerfully writing checkpoints nobody will read. That also sets the contract with the rest of the graph: - the step is not complete until its **declared output** — the fitted model artifact — is published atomically, exactly like any other step; - checkpoints are internal scratch, not the output, and are discarded once the step completes; - resuming does not weaken the output guarantee, because the published artifact is still produced by one atomic publish at the end. ## How often to checkpoint Checkpointing is a trade between write overhead and redone work. If a checkpoint costs *C* to write and failures arrive on average every *M*, the spacing that minimises total waste is near **sqrt(2 x C x M)**. With C = 60 seconds and M = 6 hours (21,600 seconds): - spacing ≈ sqrt(2 x 60 x 21,600) = sqrt(2,592,000) ≈ 1,610 seconds ≈ **27 minutes**; - write overhead ≈ 60 / 1,610 ≈ **3.7 percent** of runtime; - expected redone work per failure ≈ half the interval ≈ **13 minutes**. So on a six-hour fit, roughly thirteen checkpoints cost about fourteen minutes in total and cap the loss from any single crash at about a quarter of an hour. Checkpoint every minute and the overhead alone exceeds an hour; checkpoint hourly and the average crash costs thirty minutes. For a step that finishes well inside one interval, checkpointing is not worth writing at all — rerun it. ## What resume must not change | Aspect | Must match the interrupted run | Why | |---|---|---| | Declared input identity | Yes | Resuming onto different data produces a model fitted to two datasets | | Hyperparameters | Yes | A mid-run parameter change makes the curve uninterpretable | | Random-number state | Yes | A reset stream re-draws choices already made | | Input cursor | Yes | A reset cursor re-consumes rows and skews the fit | | Worker placement | No | Any worker may pick the step up, provided the checkpoint is readable | | Wall-clock time | No | The record is bound to the run, not to when it resumed | The rule behind the table: a resume is legitimate only when the step it re-enters is the same step, fitting the same declared input, under the same parameters. Anything else is a new run wearing an old run's weights.
- The fit step runs across several workers — what changes about checkpointing and restart?The checkpoint becomes a coordinated event: every worker writes its part at the same boundary and the checkpoint counts as complete only when all parts are published. The graph schedules the step as one all-or-nothing unit, so a single worker's death restarts the whole step from the last coordinated checkpoint rather than re-running one shard. How the workers divide the computation is a separate subject; the graph owns the gang-scheduled unit and the shared resume point.
- When is checkpointing not worth adding?When the step's runtime is small relative to the optimal spacing. With a 60-second checkpoint and failures every six hours the spacing is around 27 minutes, so a step that finishes in ten minutes should simply be rerun — the write overhead would exceed the work it protects.
- Does resuming weaken the guarantee the step gives the rest of the graph?No, provided checkpoints stay internal. The step's declared output — the fitted artifact — is still produced by a single atomic publish at the end, so downstream steps see one complete artifact or none. Checkpoints are scratch state, cleaned up once the step completes.
saying these in an interview costs you the question
- Saves model weights only and calls the step resumable
- Overwrites one checkpoint file in place each time
- Stores the seed rather than the generator's current state
- Keys the checkpoint by attempt, so retries never find it
- Resumes onto a different input window without noticing
- Treats the latest checkpoint as the step's published output