A Keras fit() run is killed at epoch 40 of 100 — which callback resumes it automatically?
answer
- one callback exists purely for crashes
- checkpoints are artifacts, not resume points
- optimizer state is what a weights file loses
- the epoch index has its own fit argument
- backup is deleted on a clean finish
basics
~20 skeras.callbacks.BackupAndRestore(backup_dir=...). It snapshots model weights, optimizer state and the epoch counter as training proceeds, and on the next fit() call with the same backup_dir it restores them and continues from where it stopped. ModelCheckpoint only writes files; resuming from those is manual.
solid answer
~50 s`BackupAndRestore` is the fault-tolerance callback; `ModelCheckpoint` is the artifact callback, and confusing them is the usual mistake. BackupAndRestore writes a backup of the full training state — weights, optimizer state and the epoch index — into `backup_dir`, and at the start of the *next* `fit()` it detects that backup, restores it, and resumes at the interrupted epoch with no arguments changed on your side. On a clean finish it deletes the backup, so a completed run leaves nothing behind. Resuming from a ModelCheckpoint file is a different, manual flow: `keras.saving.load_model(path)` — which recovers the optimizer state only if the checkpoint was a full `.keras` save, not `save_weights_only=True` — then `fit(..., initial_epoch=40)` so epoch-indexed schedules and log lines line up. Neither approach restores your data-pipeline position, so an epoch's sample ordering after resume is not what it would have been.
code
python · 11 linesimport keras
model = keras.Sequential([keras.Input(shape=(8,)), keras.layers.Dense(1)])
model.compile(optimizer="adam", loss="mse")
callbacks = [
keras.callbacks.BackupAndRestore(backup_dir="/mnt/state/backup"),
keras.callbacks.ModelCheckpoint("art/best.keras", monitor="val_loss", save_best_only=True),
keras.callbacks.ModelCheckpoint("art/last.keras"),
keras.callbacks.CSVLogger("art/history.csv", append=True),
]go deeper
Know that Keras has a dedicated callback for surviving interruptions, and that reloading a checkpoint by hand also needs the initial_epoch argument so epoch numbering continues rather than restarting.
Explain what a resume actually needs — weights, optimizer state and the epoch index — and why a weights-only checkpoint restores only the first of the three.
Show the production reasoning: durable storage for the backup directory, a last-epoch checkpoint alongside the best-only artifact, CSVLogger with append=True, and the stateful-callback resets that let a restarted run clobber a better model.
Own the fault-tolerance budget — checkpoint cadence against storage cost and expected preemption rate, whether jobs are designed to be killed at all, and what reproducibility guarantee you are willing to publish for resumed runs.
## Two callbacks that both write files, for different reasons On a preemptible instance or a long training job, a crash at epoch 40 of 100 should not cost 40 epochs. Keras has one callback whose entire purpose is that, and another that people misuse for it. `keras.callbacks.ModelCheckpoint` exists to produce an **artifact**: the model you will evaluate, deploy or promote. `keras.callbacks.BackupAndRestore` exists to produce a **resume point**: temporary state that lets an interrupted `fit()` continue and is deleted once it is no longer needed. ## BackupAndRestore `keras.callbacks.BackupAndRestore(backup_dir="/path/to/backup")` is passed to `fit()` like any callback. During training it periodically saves the training state into that directory. When `fit()` is called again with a BackupAndRestore pointing at the same directory, the callback finds the existing backup at the start of training, restores the model weights and optimizer state, and sets the loop's starting epoch to where the interruption happened. You do not change your call site: the same script, rerun, picks up mid-run. The backup includes what a resume genuinely needs and a bare weights file does not: the optimizer's internal state. Adam keeps first- and second-moment estimates per parameter; a scheduler's or optimizer's iteration counter drives step-based rate schedules. Restart with the weights alone and those all reset, so the first epochs after a resume behave like the first epochs of a fresh run — a visible loss bump that people misdiagnose as a data problem. On successful completion the callback removes the backup directory contents, since keeping a resume point for a finished run only invites a later process to resume something already done. `backup_dir` must therefore point somewhere durable across the interruption — a mounted volume, not container-local scratch that dies with the pod — and must not be the same directory as your checkpoint artifacts. ## The manual route, and initial_epoch If all you have is a ModelCheckpoint file, resuming is a script you write: 1. `model = keras.saving.load_model("ckpt/last.keras")`. A full `.keras` save carries architecture, weights *and* optimizer state, so the optimizer resumes warm. A `save_weights_only=True` file carries weight values only — trainable kernels and biases plus non-trainable weights such as BatchNormalization moving statistics — and no optimizer slots. 2. `model.fit(..., epochs=100, initial_epoch=40)`. `initial_epoch` does not change how many epochs run in the sense of a countdown; it sets the starting index, so the loop runs epochs 40 through 99. Skip it and you rerun 100 more epochs *and* every epoch-indexed learning-rate schedule restarts from epoch 0, quietly undoing the decay. For this to work, ModelCheckpoint must have been saving on a cadence you can afford to lose — `save_best_only=True` alone is not a resume point, because the best epoch may be far behind the current one. Teams that resume manually usually run two checkpoint callbacks: one best-only artifact, one last-epoch file with a constant path. ## What neither of them restores Data-pipeline position. Neither mechanism knows where in your dataset iterator the interruption happened, so after a resume the epoch starts from the beginning of the data with its own shuffle. For most training this is harmless, but it means "resumed run" is not bit-identical to "uninterrupted run", and if you are claiming exact reproducibility you have to say so. Random-number state is in the same category: seeds are re-consumed from wherever the new process starts. Stateful callbacks are the other gap in the manual route. A fresh `EarlyStopping` has an empty `best` and a zeroed wait counter; a fresh `ModelCheckpoint(save_best_only=True)` pointed at an existing file will overwrite a better model from the earlier run on its very first epoch, because it has no memory of what "best" meant before. Writing each run to a fresh directory avoids that whole class of problem. ## Logs across a restart `CSVLogger(filename, append=True)` continues an existing CSV rather than truncating it, which is what you want across a resume; the default `append=False` starts a fresh file and destroys the earlier record. The `History` object returned by the second `fit()` covers only the epochs that second call ran, so a full training curve has to be reassembled from the CSV or from your tracking system, not from History. ## Choosing Use BackupAndRestore when the job runs on infrastructure that can kill it — spot or preemptible instances, shared clusters, anything with a wall-clock limit — and durable storage is available for `backup_dir`. Use the manual load plus `initial_epoch` route when the restart is deliberate: extending a run that finished, changing hyperparameters mid-course, or branching from a specific epoch you chose.
- What exactly does initial_epoch change in a Keras fit() call?It sets the epoch index the loop starts from, so `fit(epochs=100, initial_epoch=40)` runs epochs 40 through 99 rather than another 100. That matters beyond bookkeeping: anything keyed on the epoch number — a LearningRateScheduler function, log lines, checkpoint filename templates — sees the true index instead of restarting at zero and re-running a decay curve you already worked through.
- Why does loss jump after resuming from a save_weights_only checkpoint?Because the optimizer restarted cold. A weights-only file holds weight values and nothing else, so Adam's moment estimates and the iteration counter are reinitialised. The first steps after resume take effectively unconditioned updates and any step-based schedule restarts, producing a bump that looks like a data bug. A full .keras save carries the optimizer state and avoids it.
- You resume with a fresh ModelCheckpoint(save_best_only=True) pointed at the existing file. What goes wrong?The callback's internal best value starts empty in the new process, so the first epoch after resume counts as an improvement and overwrites the file — replacing a genuinely better model from before the crash. Write each run into a fresh directory, or template the filepath, so no previous artifact can be clobbered by a restarted run's first epoch.
- Is a resumed run reproducible against an uninterrupted one?No, and it is worth saying so plainly. Neither BackupAndRestore nor a manual reload restores the data pipeline's position or the random-number stream, so after a resume the epoch re-iterates the dataset with its own shuffle. Metrics come out statistically equivalent, not identical. Any claim of exact reproducibility has to be scoped to uninterrupted runs.
saying these in an interview costs you the question
- Treats save_best_only checkpoints as a resume point
- Omits initial_epoch and reruns the full epoch count
- Assumes a weights-only file restores optimizer state
- Points backup_dir at ephemeral container-local storage
- Expects a resumed run to be bit-identical to an uninterrupted one