skip to content

In Keras, what does ModelCheckpoint(save_best_only=True) change about when a file is written?

level: juniorimportance: must knowfreq 70%

answer

  1. writes only when the number improves
  2. monitor plus mode decides improvement
  3. default is write every epoch
  4. Keras 3 wants a .keras path
  5. missing monitor key warns, then skips

basics

~20 s

With save_best_only=True, ModelCheckpoint writes a file only on epochs where the monitored value improves on the best seen so far. With the default False it writes every epoch. The monitored value is named by the monitor argument, which defaults to val_loss.

solid answer

~40 s

`keras.callbacks.ModelCheckpoint` is passed to `fit(callbacks=[...])` and saves the model at the end of each epoch (or every N batches with `save_freq=N`). By default `save_best_only=False`, so it writes unconditionally. Setting it to `True` makes the callback compare `logs[monitor]` against the best value it has seen and write only on an improvement — where "improvement" means lower for `mode="min"`, higher for `mode="max"`, and `mode="auto"` guesses from the metric name. Two practical points in Keras 3: the `filepath` must end in `.keras` for a full-model save, or `.weights.h5` when `save_weights_only=True`, and the path is a format template, so `"m.{epoch:02d}-{val_loss:.3f}.keras"` gets filled from the epoch index and the logs dict. If the monitored key is absent from logs — a typo, or `val_loss` with no validation data — the callback warns and silently skips saving.

code

python · 16 lines
python
import keras

model = keras.Sequential([keras.Input(shape=(4,)), keras.layers.Dense(1)])
model.compile(optimizer="adam", loss="mse")

best_only = keras.callbacks.ModelCheckpoint(
    filepath="ckpt/best.keras",
    monitor="val_loss",
    mode="min",
    save_best_only=True,
    verbose=1,
)

every_epoch = keras.callbacks.ModelCheckpoint(
    filepath="ckpt/model.{epoch:02d}-{val_loss:.3f}.keras",
)

go deeper

for a junior

Know that without save_best_only=True you keep the last epoch, not the best one, and that monitor names a key such as val_loss or val_accuracy that must actually be produced by the run.

for a middle

Explain how mode decides the comparison direction, how the filepath template is filled from the epoch index and the logs dict, and why a missing monitor key warns and skips instead of raising.

for a senior

Show the operational side: one constant path versus templated paths and disk growth, the difference between a full .keras save and a weights-only file when resuming, and how you detect a run that quietly checkpointed nothing.

for a principal

Own the policy — what artifact a training job is contractually required to emit, whether best-only or every-N-batches suits the failure modes of your infrastructure, and how checkpoint naming and retention feed the downstream promotion process.

## What the callback does `keras.callbacks.ModelCheckpoint` is the standard way to persist a model *during* a `fit()` run rather than after it. You never call it yourself; you construct it and hand it to `model.fit(..., callbacks=[ckpt])`, and Keras invokes its hooks around each epoch and batch. Its job is narrow: decide *whether* to save right now, and *what* to write. ``` ckpt = keras.callbacks.ModelCheckpoint( filepath="ckpt/best.keras", monitor="val_loss", mode="min", save_best_only=True, ) ``` ## save_best_only, and what "best" means With the default `save_best_only=False`, the callback writes at the end of every epoch. If `filepath` is a constant string, each write overwrites the previous file, so you end up with the *last* epoch, not the best one. That is the classic beginner outcome: training visibly overfit after epoch 12, and the file on disk is epoch 40. With `save_best_only=True`, the callback keeps an internal `best` value. At the end of each epoch it reads `logs[self.monitor]` and compares. The comparison direction comes from `mode`: - `mode="min"` — an improvement is a *smaller* value. Correct for losses and error metrics. - `mode="max"` — an improvement is a *larger* value. Correct for accuracy, AUC, F1. - `mode="auto"` (the default) — Keras infers the direction from the metric name, treating names that look like accuracy/AUC as `max` and everything else as `min`. `auto` is right most of the time and wrong exactly when your metric has an unusual name. If you have a custom metric where higher is better and the name does not advertise it, set `mode` explicitly rather than trusting the guess. ## The monitor key must actually exist The monitored value is looked up in the `logs` dict Keras passes to the callback at epoch end. That dict contains the epoch's training metrics under their metric names, plus validation metrics prefixed with `val_` — but the `val_` entries only exist when `fit()` was given `validation_data` or `validation_split`. So `monitor="val_loss"` with no validation data, or `monitor="val_acc"` when the compiled metric is named `accuracy` (giving `val_accuracy`), both miss. Keras does not raise; it emits a warning saying the monitored value is not available and skips the save. A run can therefore complete with an empty checkpoint directory and no error. When a checkpoint file is mysteriously missing, check the metric name first. ## The filepath is a template `filepath` is run through Python string formatting with the epoch number and the contents of `logs`, so `"model.{epoch:02d}-{val_loss:.3f}.keras"` produces `model.07-0.312.keras`. A templated path keeps every write as a separate file — useful for inspecting a training trajectory, dangerous for disk usage on a long run. A constant path keeps exactly one file. Combining a constant path with `save_best_only=True` is the usual production choice: one file, always the best epoch. ## Keras 3 extension rules and what is inside the file Keras 3 validates the suffix. With `save_weights_only=False` (the default) the path must end in `.keras`, the backend-agnostic v3 archive; with `save_weights_only=True` it must end in `.weights.h5`. Passing the wrong suffix raises rather than guessing. The distinction matters beyond naming: a full `.keras` save carries architecture, weights and the optimizer's state, so reloading gives you an optimizer that can continue training sensibly. A weights-only file carries the weight values alone — trainable kernels and biases, plus non-trainable weights such as BatchNormalization moving statistics — but no optimizer slots, so resuming from it restarts the optimizer's momentum from scratch. ## save_freq `save_freq="epoch"` is the default. Passing an integer switches to "every N training batches", which is how you checkpoint inside a very long epoch. Note that `save_best_only` still needs a monitored value to compare, and validation metrics only refresh once per epoch, so mid-epoch best-only checkpointing on a `val_` metric compares against a stale number. ## What it does not do ModelCheckpoint writes artifacts; it does not resume anything. Restarting a killed run means loading the file yourself and passing `initial_epoch` to `fit()`, and the callback keeps no memory of its `best` value across processes — a fresh run starts with an empty best and will happily overwrite a better file from the previous run.

  • What actually happens if you set monitor to a metric name that never appears in the logs?
    Nothing fatal. At epoch end the callback looks the key up in the logs dict, finds it missing, emits a warning that the monitored value is unavailable, and skips the save for that epoch. Repeat for every epoch and the run finishes with no checkpoint written and no exception raised — which is why a silently empty checkpoint directory usually means a metric-name typo or missing validation data.
  • How do save_best_only and save_weights_only interact, and when would you pick weights-only?
    They are independent: one decides *when* to write, the other decides *what*. Weights-only makes sense when you already have the architecture in code and want small, fast writes — but it drops the optimizer state, so a resumed run restarts momentum and any optimizer accumulators from zero. For anything you intend to continue training from, save the full `.keras` model.
  • You keep one constant filepath with save_best_only=True, then restart training in a new process. What is the risk?
    The callback's internal `best` starts empty in the new process, so the first epoch of the restarted run counts as an improvement and overwrites the file — potentially replacing a genuinely better model from the earlier run. Either write to a fresh directory per run, or template the filepath so old files survive.

saying these in an interview costs you the question

  • Believes save_best_only=True is the default
  • Assumes the final epoch's weights are the best
  • Monitors val_loss with no validation data supplied
  • Uses mode='min' while monitoring val_accuracy
  • Thinks a weights-only file also stores optimizer state

context