skip to content

When would you use Keras save_weights instead of saving the whole model?

level: middleimportance: should knowfreq 58%

answer

  1. values only, no architecture
  2. the filename suffix is enforced
  3. target model must already be built
  4. skip_mismatch is the transfer-learning switch

basics

~20 s

Use save_weights when the architecture already lives in code you trust: it writes only variable values to a .weights.h5 file, so loading needs a model that is already built with the identical structure. Whole-model .keras saving is for artifacts that must reconstruct themselves.

solid answer

~40 s

`model.save_weights("ckpt.weights.h5")` stores variable values only, and Keras 3 requires that exact `.weights.h5` suffix. Reloading is `model.load_weights("ckpt.weights.h5")` on a model you have already constructed and built, because the values are matched into variables that must already exist - a subclassed model that has never been called has no variables yet and the load fails on a mismatch. That is the tradeoff: weights-only files never re-instantiate your classes, so they sidestep serialization of custom code entirely, but they are useless without the build code. I reach for them for in-run checkpoints and for transfer learning, where `skip_mismatch=True` lets a partially different architecture take the layers that do line up. For a hand-off artifact, `.keras` is the right call.

code

python · 9 lines
python
import keras

model = keras.Sequential([keras.layers.Dense(8), keras.layers.Dense(1)])
model.build((None, 4))
model.save_weights("ckpt.weights.h5")

same = keras.Sequential([keras.layers.Dense(8), keras.layers.Dense(1)])
same.build((None, 4))
same.load_weights("ckpt.weights.h5")

go deeper

for a junior

Remember that a weights file holds numbers only, that the filename has to end in .weights.h5, and that you must construct the same model in code before loading into it.

for a middle

Explain the build requirement: variables are created lazily in subclassed models, so load_weights has nothing to fill until build() or a first forward pass has run.

for a senior

Show judgment about which artifact to hand over. Weights files are fine inside one run next to pinned code; anything crossing a process or team boundary should be a whole-model archive.

for a principal

Own the drift risk. Weights-only checkpoints couple an artifact to a source revision, so decide how that pairing is recorded and enforced before a partial load quietly degrades a production model.

## Two files, two contracts A `.keras` archive answers "rebuild this model". A weights file answers "here are the numbers". The second is only meaningful next to code that produces the same variables in the same shapes. `model.save_weights(filepath)` in Keras 3 insists the filename end in `.weights.h5`; anything else raises. The file is HDF5 containing the model's variable values, keyed by their position in the layer tree. ## Why loading needs a built model `model.load_weights(path)` does not create layers. It walks the existing model and assigns saved values into variables that are already there. A Sequential or Functional model built with an input shape already has its variables, so loading works right after construction. A subclassed model creates its variables lazily on the first `build()` or the first forward pass, so this fails: ``` model = MyModel() model.load_weights("ckpt.weights.h5") # nothing to load into ``` Call `model.build(input_shape)` or run one batch through it first. The error is about a structure or variable mismatch, and the cause is almost always "the model was not built", not "the file is bad". `load_weights` also accepts `skip_mismatch=True`, which skips layers whose weight count or shapes do not line up instead of raising. That is the transfer-learning switch: load a backbone's weights into a model whose head has a different number of classes and let the head stay randomly initialized. Use it deliberately - with it on, a wholesale mismatch loads almost nothing and trains silently from noise. ## When each one wins Weights-only is the better fit when: - The architecture lives in versioned source you always have. Then the archive's config adds nothing. - You have custom layers and do not want to make them serializable yet. A weights file never re-instantiates a class, so registration and get_config do not enter the picture. - You want frequent, cheap in-run checkpoints and will restart from the same script. - You are doing partial or cross-model loading, where skip_mismatch matters. Whole-model `.keras` is the better fit when: - Someone else, or some other process, must load the model without your training script. - You want optimizer state for resuming. - The artifact has to outlive the repository revision that built it. ## The failure mode to name in an interview The dangerous case is a weights file whose architecture code has drifted: someone changes a layer width, some shapes still match, `skip_mismatch=True` is set from an old copy-paste, and the model loads partially. Nothing raises; evaluation is merely bad. Guard it by keeping weights files next to the exact code revision that wrote them, by logging how many layers were actually restored, or by preferring `.keras` for anything you will not reload within the hour.

  • What does skip_mismatch=True actually skip, and when is it dangerous?
    It skips any layer whose weight count or shapes disagree with the file instead of raising. That is useful when reusing a backbone under a new head. It is dangerous because a badly mismatched pair now loads almost nothing quietly, so the model trains from near-random values while looking loaded. Log how many layers were actually restored.
  • A subclassed model raises on load_weights immediately after construction. What is the fix?
    Build it first. A subclassed model creates variables lazily, so right after __init__ there is nothing to load into. Call model.build(input_shape) or push one batch through it, then load. The same code works unchanged for Sequential or Functional models because those build eagerly once given an input shape.

saying these in an interview costs you the question

  • Expecting load_weights to recreate the layers
  • Saving weights to model.h5 and expecting it to work
  • Turning on skip_mismatch to silence a real error
  • Shipping a weights file with no architecture code

context