skip to content

Why does model.summary() fail on a subclassed Keras model before it runs?

level: middleimportance: must knowfreq 64%

answer

  1. no declared input means no known shapes
  2. weights are created lazily, not in init
  3. the built flag flips on first data
  4. build takes a batch dimension of None
  5. Functional pins shapes at construction

basics

~20 s

A subclassed keras.Model has no declared input, so Keras cannot know any shapes until data flows through call(). Until then the model is unbuilt, its sublayers have no weights, and summary() raises a ValueError. Calling model.build(input_shape=...) or running one batch fixes it.

solid answer

~50 s

Sequential and Functional models start from `keras.Input`, which pins the input shape, so Keras propagates shapes through the whole topology at construction time. A subclassed model has no such declaration: `call()` is arbitrary Python, and Keras cannot tell what shapes it will receive without executing it. So the model reports `model.built == False`, its sublayers have not created their weight tensors yet, and `model.summary()` raises a ValueError saying the model has not been built. Two things fix it — call the model once on a real batch, or call `model.build(input_shape=(None, 32))` with the batch dimension as `None`. This deferred build is not just a cosmetic problem: because parameters do not exist yet, an unbuilt model also has nothing meaningful to hand an optimizer or to save weights from. In Keras 3 the same rule applies on all three backends, since the build state belongs to Keras rather than to the backend.

code

python · 19 lines
python
import keras
import numpy as np


class MLP(keras.Model):
    def __init__(self):
        super().__init__()
        self.hidden = keras.layers.Dense(64, activation="relu")
        self.out = keras.layers.Dense(10)

    def call(self, inputs):
        return self.out(self.hidden(inputs))


model = MLP()
print(model.built)               # False - no shapes known yet
model.build(input_shape=(None, 32))   # or: model(np.zeros((1, 32)))
print(model.built)               # True
model.summary()

go deeper

for a junior

Know that a subclassed model has no weights until it sees data or you call build(), and that summary() errors until then. Be able to name both fixes: one dummy batch, or model.build(input_shape=...).

for a middle

Explain that layers create weights in build(input_shape), not init, and that keras.Input is what lets Sequential and Functional models infer every shape up front while an arbitrary call() cannot be inspected without running it.

for a senior

Discuss the downstream consequences — optimizer variable lists, weight saving, missing per-layer output shapes and lost graph surgery — and the habit of forcing an explicit build so shape errors surface before a training run burns time.

for a principal

Frame it as a contract question: an imperative model defers its shape contract to runtime, so the team must supply the discipline the framework no longer enforces, through build calls, smoke tests and reviewed shape assertions.

## Declarative versus imperative model definitions Keras has two fundamentally different ways of knowing what a model looks like. **Declarative** — Sequential and Functional. You begin with `keras.Input(shape=...)`, which fixes the per-sample shape. Every subsequent layer call happens on a symbolic tensor, so Keras can ask each layer "given this input shape, what is your output shape?" and thread the answer forward. By the time `keras.Model(inputs, outputs)` returns, every layer's output shape and parameter count is known and every weight has been created. **Imperative** — subclassing `keras.Model` and writing `call(self, inputs)`. The forward pass is ordinary Python. It may contain loops, conditionals, or shape manipulation that depends on the data. Keras has no way to reason about that without running it, and it has nothing to run it *on*, because no input shape was ever declared. ## What "built" means A Keras layer creates its weights in `build(input_shape)`, not in `__init__`. A `Dense(64)` knows it wants 64 output units, but the kernel it needs is shaped `(input_features, 64)` — and `input_features` is unknown until something arrives. So weight creation is deferred, and each layer carries a `built` flag. A model is built when its own build has run and its sublayers' weights exist. This is why the sequence below is the normal lifecycle of a subclassed model: 1. `__init__` — sublayers are instantiated and tracked as attributes. No weights exist yet. `model.built` is `False`. 2. First call with real data, **or** an explicit `model.build(input_shape)` — shapes become known, weights are created, `model.built` becomes `True`. 3. From that point `model.summary()`, `model.weights`, `model.count_params()` and weight saving all work. `model.summary()` before step 2 raises a `ValueError` telling you the model has not been built and asking you to build it or call it on a batch of data. That message is the single most common first encounter developers have with the concept. ## Building explicitly `model.build(input_shape=(None, 32))` takes the **full** shape including the batch dimension, with `None` for the batch. That trips people up, because `keras.Input` takes the shape **without** it. Two adjacent APIs, two conventions: - `keras.Input(shape=(32,))` — per-sample shape. - `model.build(input_shape=(None, 32))` — batch-inclusive shape. For a model with several inputs, you pass the corresponding structure of shapes. ## Why the difference matters beyond summary() Even after a subclassed model is built, it never gains the *static graph* that Functional models have. Shape inference happened once, for one concrete input shape, by executing the forward pass. The consequences: - `model.summary()` on a built subclassed model can print layers and parameter counts, but per-layer output shapes are frequently reported as unknown, because there is no recorded node graph mapping tensors between layers. - `keras.utils.plot_model` has little topology to draw. - `model.get_layer("name").output` has nothing to return — the layer was never called in a recorded graph, only inside Python. - Shape mistakes surface at the first forward pass rather than at definition time, so a typo in a hidden dimension costs you a run rather than a line of feedback. ## Practical habits - Give subclassed models an explicit `build()` call, or a smoke test that pushes one dummy batch through them, right after construction. It turns a class of runtime errors into an import-time error. - Do not create weights inside `call()` by instantiating layers there; layers must be attributes created in `__init__` so Keras can track them. A layer created per call produces a model whose parameter count changes every step and whose optimizer state is meaningless. - Remember the optimizer needs the variable list. Attaching an optimizer to a model whose weights do not exist yet is a common source of confusing errors; building first removes the ambiguity. - If you need `summary()` and shape output in logs for a subclassed model, the practical fix is usually structural: keep the novel piece as a custom `Layer` and assemble the model itself with the Functional API, which restores full shape inference. ## The version note This is Keras 3 behaviour and it is backend-independent. The built/unbuilt distinction is Keras's own bookkeeping, so switching `KERAS_BACKEND` between TensorFlow, JAX and PyTorch does not change when weights are created or when `summary()` starts working.

  • Why does build() take (None, 32) while keras.Input takes (32,)?
    `keras.Input` describes a single sample, so the batch dimension is implicit and always flexible. `build(input_shape=...)` receives the shape a layer will actually see, which includes the batch axis, and `None` marks it as variable. The mismatch is a genuine API wrinkle worth stating out loud in an interview rather than guessing.
  • Even after building, why can summary() still show unknown output shapes for a subclassed model?
    Building runs the forward pass once to create weights, but it does not record a node graph linking layer outputs to layer inputs. Without that topology Keras has no stored mapping from one layer's output to the next layer's input, so it can report parameter counts but often not per-layer output shapes. Functional models keep that graph, which is why they always can.
  • What goes wrong if you instantiate a layer inside call() instead of __init__?
    A fresh layer with fresh weights is created on every forward pass. Nothing is tracked as a model parameter, the optimizer never sees consistent variables, and the model cannot learn. Layers must be attributes assigned in `__init__` so Keras can track them, with only their invocation happening in `call()`.

saying these in an interview costs you the question

  • Thinking weights are created in the layer constructor
  • Passing build() a shape without the batch dimension
  • Believing summary() triggers the build itself
  • Claiming subclassed models can never print a summary
  • Saying the behaviour differs between the TensorFlow and PyTorch backends

context