In Keras, when is Sequential not enough and you need the Functional API?
answer
- one straight line versus a graph
- branches, merges, several inputs
- symbolic placeholders, then a model
- same layer instance called twice
- shape known before data arrives
basics
~20 skeras.Sequential models a straight stack: one input, one output, each layer feeding the next. Anything with branches, several inputs or outputs, skip connections, or one layer instance reused twice needs the Functional API, which wires an explicit graph of layer calls.
solid answer
~50 s`keras.Sequential` takes a list of layers and chains them, so it can only express a linear pipeline with exactly one input tensor and one output tensor. The Functional API is declarative instead: you create symbolic entry points with `keras.Input(shape=...)`, call layers on those symbolic tensors to record a graph, and close it with `keras.Model(inputs=..., outputs=...)`. Because each call returns a symbolic `KerasTensor` rather than a value, you can fan out, merge branches with `keras.layers.Concatenate` or `keras.layers.Add`, hand the same layer instance two different tensors to share its weights, and return a list of outputs. In Keras 3 both are still fully declarative, so both give you `model.summary()` with real output shapes, `model.get_layer()`, and shape errors reported while you build rather than at the first batch. Reach for Functional the moment the architecture stops being a line.
code
python · 9 linesimport keras
model = keras.Sequential([
keras.Input(shape=(28, 28, 1)),
keras.layers.Conv2D(32, 3, activation="relu"),
keras.layers.GlobalAveragePooling2D(),
keras.layers.Dense(10, activation="softmax"),
])
model.summary()go deeper
Be able to say plainly that Sequential is a single straight stack with one input and one output, and that branches, multiple inputs or multiple outputs mean the Functional API. Know that Functional models start from keras.Input.
Explain that layer calls on a keras.Input return symbolic tensors that record a graph, that keras.Model(inputs, outputs) closes it, and that this static graph is what makes summary, shape checking and layer surgery possible.
Show you have hit the limits in real work: merged multi-modal inputs, residual connections, shared towers, and knowing when the declarative graph stops being expressive enough and a custom layer or subclassed model takes over.
Own the readability and tooling argument. Discuss why a declarative topology is easier to review, export and instrument than imperative code, and where you draw the line so a codebase does not accumulate three styles of model definition.
## What Sequential actually is `keras.Sequential` is a thin convenience wrapper over the same machinery the Functional API uses. You hand it an ordered list of layers, and it wires layer *n*'s single output into layer *n+1*'s single input. That contract is the whole story: **one input tensor, one output tensor, no branching**. If you pass `keras.Input(shape=...)` as the first element, Keras knows the input shape immediately and builds every layer's weights at construction time; without it, the stack stays unbuilt until the first batch or an explicit `model.build()`. Sequential is the right tool more often than people admit — a convolutional stack ending in a classifier head, an MLP, a preprocessing chain. It reads well and it is hard to get wrong. ## What the Functional API adds The Functional API is a small symbolic language. `keras.Input(shape=(32,))` does not allocate data; it returns a `KerasTensor`, a placeholder that carries a shape and a dtype but no values. Calling a layer on a `KerasTensor` returns another `KerasTensor` and records a node in a graph: *this layer was applied to that tensor*. When you finally write `keras.Model(inputs=..., outputs=...)`, Keras walks backwards from the outputs to the inputs and freezes that graph as the model's topology. Because the graph is data, not control flow, four things become possible that Sequential cannot express: 1. **Multiple inputs.** An image tower and a tabular tower merged before the head. `keras.Model(inputs=[image, meta], outputs=...)`. 2. **Multiple outputs.** One trunk, two heads — say a regression output and a classification output — returned as a list or dict. 3. **Non-linear topology.** Residual/skip connections (`keras.layers.Add()([x, shortcut])`), inception-style parallel branches, U-Net style concatenations from encoder to decoder. 4. **Shared layers.** Call the *same* layer instance on two different tensors and both calls use the same weights — the standard way to build siamese/twin-tower models. ## Why the explicit graph is worth having Because the topology is known before any data flows, Keras can propagate shapes through it. That buys you concrete, immediate benefits: - `model.summary()` prints each layer's real output shape and parameter count. - Shape mismatches raise while you are typing the architecture, not on the first training batch. - `keras.utils.plot_model(model)` can draw the graph. - `model.get_layer("name").output` gives you a handle on any intermediate tensor, so you can build a feature extractor or splice a new head onto an existing trunk without touching the original code. - The topology is plain configuration, so it round-trips through save/load without custom code. ## The failure mode to recognise The usual interview trap is a candidate trying to force a branch into Sequential — for example calling `model.add()` twice hoping for two heads, or trying to add two Sequential models together. Sequential simply has nowhere to put a second output. The second trap is treating a `KerasTensor` like real data inside the Functional build: `if x > 0:` or `x.numpy()` on a symbolic tensor fails, because there is no value yet. Data-dependent control flow is exactly the point where the declarative APIs run out and a subclassed model or a custom layer takes over. ## Choosing in practice A workable rule: start Sequential; move to Functional the first time the architecture stops being a line; keep the novel arithmetic inside a custom layer that the Functional graph then wires up. Converting a Sequential model to Functional is mechanical — you already have the layer list, you just name an input and thread it through — so starting with Sequential costs nothing. One last note on Keras 3: both APIs are backend-agnostic. The same Sequential or Functional model definition runs on the TensorFlow, JAX or PyTorch backend, selected by the `KERAS_BACKEND` environment variable, because the layer graph is Keras's own representation and not a backend graph.
- What exactly does a layer call return when you apply it to a keras.Input?A `KerasTensor` — a symbolic placeholder carrying shape and dtype but no values. The call records a node in the layer graph rather than computing anything. That is why you cannot inspect its values or branch on them with a Python `if` while building; the graph is only executed later, when real data flows through the model.
- How do you share weights between two towers of a Functional model?Instantiate the layer once and call it on both tensors. `shared = keras.layers.Dense(64)`, then `a = shared(left)` and `b = shared(right)` — both calls use the same weight objects, so the tower is trained jointly. Creating two `Dense(64)` instances instead would give two independent weight sets, which is the classic siamese-model bug.
- Can you nest a Sequential model inside a Functional graph?Yes. A `keras.Model` is itself a layer, so a Sequential trunk can be called on a `KerasTensor` like any other layer and its output wired onward. This is the normal way to reuse a pretrained backbone: build the Functional graph around it and attach new heads, while the nested model keeps its own layers and weights.
Sequential is a single-track railway line; the Functional API is a rail network where lines split, merge and share the same platform.
saying these in an interview costs you the question
- Claiming Sequential supports multiple inputs if you call add() twice
- Saying the Functional API is only for very large models
- Thinking a Functional model trains faster than a Sequential one
- Calling two Dense layers with equal size "shared weights"
- Trying to use a Python if on a symbolic KerasTensor while building