Why does a stack of fully-connected layers with no nonlinearity collapse to one layer?
answer
- composition of two maps of one kind
- closed under composition
- the biases compose into one bias
- product of the weight matrices
- narrow middle means limited rank
basics
~20 sComposing affine maps yields another affine map: W2(W1x + b1) + b2 equals (W2W1)x + (W2b1 + b2). Depth with no nonlinearity between the layers buys more parameters and a different optimisation path, never a larger family of functions.
solid answer
~40 sEach fully-connected layer with no activation computes an affine map, and affine maps are closed under composition. Two of them give `W2(W1x + b1) + b2 = (W2 W1)x + (W2 b1 + b2)`, a single layer with weight matrix `W2 W1` and bias `W2 b1 + b2`, and the argument telescopes to any depth. So a ten-layer stack with nothing between the layers represents exactly the same functions as one layer of the same widths. The symptom is unmistakable: a three-layer network on a two-input parity task scoring identically to a plain linear model, run after run, however wide you make it. The fix is a nonlinearity *between* consecutive linear maps. Note too that a narrow middle layer makes `W2 W1` low-rank, so such a stack is strictly weaker than one unrestricted layer.
go deeper
Be ready to say that two activation-free layers multiply into one weight matrix, so depth alone adds nothing, and that a nonlinearity between layers is what makes a network more than a linear model.
Expect to derive the composite weight and bias on the board, name the maps as affine, and explain why bias terms do not introduce any curvature into the function.
Demonstrate the diagnosis: a deep model matching a linear baseline exactly, run after run, points at a missing nonlinearity long before it points at learning rate or capacity. Know how to confirm it by comparing against a fitted single linear layer.
Own the distinction between parameter count and expressive power when reviewing designs and budgets. Low-rank factorised layers are a deliberate compression trade, and a team that reads them as capacity will keep buying compute that cannot help.
## The algebra A fully-connected layer without an activation computes `y = W x + b` for a weight matrix `W` and a bias vector `b`. This is an *affine* map. Compose two of them: ``` y = W2 (W1 x + b1) + b2 = (W2 W1) x + (W2 b1 + b2) ``` Define `W' = W2 W1` and `b' = W2 b1 + b2`. The composite is `y = W' x + b'` — one affine map. The argument applies again to `W' ` and a third layer, and by induction to any depth: an arbitrarily deep stack of activation-free layers is one affine map from input to output. This is a statement about the *set of representable functions*, and it is exact, not approximate. Depth here is not a weak improvement; it is zero improvement in expressive power. ## Why the biases do not save it A common wrong answer is that biases introduce nonlinearity. They do not. A bias shifts, and shifts compose into shifts: the composite bias is `W2 b1 + b2`, still just a vector added at the end. The map remains affine, its level sets remain flat, and its decision boundary — if you thresholded the output — remains a hyperplane. ## Why width does not save it either Making the hidden layers wider adds parameters but changes nothing about the function class, because the product `W2 W1` is still a single matrix of the input-to-output shape. Worse, going *narrow* actively hurts. If `W1` maps 100 inputs to 5 hidden units and `W2` maps those 5 to 100 outputs, then `rank(W2 W1) <= 5`. The composite is not merely an affine map; it is a *rank-constrained* affine map. A single unrestricted layer of the same input/output widths can express strictly more. This is exactly why low-rank factorisation is a legitimate compression trick — you deliberately trade capacity for parameters — and exactly why it is not a source of capacity. ## The symptom in practice The canonical way this shows up: someone builds a three-layer network for a two-input parity task — output 1 when exactly one of two binary flags is on — trains it carefully, tunes the learning rate, widens the hidden layers, and reports that accuracy sits at chance-like 50–75% and refuses to move, matching a plain linear model exactly. The instinct is to add depth or capacity. The correct diagnosis is that no nonlinearity is applied between the layers, so the whole network *is* a linear model; the target function is not affine, and no setting of the parameters can reach it. Diagnostics that confirm it quickly: multiply the learned weight matrices together and check that the product reproduces the network's outputs on held-out inputs; or fit a single linear layer and observe that it matches the deep model's loss almost exactly, run after run. ## Where the nonlinearity must go Between *consecutive* linear maps. A frequent half-fix is to apply a nonlinearity only at the output — for example squashing the final score into a probability. That leaves the entire hidden body collapsible: the network is one affine map followed by one squashing function, which is a generalised linear model, not a deep one. Its decision boundary is still a hyperplane in the input space. Every hidden layer needs its own nonlinearity for depth to mean anything. ## The nuance worth knowing Identical function class does not mean identical training. A deep linear stack is a different *parameterisation* of the same function set, and reparameterisation changes the loss surface's geometry, the conditioning of the problem and therefore the trajectory that gradient descent takes. Deep linear networks are studied precisely because they have non-trivial training dynamics while having trivial expressive power. But the ceiling is the ceiling: whatever path training takes, the endpoint is an affine map, so a task that requires a curved decision boundary remains out of reach. ## How to answer the interview version State the composition identity with the explicit composite weight and bias. Say the word *affine*, and note that biases compose into a bias. Add the rank observation about narrow middle layers to show you understand the collapse can be a restriction rather than a wash. Close with the operational point: the fix is a nonlinearity between every adjacent pair of linear maps, not more layers and not more units.
- Does making the hidden layers wider or narrower change that conclusion?Wider changes nothing about expressive power — the product of the weight matrices is still one matrix of the same input-to-output shape. Narrower makes it strictly worse: if the middle layer has 5 units, the composite matrix has rank at most 5, so the stack expresses less than a single unrestricted layer. That is why low-rank factorisation is used as a compression technique, and never as a way to add capacity.
- If the function class is identical, is training a deep linear stack the same as training one layer?No. It is a different parameterisation of the same function set, so the loss surface, its conditioning and the trajectory gradient descent follows all differ — deep linear networks have genuinely non-trivial training dynamics. But the reachable functions are unchanged, so on a task that needs a curved boundary the deep stack cannot beat the single layer no matter how it trains.
- Someone applies a nonlinearity only at the output layer. Is that enough?No. Every hidden layer collapses into one affine map, so the network is a single affine map followed by one squashing function — a generalised linear model whose decision boundary is still flat in the input space. Depth only counts when a nonlinearity sits between each adjacent pair of linear maps.
saying these in an interview costs you the question
- Claims more layers always means more capacity
- Says the bias terms make the stack nonlinear
- Thinks an output nonlinearity rescues the hidden layers
- Believes the collapse only happens for equal-width layers
- Equates more parameters with more expressive power