Why does a residual block add its input back to its output instead of just stacking layers?
answer
- the deeper plain net trained worse
- not overfitting, an optimization failure
- identity should be the cheap default
- F(x) + x, drive F toward zero
- shortcut derivative is one, gradients add
basics
~20 sA residual block computes F(x) + x, so its layers only learn the change to make to the input. Behaving like an identity then just means pushing F toward zero, which plain stacked layers struggle to fit.
solid answer
~50 sDeep plain CNNs suffer a degradation problem: a 56-layer plain stack reached higher **training** error than a 20-layer one, so this is an optimization failure, not overfitting. A residual block reframes the layers as `y = F(x) + x`: the weights model the residual, the correction to the input, rather than the whole mapping. That makes the identity the cheap default — driving `F` toward zero is far easier for gradient descent than fitting an identity out of a stack of convolutions and nonlinearities — so extra depth cannot easily make things worse. The addition also helps backwards: the shortcut's local derivative is one, so the gradient arriving at the block's input is the incoming gradient plus what came through `F`, rather than only a product of layer terms. That is why the pattern unlocked networks a hundred layers deep and beyond.
go deeper
Be ready to state the formula y = F(x) + x, say that the layers learn a correction to the input, and name the degradation result: the deeper plain network trained worse, which is not overfitting.
Explain both halves of the argument — the identity becomes trivially representable in the forward pass, and the shortcut's derivative of one gives an additive gradient path back to earlier layers. Say why addition, not concatenation.
Show you know the limits: shortcuts fix optimization, not capacity, and depth still has diminishing returns. Mention starting blocks as exact identities via a zero-initialized residual scale when training very deep stacks.
Own the design tradeoff: residual topology lets you over-provision depth cheaply and prune later, but every block you add costs inference latency forever. Be able to argue depth versus width versus resolution as a budget decision.
## The problem it was invented for In principle, making a network deeper should never hurt the error it can reach on the **training set**. A deeper network contains the shallower one as a special case: keep the first 20 layers, make the extra 36 layers compute the identity, and you have exactly the shallow model. So the deeper model's best achievable training error is at most the shallow model's. Empirically, plain stacks do not behave that way. A 56-layer plain convolutional network trains to *higher* training error than a 20-layer one built the same way. This is the **degradation problem**, and the word that matters is *training*. Overfitting would look like low training error with high test error; here the training error itself is worse. The deeper network is not failing to generalize, it is failing to optimize — gradient descent cannot find the solution we know exists. The reason the argument above fails in practice is that "make these layers compute the identity" is easy to *state* and hard to *fit*. A stack of convolutions with nonlinearities in between has no natural setting of its weights that reproduces its input exactly, so the optimizer has to discover one, in a high-dimensional space, while also fitting the data. ## The reformulation A residual block changes what the weights are asked to represent. Instead of a stack computing some desired mapping `H(x)` directly, the block computes ``` y = F(x) + x ``` where `F` is the stacked layers (in the basic form: 3x3 conv, nonlinearity, 3x3 conv) and `x` reaches the output untouched along a **shortcut** — a parameter-free elementwise addition, not a concatenation. The trainable part now models `H(x) - x`: the *residual*, the change to make to the input. The payoff is that the identity is now the cheap default. If the best thing a block can do is nothing, the optimizer only has to drive `F`'s output toward zero — shrinking weights toward zero is exactly the direction gradient descent finds easily. Depth stops being a liability: a block that has nothing useful to add can quietly get out of the way. In practice, the learned residuals in a trained network are small, which is evidence that most blocks really are making modest refinements to a signal that is largely carried by the shortcut. A common initialization trick leans on this directly: scale the residual branch's output by a learnable factor initialized at zero, so every block starts as an exact identity and the network begins training as if it were shallow, then grows into its depth. ## The gradient view Differentiate the block. With `y = F(x) + x`, the derivative of the output with respect to the input is `F'(x) + 1` (in the vector case, the branch Jacobian plus the identity matrix). So the gradient handed back to earlier layers is ``` g_in = g_out * (1 + F'(x)) = g_out + g_out * F'(x) ``` The first term is the incoming gradient passed through **unchanged**. Whatever happens inside the branch, that additive term survives, so there is always a route by which the training signal reaches early layers without being scaled down repeatedly. This is the "gradient highway" argument, and it is a *sum*, not a product — that is the whole point of using addition rather than a multiplicative gate. Be careful not to over-claim it. Residual connections do not make gradients constant, and they are not a substitute for sane initialization or normalization; they change the shape of the optimization landscape so that a good solution is reachable. ## The unravelled-ensemble reading A useful second interpretation: expand `y = F(x) + x` recursively over a stack of blocks and you get a sum over every subset of blocks — a network of `n` residual blocks is, algebraically, an ensemble of `2^n` paths of varying length. Evidence for this reading is that you can delete a trained residual block at test time and performance degrades gracefully, whereas deleting a layer from a plain stack is catastrophic. **Stochastic depth** exploits it during training: randomly drop whole residual blocks (their branch only — the shortcut always remains), which shortens the effective network, regularizes, and speeds up training. The mental model is less "one very deep chain" and more "many shallow paths that share weights". ## What it does not fix Residuals address optimization, not capacity or generalization. They do not make a model immune to overfitting, they do not remove the need for good data or augmentation, and they do not by themselves let you scale depth without limit — very deep residual stacks still show diminishing returns per added block. And the addition imposes a hard constraint: the shortcut tensor and the branch output must have the same shape, which is why stage transitions need a projection.
- Is the degradation that shortcuts fix a form of overfitting?No. Overfitting means low training error with a gap to test error. Here the deeper plain network's *training* error is worse than the shallow one's, so the model is not even fitting the data it can see. It is an optimization failure — the solution exists in the parameter space and gradient descent does not find it.
- Why addition rather than concatenating the shortcut onto the branch output?Addition keeps the block's output shape identical to its input, so blocks stack indefinitely at fixed width, and it gives the shortcut a local derivative of exactly one. Concatenation grows the channel count with every layer, so later layers get wider and the block is no longer a drop-in repeat unit.
- What does stochastic depth suggest about how a deep residual stack actually behaves?You can drop whole residual branches at random during training, and even delete blocks from a trained network, with only graceful degradation. That points to the unravelled view: the stack behaves like an ensemble of many shorter paths sharing weights, not one rigid chain where every layer is load-bearing.
- If a block's best behaviour is the identity, why keep it at all rather than making the network shallower?You do not know in advance which blocks are useful, and it is data-dependent. The point of the reformulation is that you can afford to over-provision depth: blocks that earn their keep learn a non-zero residual, and the rest cost compute but not accuracy. Prune later if inference cost matters.
Editing a draft instead of rewriting it from scratch. If the draft is already correct, the edit is "change nothing" — trivial. Rewriting the whole page from memory just to reproduce it is far harder.
saying these in an interview costs you the question
- Says residual connections mainly prevent overfitting
- Claims the shortcut branch has its own weights
- Says the layers learn the output rather than the correction
- Thinks the gradient is multiplied, not added, along the shortcut
- Confuses the additive shortcut with concatenating features