In a ResNet, when can a shortcut be a plain identity and when must it be a projection?
answer
- elementwise addition has a shape rule
- inside a stage nothing changes shape
- transitions halve size, double channels
- 1x1 convolution with matching stride
- one projection per stage, the first block
basics
~20 sA ResNet shortcut can stay a plain identity only while the branch preserves spatial size and channel count, since the two are added elementwise. Where a stage strides down and widens, it becomes a 1x1 convolution.
solid answer
~50 sThe shortcut is added elementwise, so it must match the branch output in height, width and channels. Inside a stage that holds shape constant, that is free — the shortcut is the identity, with no parameters and no compute. At a stage transition, where a stride-2 convolution halves spatial size and the channel count typically doubles, identity is impossible and you insert a **projection shortcut**: a 1x1 convolution with stride 2 mapping the old channel count to the new one. The original paper compared three policies — parameter-free identity with zero-padded extra channels, projections only where dimensions change, and projections on every shortcut. The middle one is the standard choice: projections everywhere add parameters and clutter the clean identity path for little gain. Note that in a bottleneck design the first block of a stage also needs a projection purely to match the four-fold channel expansion, even at stride 1.
go deeper
Remember that the shortcut is added, not concatenated, so both tensors must have identical shape. Know that a 1x1 convolution is what fixes a mismatch.
Explain exactly which blocks carry projections — the first of each stage — and why: stride-2 downsampling plus a channel increase, or in bottleneck designs the four-fold expansion alone.
Bring the details that matter in practice: the stride-2 1x1 sees a quarter of the positions, average-pooling before a stride-1 projection is a cheap improvement, and projections everywhere costs the exact-identity property for nothing.
Frame it as where you are willing to spend parameters on the skip path at all. Be able to argue when a backbone's stage layout should change and what that costs across training, transfer and inference latency.
## Why the question exists at all A residual block computes `y = F(x) + x`. The `+` is elementwise, so the shortcut tensor and the branch output must agree in **every** dimension: height, width and channel count. Whenever the residual branch changes any of them, a bare identity is not merely suboptimal — it does not typecheck. Inside a stage this never happens by design. A stage is a run of blocks at fixed resolution and fixed width, so the branch's convolutions use stride 1 and the same output channel count as their input, and the shortcut is a literal pass-through: no parameters, no multiply-accumulates, no memory beyond keeping `x` alive until the addition. ## Where it does happen: stage transitions A convolutional backbone reduces resolution in stages and compensates by widening. A typical transition halves height and width — the first convolution of the stage's first block uses stride 2 — and doubles channels, say from 128 to 256. The branch now emits a tensor of half the spatial size and twice the width, and the incoming `x` matches neither. The fix is a **projection shortcut**: a 1x1 convolution with stride 2, mapping `C_in` channels to `C_out`. A 1x1 kernel does no spatial mixing, so it is the cheapest way to re-map channels, and the stride handles the spatial change. Its parameter count is `C_in * C_out` — real, but small next to the 3x3 convolutions in the branch. ## The three policies, and why the middle one won The ResNet paper evaluated three options for handling dimension changes: - **Parameter-free identity.** Subsample the shortcut spatially (take every other position) and pad the missing channels with zeros. Costs nothing, adds no parameters, but the new channels enter with no signal at all and the subsampling is crude. - **Projections only where dimensions change.** Identity everywhere else. Slightly better accuracy than the parameter-free option, at a very small parameter cost. - **Projections on every shortcut.** Marginally better again, but every block now pays parameters and compute on the shortcut, and the identity path — the thing that made the whole design work — is no longer an identity anywhere. The middle policy is what deployed ResNets use. The reasoning generalizes: keep the shortcut an exact identity wherever you can, and spend parameters on it only where the shape forces you to. The clean path is the asset. ## The bottleneck wrinkle In a bottleneck design the block's output width is four times its internal 3x3 width. So the **first** block of every stage needs a projection even when it does not downsample: at the earliest stage, a block may take 64 channels in and emit 256, at stride 1. The channel mismatch alone forces the 1x1 projection. Later blocks in that stage take 256 in and emit 256, so they use plain identity shortcuts. A good candidate can say which blocks in a backbone carry projections: exactly one per stage, the first. ## A known wart in the downsampling shortcut A 1x1 convolution with stride 2 looks at one position in every 2x2 window and ignores the other three. Three quarters of the input is simply not seen by the shortcut. Common refinements move the stride off the 1x1 — for example, apply a 2x2 average pooling first and then a stride-1 1x1 projection, so the shortcut summarizes the window rather than sampling one corner of it. A related tweak inside the branch places the stride on the 3x3 convolution rather than the leading 1x1, for the same reason. These are cheap changes that tend to help slightly; they are exactly the kind of detail an interviewer uses to tell reading from experience. ## Practical consequences - **Debugging.** A shape-mismatch error at the addition almost always means you changed stride or width in the branch and forgot to update the shortcut. The shortcut's stride must equal the product of the branch's strides. - **Transfer and surgery.** If you widen a stage or change its stride when adapting a backbone, the projection is the layer you must re-create; the identity blocks are unaffected. - **Counting cost.** Projection shortcuts are a small share of a backbone's parameters but they sit on the critical path at every stage boundary, so they are not free at inference. - **Do not make everything a projection** as a defensive habit. You lose the exact-identity property, gain parameters and gain nothing measurable.
- How many projection shortcuts does a four-stage ResNet contain?One per stage, in that stage's first block — four in total for a standard four-stage backbone. Every other block keeps a parameter-free identity, because within a stage the resolution and width are constant. If someone answers "one per block", they are describing the projections-everywhere variant, which is not the standard configuration.
- Why not solve the width mismatch by zero-padding the shortcut's extra channels instead?It works and costs no parameters — it was the paper's option A. But the added channels start with no information flowing along the shortcut, and the spatial subsampling is a crude every-other-pixel pick. The measured accuracy is slightly worse than a 1x1 projection, whose parameter cost is negligible next to the branch.
- You change a stage's branch to stride 2 and training crashes at the addition. What went wrong?The shortcut still has stride 1, so the two tensors disagree spatially and the elementwise add fails. The shortcut's stride must equal the product of the strides inside the branch, and its output channel count must equal the branch's. Fix the projection, not the branch.
saying these in an interview costs you the question
- Thinks the shortcut can add tensors of different shapes
- Says every ResNet shortcut is a 1x1 convolution
- Puts a 3x3 convolution on the shortcut path
- Forgets the channel-expansion mismatch in bottleneck blocks
- Believes projections are needed only when stride changes