Why does an inverted residual block expand the channel count with a 1x1 before projecting back?
answer
- narrow ends, wide middle
- reverse of the classic bottleneck
- width is cheap where cost is linear
- the last projection gets no activation
- clipping a narrow tensor is unrecoverable
basics
~20 sBecause the block's middle operator is cheap per channel, so extra width there costs little while giving the nonlinearity room to work. The expensive parts and the residual stay on the narrow ends, which keeps parameters and stored activations small.
solid answer
~50 sA classic residual bottleneck goes wide, squeezes to a narrow middle with a 1x1, does its spatial work, then expands back with a 1x1, and the skip connects the *wide* tensors. The inverted residual reverses that: the block's inputs and outputs are narrow, a 1x1 expands the channels by a small factor, the cheap spatial operator runs in that wide middle, and a second 1x1 projects back down - with the skip connecting the *narrow* tensors. This works because the middle operator's cost grows linearly with channel count rather than quadratically, so width is affordable exactly where it buys nonlinear capacity, while the two 1x1 mixing layers and the residual tensors are kept thin. The final projection is deliberately left linear: applying a rectifier right after squeezing into a low-dimensional space destroys information that the narrow representation has no spare directions to carry.
go deeper
Recall the block's profile: narrow in, expand with a 1x1, cheap spatial work in the wide middle, project back down, skip from narrow to narrow. Knowing that the last step has no activation is already a good answer at this level.
Explain both halves of the argument - width is affordable only where cost grows linearly with channels, and the nonlinearity needs a wide space to be information-preserving. Be able to contrast the profile with a classic residual bottleneck.
Show you would reason about the intermediate tensor as the memory peak, pick expansion factors per stage rather than globally, and notice when a block drops its residual because of stride or a channel change. Expect questions about what breaks when someone re-adds the activation.
Own the tradeoff between the block's capacity dial and the deployment constraint that actually binds - parameters, arithmetic, or peak activation memory - and be ready to justify why the team standardises on one block family rather than hand-tuning per model.
## Two ways to build a bottleneck A classic residual bottleneck block, as used in the deeper ResNet variants, is wide at its ends. It takes a wide tensor, uses a 1x1 convolution to *reduce* the channel count, performs the spatial convolution in that cheap narrow middle, then uses a second 1x1 to *expand* back to the wide width so the residual can be added. The design assumes the expensive operator is the spatial convolution, so the way to save is to run it on few channels. An inverted residual block turns that inside out. Its inputs and outputs are the *narrow* tensors. The first 1x1 convolution *expands* the channel count by a small factor, the spatial work happens in the wide middle, and the second 1x1 *projects* back down to the narrow width. The residual connection joins the narrow input to the narrow output. That reversal - narrow ends, wide middle - is what the word "inverted" refers to. ## Why expanding is affordable The design only makes sense when the middle spatial operator is one whose cost scales linearly with the number of channels rather than with the square of it, because each channel is processed independently rather than mixed with all the others. Under that condition, doubling or sextupling the middle width multiplies the middle operator's cost by the same small factor, not by its square. All the quadratic channel-mixing cost is concentrated in the two 1x1 convolutions, and those are attached to a narrow side, so their cost is (narrow width) x (expanded width) rather than (wide) x (wide). So the block spends its width where width is cheap and stays thin where width is expensive. That is the whole economic argument. ## Why the nonlinearity wants a wide space There is a representational argument on top of the cost argument. A rectifier-style activation zeroes every negative coordinate. In a high-dimensional space that is a mild operation: the information carried by a zeroed coordinate is typically still recoverable from the many other coordinates the representation lives in, because the useful signal occupies a lower-dimensional subspace than the tensor it sits in. Squeeze the representation down until it has almost no redundant directions, then apply the same rectifier, and the zeroing is no longer recoverable - each destroyed coordinate was carrying signal nothing else duplicates. This gives the block its second rule: put the activations in the expanded middle, where they have room, and leave the final projection **linear** - no activation after the down-projection. This is the linear bottleneck. It looks counterintuitive to anyone who has learned "more nonlinearity is better", but the narrow tensor is the block's carrier of information between blocks, and clipping it costs more than the extra nonlinearity gains. ## Where the residual attaches, and why The skip runs between narrow tensors. Two benefits follow. First, the tensor that must be kept alive from the start of the block to its end - and therefore held in memory during the forward pass, and during backpropagation - is the small one, not the expanded one. In a network built from many such blocks, peak activation memory is dominated by what crosses block boundaries, and here that is thin. Second, the added tensors are the ones the linear bottleneck has kept information-preserving, so the residual is adding two clean, uncompressed-by-a-rectifier representations. The skip is only used when it can be: the block must have stride 1 and matching input and output channel counts. A block that downsamples spatially or changes width omits the residual. ## The expansion factor as a dial The expansion factor - how many times wider the middle is than the ends - is the block's main capacity knob. A larger factor gives the nonlinearity more room and typically raises accuracy, but it multiplies the parameter and arithmetic cost of both 1x1 convolutions and inflates the intermediate activation tensor, which is the block's peak-memory term. Small integer factors are the usual choice, and the factor can reasonably differ between early blocks (where spatial resolution is high, making the wide intermediate expensive in memory) and late ones. ## The common misreads One is thinking "inverted" describes the skip connection running in some unusual direction; it does not, it describes the wide/narrow profile of the block. Another is adding an activation after the projection because it seems free - it is not free, it is the specific thing the design removes. A third is treating the expansion as a cost-free way to add capacity: the wide intermediate tensor is exactly where the block's activation memory peaks, and at high input resolution that can bind before parameters do.
- Why is there no activation after the final 1x1 projection in an inverted residual block?Because the projection lands in a low-dimensional space, and a rectifier there zeroes coordinates whose information nothing else duplicates. In a wide tensor the same clipping is survivable, since the signal occupies a subspace other coordinates still carry; in a narrow one there is no such redundancy. Leaving the bottleneck linear preserves what the block passes on, which is why the design is called a linear bottleneck.
- Where does the residual connection attach, and what does that buy?It connects the narrow input of the block to the narrow output, skipping over the expanded middle entirely. That keeps the tensor which must survive across the block small, so the peak activation memory contributed by block boundaries stays low, and it means the added tensors are the linear, information-preserving ones. The skip is only present when the block has stride 1 and equal input and output channel counts; downsampling blocks drop it.
- What does raising the expansion factor trade off?It gives the middle nonlinearity more room and usually improves accuracy, at the price of scaling both 1x1 convolutions - the block's parameter-heavy parts - and inflating the intermediate activation tensor. That intermediate is the block's memory peak, and at high spatial resolution it can be the binding constraint even when the parameter count looks fine, so early blocks often justify a smaller factor than late ones.
A workshop with a narrow doorway and a large floor: you widen the room, not the doorway, because floor space is cheap and doorways are what you pay for - and you avoid trimming anything while it is squeezed in the doorway, because there is no room to lose a piece there.
saying these in an interview costs you the question
- Says it is inverted because the skip connection runs backwards
- Adds an activation after the projection because more nonlinearity helps
- Describes it as narrowing in the middle like a classic bottleneck
- Treats the expansion as free capacity with no memory cost
- Attaches the residual to the expanded tensors
- Cannot say why widening the middle is cheaper than widening the ends