Why is a depthwise-separable convolution roughly nine times cheaper than a dense 3x3 layer?
answer
- two jobs done in two passes
- one spatial filter per input channel
- the 1x1 does all the mixing
- two terms, k squared and C_out
- at k=3 it lands near one ninth
basics
~20 sIt factorises a dense layer into a spatial stage with one k x k filter per input channel and a pointwise stage that mixes channels. Cost falls to 1/C_out + 1/(k*k) of dense, about 1/9 at k=3.
solid answer
~50 sA dense `k x k` layer holds `k*k*C_in*C_out` weights because every filter spans every input channel. A separable version splits that job in two: a depthwise stage with one `k x k` filter per input channel, costing `k*k*C_in`, then a pointwise 1x1 stage costing `C_in*C_out` that recombines the channels. The ratio is `(k*k*C_in + C_in*C_out) / (k*k*C_in*C_out) = 1/C_out + 1/(k*k)`. For a 3x3 layer with 128 input and 256 output channels that is `1/256 + 1/9 = 0.115`, so 33,920 weights instead of 294,912 — 8.7x fewer, and the multiply-accumulates fall by the same ratio since both stages run at the same resolution. With `k = 3` the `1/(k*k)` term dominates, which is why the answer is "about nine times" for any reasonably wide layer. The capacity you give up is the ability to learn spatial and cross-channel structure jointly in one filter; the factorisation assumes the two can be learned separately.
code
python · 11 linesk, c_in, c_out = 3, 128, 256
dense = k * k * c_in * c_out # one k x k plane per (in, out) pair
depthwise = k * k * c_in # one k x k filter per input channel
pointwise = c_in * c_out # one weight per (in, out) pair
separable = depthwise + pointwise
print(dense, separable) # 294912 33920
print(round(dense / separable, 2)) # 8.69 -> the "about 9x"
print(round(separable / dense, 5)) # 0.11502
print(round(1 / c_out + 1 / (k * k), 5)) # 0.11502 -> same ratiogo deeper
Recall the two stages in order: a per-channel k x k filter that does not mix channels, then a 1x1 that does all the mixing. Know the headline is roughly nine times cheaper at 3x3.
Derive the ratio 1/C_out + 1/(k*k) from the two parameter counts on the spot, and be able to plug in real channel numbers and get the right figure without hesitating.
Say which stage the cost actually lives in, what expressivity the factorisation trades away, and how you would spend the freed budget — more width, more depth, or higher input resolution.
Own the decision of where separability belongs in a backbone at all: which stages tolerate it, what accuracy you are willing to trade, and why a headline FLOP reduction is an upper bound on the benefit you will ship.
## The dense baseline A convolution with kernel `k x k`, `C_in` input channels and `C_out` output channels has one `k x k` plane of weights for every (input channel, output channel) pair: ``` params_dense = k * k * C_in * C_out macs_dense = H_out * W_out * params_dense ``` Every filter does two jobs at once: it aggregates a `k x k` spatial neighbourhood *and* it combines all `C_in` channels. Depthwise-separable convolution is the hypothesis that those two jobs can be done sequentially by two cheaper layers with little loss. ## The factorisation **Stage 1, depthwise.** One `k x k` filter per input channel, each applied to its own channel only. This is a grouped convolution with the group count equal to `C_in`. It produces `C_in` output channels (with a depth multiplier of 1) and does no channel mixing at all: ``` params_depthwise = k * k * C_in ``` **Stage 2, pointwise.** A dense 1x1 convolution mapping `C_in` channels to `C_out`, which does no spatial aggregation and all of the channel mixing: ``` params_pointwise = C_in * C_out ``` A normalization and a nonlinearity are normally applied after each stage, so the block is not a linear collapse of the two. ## The ratio ``` params_sep / params_dense = (k*k*C_in + C_in*C_out) / (k*k*C_in*C_out) = 1/C_out + 1/(k*k) ``` Two terms, each with a clean reading. `1/(k*k)` is what you save by not repeating the spatial kernel for every output channel. `1/C_out` is what you save by not repeating the channel mixing at every spatial offset. With a 3x3 kernel and any width beyond a few dozen channels, `1/C_out` is small and the ratio sits just above `1/9`. Worked concretely for `k = 3`, `C_in = 128`, `C_out = 256`: ``` dense = 9 * 128 * 256 = 294,912 depthwise = 9 * 128 = 1,152 pointwise = 128 * 256 = 32,768 separable = 33,920 ratio = 33,920 / 294,912 = 0.115 -> 8.69x fewer ``` Because both stages operate at the same output resolution, the multiply-accumulate ratio is identical: multiply either parameter count by `H_out * W_out`. The kernel size drives the answer. At `k = 1` the factorisation saves nothing (there is no spatial kernel to factor out). At `k = 5` the ratio is `1/25 + 1/C_out`, roughly 23x for a wide layer. Larger kernels make separability more attractive, which is why compact designs pair depthwise stages with kernels they would never afford densely. ## Where the parameters end up Note the split in the worked example: the depthwise stage holds 1,152 weights and the pointwise holds 32,768 — over 96 percent of the block. The stage people talk about is the cheap one; the stage that actually costs is the 1x1. That matters when reasoning about what to shrink next, and it explains why designs that further attack cost tend to attack the pointwise stage. ## What you give up The factorisation is a structural constraint, not a free lunch. A dense filter can learn a spatial pattern that exists only in a particular *combination* of input channels. A separable block cannot: the spatial filtering happens per channel, before any mixing, so the spatial pattern each channel is scanned for is fixed independently of how channels will later be combined. In practice the loss is modest at typical widths, and the saved budget usually buys back more than it costs — more channels, more layers, or higher input resolution — but on small or heavily structured problems the dense layer can genuinely win. A second consequence is optimisation behaviour: a depthwise filter receives gradient from a single channel's activations, so a dead or badly scaled channel has nowhere to hide. Depthwise stages are more sensitive to weight decay and to aggressive regularisation than dense ones, because there are so few weights per channel that any shrinkage is felt directly. ## Stride and shape When the block downsamples, the stride is applied in the depthwise stage and the pointwise stays at stride 1. The usual output-size arithmetic still applies to the depthwise stage: `out = floor((n + 2p - k)/s) + 1`. The pointwise stage never changes spatial size; it only sets the output depth. ## The one-line answer Separating a convolution into a per-channel spatial filter followed by a 1x1 channel mixer replaces a `k*k*C_in*C_out` cost with `k*k*C_in + C_in*C_out`, a ratio of `1/C_out + 1/(k*k)`, which for a 3x3 layer of ordinary width is about one ninth.
- Which of the two stages holds most of the block's parameters, and why does that matter?The pointwise 1x1 does. In a 3x3 block with 128 input and 256 output channels it holds 32,768 of the 33,920 weights, over 96 percent, because its cost is `C_in * C_out` while the depthwise stage is only `k*k*C_in`. It matters because any further cost reduction has to attack the 1x1; shaving the depthwise stage buys essentially nothing.
- For which kernel size does the separable factorisation stop being worth it?At `k = 1` it saves nothing, since there is no spatial kernel to factor out and the depthwise stage degenerates to a per-channel scalar. The saving grows with `k*k`: about 9x at 3x3 and about 23x at 5x5 for a wide layer. That is why compact designs can afford larger depthwise kernels that would be unaffordable densely.
- What representational capacity does the factorisation give up?A dense filter can learn a spatial pattern that only exists in a specific combination of input channels. A separable block filters each channel spatially *before* any mixing, so the per-channel spatial patterns are chosen independently of how the channels will later be combined. The lost expressivity is usually repaid by spending the saved budget on width, depth, or resolution.
Instead of hiring one specialist for every (input, output) pair, you hire one blur specialist per input channel and then hold a single meeting where all channels are combined — two cheap passes instead of one enormous one.
saying these in an interview costs you the question
- Claims the saving is a factor of C_out rather than about k squared
- Says the depthwise stage holds most of the parameters
- Forgets that a pointwise stage is needed to mix channels at all
- Treats the factorisation as lossless in capacity
- Thinks the depthwise stage changes the channel count