You swapped 3x3 convolutions for depthwise-separable blocks, cut FLOPs 8x, but latency only halved — why?
answer
- arithmetic removed, data movement unchanged
- each element used only k-squared times
- the cheap stage was never the slow one
- one layer became two full passes
- the 1x1 still holds nearly all the math
basics
~10 sA depthwise stage removes arithmetic but not data movement: it streams the full activation tensor while doing only k-squared multiply-accumulates per element. Runtime becomes memory-bound, and one layer becoming two adds further passes.
solid answer
~40 sFLOPs are an upper bound on speed-up, not a prediction of it. A dense 3x3 layer loads each input element once and reuses it across all `C_out` filters, so it does a lot of arithmetic per byte moved and keeps the multiply units busy. The depthwise stage reuses nothing across channels: each element feeds only `k*k` multiply-accumulates, so the stage streams the whole tensor in and out for very little math and becomes memory-bandwidth-bound. Meanwhile the pointwise 1x1 holds over ninety percent of the block's remaining arithmetic, so your FLOP saving mostly deleted work that was never the bottleneck in time. One layer also became two, each with its own normalization and activation pass over the full tensor. Benchmark latency on the target device and treat the FLOP ratio as a best case.
go deeper
Take away the headline: fewer multiply-accumulates does not automatically mean a faster model, because moving activations through memory also takes time.
Be able to explain the reuse difference — a dense filter uses each loaded input element across every output channel, a depthwise filter uses it only k-squared times — and say which stage of the block still holds the arithmetic.
Show the diagnosis: profile the two stages separately, attribute wall-clock against multiply-accumulates, and act on the stage that holds the time rather than the one with the impressive ratio.
Own the standard your team designs against. Decide whether compact-model targets are stated in FLOPs or in measured latency on the deployment device, and be ready to defend keeping a dense layer that a FLOP-driven review would have factorised.
## The mismatch FLOPs (or multiply-accumulates) count arithmetic. Latency counts the time a device takes to produce an output, which is set by whichever resource saturates first — arithmetic units, memory bandwidth, or fixed per-layer overhead. A design change that removes arithmetic only shortens the run if arithmetic was the binding constraint. For depthwise convolution it usually is not. ## Why a dense convolution is arithmetic-heavy In a dense `k x k` layer, one input element at position `(c, h, w)` is read once and then participates in `k * k * C_out` multiply-accumulates — every output channel's filter wants it, at every offset that covers it. Arithmetic done per byte of activation loaded is therefore large, growing with `C_out`. That is a workload the multiply units can be kept busy on, and it is why dense convolutions can be expressed as large dense matrix products and executed efficiently. ## Why a depthwise convolution is not In the depthwise stage each input element belongs to exactly one channel and is consumed by exactly one filter, so it participates in only `k * k` multiply-accumulates — nine for a 3x3, independent of width. The layer still reads the entire input tensor and writes an output tensor of the same size. Arithmetic per byte moved has collapsed by a factor of `C_out`, so the stage spends most of its time waiting for memory rather than computing. It is **memory-bandwidth-bound**: shaving its FLOPs further would change nothing, because its FLOPs were not what made it slow. There is a second, related effect. A depthwise layer decomposes into `C_in` tiny independent convolutions, each over a single channel plane, with almost no data reuse to amortise. That maps poorly onto machinery built for large dense matrix products, and on a wide parallel device it can also leave units idle when each per-channel problem is small. ## Where the FLOPs actually were Run the numbers on a 3x3 block with 128 input and 256 output channels. Dense: 294,912 weights, so 294,912 multiply-accumulates per output position. Separable: 1,152 in the depthwise stage and 32,768 in the pointwise, 33,920 total. The 8.7x FLOP reduction is real — but 96.6 percent of the remaining arithmetic sits in the pointwise 1x1, and the stage you made "free" was already contributing only about 0.4 percent of the dense block's math. In other words, your headline FLOP saving came from deleting arithmetic, while your latency is now set by a stage whose arithmetic you did not touch plus a stage that is bound by bandwidth. ## The costs the FLOP count never showed - **Two layers instead of one.** Each stage typically carries its own normalization and nonlinearity, so the full activation tensor is read and written several more times than before. Those passes do almost no arithmetic and cost pure bandwidth. - **Fixed per-layer overhead.** Every layer launch has a floor cost. Doubling the layer count in every block multiplies that floor, which is most visible at small batch sizes and small spatial sizes — exactly the on-device inference regime compact models target. - **Awkward shapes.** Channel counts that are not friendly multiples for the device's tiling can leave the multiply units partly idle, and a narrow depthwise stage has no width to hide that behind. ## How to reason about it going forward 1. **Measure, do not extrapolate.** Time the actual model on the actual device at the actual batch size. A FLOP ratio is a ceiling on the achievable speed-up. 2. **Attribute the time per stage.** If the depthwise stages dominate wall-clock while contributing nearly no FLOPs, you have confirmed the bandwidth story rather than assumed it. 3. **Attack the stage that holds the time.** If the pointwise 1x1 dominates the arithmetic, that is where further cost work belongs; the depthwise stage has nothing left to give. 4. **Reduce passes over the tensor.** Fusing the normalization and activation into the convolution stage removes whole reads and writes of the activation tensor, and that saves time the FLOP count cannot see. 5. **Consider fewer, wider layers.** Sometimes a dense layer that keeps the multiply units busy finishes sooner than a factorised block with a third of the arithmetic — a genuinely counter-intuitive result that only a benchmark reveals. ## What a strong answer sounds like The candidate does not say "FLOPs are a bad metric" and stop. They name the mechanism: the depthwise stage moves as much data as before for a factor of `C_out` less arithmetic, so it is bandwidth-bound; the pointwise stage still holds the arithmetic; and splitting one layer into two adds passes and launch overhead. Then they say what they would measure to confirm it. A weak answer blames the framework, or claims the FLOP count must have been computed wrong.
- Which stage of the separable block would you profile first, and what would confirm your hypothesis?Time the two stages separately. If the depthwise stage takes a substantial share of the block's wall-clock while contributing well under one percent of its multiply-accumulates, the bandwidth story is confirmed. If instead the pointwise 1x1 dominates both time and arithmetic, the block is behaving as the FLOP count predicts and further gains have to come from narrowing it.
- What change would reduce latency without changing the FLOP count at all?Fusing the normalization and activation into the convolution stage. Those operations do almost no arithmetic but each reads and writes the entire activation tensor, so folding them in removes whole passes over memory. Reducing the number of separate layers per block has the same effect, cutting fixed per-layer launch overhead that never appears in a FLOP count.
- Could a dense convolution ever be faster than the separable block that replaces it?Yes, and it happens. A dense layer keeps the multiply units saturated and runs as one large well-shaped matrix product, while the separable block splits into a bandwidth-bound stage plus a second layer's overhead. At small spatial sizes, small batch sizes, or unfriendly channel counts the dense layer can finish sooner despite doing several times more arithmetic.
FLOPs are the number of ingredients in a recipe; latency is how long the kitchen takes. A dish with a ninth of the ingredients can still be slow if every single step means another trip to the pantry.
saying these in an interview costs you the question
- Assumes a FLOP reduction converts one-for-one into speed-up
- Blames the measurement instead of the memory traffic
- Thinks the depthwise stage holds most of the block's arithmetic
- Ignores that one layer became two full tensor passes
- Concludes FLOPs are useless rather than an upper bound