skip to content

Why can a network with 5x fewer FLOPs still be slower than its rival on a mobile CPU?

level: middleimportance: must knowfreq 70%

answer

  1. A proxy, not a stopwatch
  2. Arithmetic is not the only cost
  3. Some layers move data per multiply
  4. Utilization and per-layer overhead differ
  5. MACs versus FLOPs is 2x

basics

~20 s

FLOPs counts arithmetic only. Latency also pays for moving data, for per-layer fixed costs, and for how well each operation uses the device's parallel units, so a low-FLOP network built from poorly-supported operations can lose on the stopwatch.

solid answer

~50 s

FLOPs is an analytic count of the arithmetic one forward pass performs; latency is what a stopwatch on the target device reports. The two agree only when every operation reaches similar hardware utilization. A network built from cheap-per-multiply operations — depthwise convolutions, many thin layers, lots of normalization, activation, concatenation and resizing — touches a great deal of data per unit of arithmetic and leaves the parallel units partly idle, so its measured time can exceed a denser rival's. At batch size one the gap is worst, because fixed per-layer costs dominate and there is little work to spread across cores. FLOPs is also reported inconsistently: many papers count multiply-accumulates and label them FLOPs, a factor-of-two difference. Treat FLOPs as a design-time proxy you can compute before training anything, and treat measured latency on the real device as the acceptance criterion.

go deeper

for a junior

Be ready to say what a FLOP count is — the arithmetic one forward pass performs, computed on paper — and that it is an estimate of cost, not a measurement of time.

for a middle

Explain the mechanics of the gap: per-operation overhead, hardware utilization, and the data each layer touches per unit of arithmetic. Know that a multiply-accumulate is two FLOPs and that papers disagree about this.

for a senior

Show the operating habit: shortlist on FLOPs, decide on measured latency on the real device in the real batching regime, and refuse to accept a speed claim that does not name its hardware.

for a principal

Own the standard your organisation reports against. Decide whether efficiency targets are written in FLOPs, milliseconds or energy, and be able to defend why a paper-comparable number and a contractual number are different artifacts.

## What a FLOP count actually is A FLOP count is an **analytic** number: you walk the network layer by layer, multiply out the arithmetic each layer performs on one input, and add it up. For a convolution with `k x k` kernels, `c_in` input channels, `c_out` output channels and an `h x w` output map, the multiply-accumulate count is `k * k * c_in * c_out * h * w`. Nothing is measured; no device is involved. That is exactly the property that makes FLOPs useful — you can compare two architectures on paper before you own a trained model — and exactly the property that makes it an unreliable predictor of time. **A naming trap first.** One multiply-accumulate (MAC) is a multiply plus an add, i.e. two floating-point operations. Many papers count MACs and print the word "FLOPs", so two published numbers can differ by 2x purely by convention. Conventions also differ on whether biases, normalization, activation functions and pooling are counted at all — they are usually dropped as "negligible", which is precisely the arithmetic that is cheap but not fast. Before comparing two published numbers, check what each one counted, or recount both yourself. ## Why the ordering flips Latency is arithmetic **plus** everything else the device must do: - **Data movement.** Every operation reads inputs and weights and writes outputs. Some operations do a lot of arithmetic per byte touched (a big dense matrix multiply); others do very little (an elementwise activation, a normalization, a concatenation, a resize). A network that shaves FLOPs by replacing dense work with many light-touch operations can spend most of its wall-clock time shuttling tensors rather than computing. - **Utilization.** Hardware reaches peak throughput only when a single operation offers enough independent work with the right shape. Thin layers, small channel counts, grouped and depthwise convolutions, and odd tensor shapes all leave lanes idle. A depthwise convolution may cost a tiny fraction of the FLOPs of the dense convolution it replaces and still take a comparable amount of time on some devices, because its effective utilization is far lower. - **Fixed per-operation cost.** Each operation has overhead that does not scale with its FLOPs. A network of 200 tiny operations pays that 200 times; a network of 40 large ones pays it 40 times. This is why the mismatch is worst at batch size one, where there is the least real work to amortize the overhead against. - **Implementation maturity.** A well-worn operation with a highly tuned implementation on that device beats an exotic operation with a generic fallback path, at equal FLOPs. - **Non-arithmetic work.** Layout changes, padding, type conversions and data-dependent control flow cost time and count zero FLOPs. Because utilization and implementation quality are properties of the **device**, the ranking is device-specific. Two models can swap places between a phone CPU, a phone's neural accelerator and a server GPU. "Model A is faster" is only a claim about the hardware you measured on. ## The worked case Candidate A has one-fifth the FLOPs of candidate B and is measurably slower on the target mobile CPU. What the FLOP count failed to price in: A's savings came from operations with low arithmetic density and low utilization, A has many more separate layers so pays overhead many more times, its normalization and activation layers contribute almost nothing to the FLOP total but a real share of the runtime, and it is being run at batch size one, the regime that most punishes small operations. B is "expensive" in a currency the device happens to spend efficiently. ## So why report FLOPs at all Because it is device-independent, computable at design time before any training run, and roughly monotone **within** one architecture family — doubling the width of the same network really does make it slower. It is a legitimate design-time proxy and a legitimate way to say "this model is in a different cost class". It is not an acceptance criterion. If the deployment constraint is stated in milliseconds, the number that closes the ticket is a measured millisecond figure on the target hardware, in the batching regime the product actually uses. ## What a strong answer sounds like Name the metrics separately — FLOPs is arithmetic, latency is time — give two concrete reasons they diverge (utilization and per-operation overhead, or data movement and operation count), note that the ranking is hardware-specific, mention the MAC-versus-FLOP convention trap, and finish with the practical rule: use FLOPs to shortlist candidates, use measured on-device latency to choose between them.

  • Papers report FLOPs and MACs almost interchangeably — why does that matter when you compare two published models?
    One multiply-accumulate is two floating-point operations, so a paper that counts MACs and writes "FLOPs" reports half the number another paper would. Conventions also differ on whether biases, normalization and activation are counted. Two numbers from different papers can therefore differ by 2x for no architectural reason. Recount both under one convention, or compare only within a single source.
  • If FLOPs mispredicts latency this badly, why is it still the headline efficiency number in papers?
    It is device-independent and reproducible, so it survives the fact that reviewers do not own the author's hardware. It can be computed at design time before a model exists. And within one architecture family it is monotone — scaling width or resolution up really does cost more time. It is a fair way to state a cost class, just not a fair way to promise a millisecond budget.
  • Does the FLOP-to-latency mismatch shrink if you move from a phone to a server accelerator?
    It changes rather than shrinks. Larger batches raise utilization and amortize per-operation overhead, which brings dense models closer to their FLOP-implied cost, but a server accelerator has its own set of well-supported and poorly-supported operations. A model tuned to be fast on a phone can be the slower option on a server, and the reverse. The ranking is a property of the hardware you measured, not of the model alone.

FLOPs is the distance on the map; latency is the drive time. Two routes of equal length differ wildly once you price in traffic lights, road surface and how many turns you make.

saying these in an interview costs you the question

  • Says FLOPs determines runtime directly
  • Assumes fewer FLOPs always means faster inference
  • Treats normalization and activation layers as free
  • Compares FLOP numbers across papers without checking the convention
  • Reports a speed claim without naming the device or batch size
  • Never proposes measuring on the target hardware

context