How does EfficientNet's compound scaling split a compute budget across depth, width and resolution?
answer
- one knob, three dimensions
- cost is linear in one, quadratic in two
- ratios fixed once on the small base
- alpha times beta-squared times gamma-squared near two
- each unit of phi doubles the FLOPs
basics
~10 sCompound scaling raises depth, width and input resolution together as fixed powers of one coefficient, so extra compute is spread over all three dimensions instead of poured into one, which saturates on its own.
solid answer
~50 sCompound scaling fixes three exponents once and then moves a single knob. Depth is multiplied by `alpha^phi`, width by `beta^phi` and input resolution by `gamma^phi`, where `phi` is the one number you change to grow the model. Convolution cost is roughly linear in depth but quadratic in both width and resolution, so the exponents are constrained by `alpha * beta^2 * gamma^2 ≈ 2`; each unit increase of `phi` then costs about 2x the FLOPs. The ratios `alpha`, `beta`, `gamma` are found once by a small grid search on the small base network under that constraint, and reused for every larger size instead of being re-searched. The motivation is empirical: pushing any single dimension alone flattens out quickly, because a higher-resolution input needs more layers to grow the receptive field and more channels to represent the finer patterns it now contains, and neither arrives if you only scale resolution.
go deeper
Be ready to name the three dimensions a convolutional network can be scaled along and say that they are grown together by one coefficient rather than one at a time.
Explain the cost arithmetic: linear in depth, quadratic in width and in resolution, and how that produces the constraint that makes one step of the coefficient cost about twice the FLOPs.
Show you know where the rule leaks in production: activation memory rather than parameter count is the wall resolution scaling hits, and a FLOP-balanced family is not automatically latency-balanced on your device.
Own the framing that a scaling rule multiplies whatever the base design is worth. Argue when a team should invest in a better base block versus simply moving along a published scaling family.
## The problem compound scaling solves You have a working convolutional network and a bigger compute budget. There are three obvious ways to spend it: - **depth** — more layers, - **width** — more channels per layer, - **resolution** — a larger input image. The usual practice was to pick one and push it. Compound scaling is the claim that this is the wrong shape of decision, and that the three dimensions should be scaled together in a fixed proportion. ## How the three dimensions cost compute A convolution layer's multiply-accumulate count is proportional to `channels_in * channels_out * kernel_area * output_positions`. Read the three dimensions off that expression: - Adding layers multiplies total cost **linearly** — twice the depth, twice the FLOPs. - Scaling width by a factor `b` multiplies *both* `channels_in` and `channels_out`, so cost grows as **`b^2`**. - Scaling input resolution by `g` multiplies the output height and width, so `output_positions` and cost grow as **`g^2`**. So the total compute multiplier of a scaled model is roughly `d * w^2 * r^2` relative to the base. Two consequences fall straight out. First, doubling width is four times the compute, not twice — a very common misstatement. Second, resolution scaling costs the same quadratic factor as width but adds **no parameters at all** to the convolution kernels; what it inflates is activations, hence peak memory during training and inference. ## The compound rule Compound scaling defines the scaled model by a single coefficient `phi`: ``` depth = alpha^phi width = beta^phi resolution = gamma^phi subject to alpha * beta^2 * gamma^2 ≈ 2, with alpha, beta, gamma >= 1 ``` Because of the quadratic terms in the constraint, the total FLOP multiplier is approximately `2^phi`. Setting `phi = 1` roughly doubles compute, `phi = 3` roughly multiplies it by eight, and so on. That is the whole point of the constraint: it turns "give me twice the budget" into a well-defined, reproducible change to all three dimensions at once. The procedure has two stages: 1. **Fix the ratios once.** With `phi = 1`, grid-search `alpha`, `beta`, `gamma` on the *small base* network under the constraint. This is cheap because the base model is small. 2. **Move `phi` only.** Every larger family member reuses the same ratios. You do not re-search the coefficients per size, and that is what makes the family a family rather than a set of unrelated models. ## Why joint scaling beats one dimension The dimensions are coupled, not independent budgets. - Raise **resolution** alone and the extra detail is available at the input but nothing downstream can use it: the receptive field of the unchanged stack now covers a smaller *fraction* of the image, and the channel count is unchanged, so the finer textures have nowhere to be represented. Accuracy climbs for a while, then flattens — while activation memory keeps climbing quadratically. On something like a satellite land-cover classifier, where the informative signal really is fine-grained, this is a seductive trap: doubling the tile resolution feels like it should always help, and the memory bill arrives long after the accuracy curve went flat. - Raise **depth** alone and you eventually hit optimization and diminishing-feature-reuse limits; the marginal layer sees the same coarse input. - Raise **width** alone on a shallow stack and you get many channels of relatively low-level features. Each dimension saturates on its own, and each saturation is relieved by moving the other two. Scaling them in proportion keeps the network balanced, so the accuracy-per-FLOP curve of the family stays steeper for longer than any single-dimension curve. ## What it does not do Compound scaling is a **scaling rule, not an architecture**. It says how to grow a base network; the base network's quality still dominates the result. Apply the same rule to a poor baseline and you get an expensive poor model. The ratios are also tied to the base design and the data: they were found by grid search on one base network, and there is no guarantee they are optimal for a different block design or a much larger `phi` than the family was validated over. Two practical cautions. Activation memory, not parameter count, is usually the constraint you hit first when resolution grows, so "it only added 5% parameters" is not the reassurance it sounds like. And the scaling rule optimizes FLOPs by construction — if your real constraint is measured wall-clock latency on a specific device, a FLOP-balanced family is a starting point, not an answer.
- Why does scaling only the input resolution stop paying off?The extra input detail has nowhere to go. With depth unchanged, the receptive field covers a smaller fraction of the image, so the network never integrates the finer structure; with width unchanged, there are no extra channels to represent it. Accuracy flattens while cost and peak activation memory keep growing quadratically in the resolution factor, which is the worst possible trade to be making.
- Does doubling the channel width double a convolution layer's FLOPs?No, it roughly quadruples them. A convolution's cost is proportional to input channels times output channels times spatial positions, and widening scales both channel counts, giving a squared factor. Depth is the only one of the three dimensions whose cost is linear, which is exactly why the constraint on the scaling coefficients carries squares on the width and resolution terms.
- Do you re-run the coefficient grid search for each larger model in the family?No. The grid search is done once on the small base network at the smallest compute step, because searching directly at large scale is prohibitively expensive. After that only the compound coefficient moves. That reuse is an approximation and part of why the family's ratios are not claimed optimal for arbitrary base architectures or arbitrarily large scale-ups.
Think of tuning a stereo: cranking only the bass gets loud and muddy long before it gets better. Compound scaling turns up bass, mid and treble together on one master dial, in ratios fixed once by ear.
saying these in an interview costs you the question
- Says just stack more layers until accuracy stops improving
- Claims doubling width doubles compute rather than quadrupling it
- Thinks higher input resolution adds convolution kernel parameters
- Re-searches the scaling ratios separately at every model size
- Treats compound scaling as an architecture rather than a scaling rule