Given a fixed parameter budget for a feedforward net, how do you decide depth versus width?
answer
- ask which constraint binds first
- depth is sequential, width is parallel
- dense width costs grow quadratically
- matched budgets, one axis at a time
- look for the knee, not the maximum
basics
~10 sLet the binding constraint decide. Depth buys parameter efficiency but adds sequential latency and harder optimization; width parallelizes well but costs full fan-in and fan-out per unit. Settle it with small matched-budget runs.
solid answer
~50 sStart by naming the constraint that actually binds — accuracy at fixed memory, latency per request, or training cost — because the axes trade differently against each of them. Depth is the parameter-efficient direction: composition compounds expressivity, so at a matched budget a deep trunk usually reaches a richer hypothesis class. It is also the expensive direction operationally, since a forward pass must traverse layers in order, so latency grows roughly linearly with depth while a wide layer's arithmetic spreads across parallel hardware. Width is more forgiving to optimize and avoids bottlenecking information, but each added unit costs its full fan-in plus fan-out. My practical rule is to move one axis at a time on a reduced-scale proxy at matched parameter counts, watch where returns saturate, and refuse to import a depth-to-width ratio from a different domain or input size.
go deeper
Recall that depth and width are separate design choices, and that a network with the same number of parameters can be arranged as few wide layers or many narrow ones.
Be ready to state the mechanics: depth compounds features and adds sequential steps, dense width costs grow with the square of the layer size, and both must be compared at a matched budget.
Show a real procedure — reduced-scale proxy runs, one axis at a time, hunting the knee of the curve — and connect the choice to the serving constraint you were actually accountable for.
Own the framing that the architecture axis is chosen by the binding business constraint, and be willing to say when the search itself is not worth the compute compared with investing in data quality.
## Frame the question before answering it "Deeper or wider?" has no context-free answer, and a strong candidate says so first. The budget is rarely just parameters — it is usually parameters *and* a latency ceiling, or a training-compute cap, or an on-device memory limit. Depth and width trade differently against each of those, so the first move is to name which constraint actually binds. ## What each axis buys **Depth buys parameter efficiency.** Layer k operates on the features layer k-1 produced, so expressible structure compounds through composition rather than accumulating one unit at a time. At a matched budget a deep trunk generally reaches a function class a shallow one cannot afford. Arithmetically, depth reuses a modest fan-in repeatedly; width is charged the full incoming and outgoing connections for every unit added. **Width buys optimization comfort and hardware utilisation.** Wide layers give the optimizer many redundant directions, which empirically makes training more forgiving. A wide matrix multiply is one large, highly parallel operation, so on parallel accelerators a wide layer is often close to free in wall-clock terms until it exhausts memory bandwidth. Wide layers also avoid information bottlenecks: a layer narrower than the signal it must carry destroys structure no later layer can recover. ## What each axis costs **Depth costs sequential latency and trainability.** A forward pass must traverse layers in order, so per-request latency scales roughly with depth no matter how much parallel hardware you own. Deep stacks are also harder to optimize, and the machinery that makes very deep stacks trainable is itself a design commitment you take on. Returns saturate: the tenth layer buys far less than the third. **Width costs quadratically in parameters.** Between two dense layers of width w the weight count is about `w*w`, so doubling width quadruples that block and halving it cuts the block to a quarter. Memory and parameter budget go first; wall-clock time often improves sublinearly on the way down, because a narrow layer under-utilises parallel hardware. ## The decision procedure 1. **Fix the constraint.** Latency-bound serving, memory-bound device, or accuracy-bound research each point at a different axis. 2. **Compare at matched budgets, not matched hyperparameters.** A 10-layer 256-wide trunk and a shallow net with the same parameter count are the comparison; anything else confounds capacity with configuration. 3. **Move one axis at a time.** Change depth with width fixed, then width with depth fixed. Joint sweeps at full scale burn compute and rarely tell you which axis produced the change. 4. **Use a reduced-scale proxy.** Shorter schedules and a subsample give the *shape* of the depth and width curves cheaply. Take the shape seriously and the absolute numbers not at all. 5. **Find where returns saturate.** Both curves flatten. The interesting information is the knee, not the maximum. 6. **Do not import a ratio.** A depth-to-width ratio tuned for another input dimensionality, dataset size or hardware target is not evidence about yours. ## The judgment part Two failure modes recur, and naming them is what separates a principal answer. The first is **cargo-culting a shape**: copying a published trunk's proportions into a problem with a fraction of the data or a tenth of the input dimension, then blaming the data when it underperforms. The second is **optimizing the wrong axis**: spending a quarter's compute pushing accuracy through depth when the product constraint was tail latency all along, which depth is the worst axis to grow. There is also an organisational call. Depth and width searches are cheap to parallelize and easy to over-run. Decide up front how much of the budget the architecture search deserves relative to data work, because on most real problems better labels move the metric more than a re-proportioned trunk does. ## A defensible default Absent a hard constraint, start at a depth known to be trainable for the problem family, set width from the input dimensionality so no layer is an obvious bottleneck, then grow the axis whose curve has not yet flattened — measuring both accuracy and the latency the product actually pays for.
- Your latency budget is halved. Do you cut depth or width first?Depth, usually. A forward pass traverses layers in order, so removing layers cuts wall-clock time close to linearly regardless of hardware. Halving width quarters the parameters in a dense block but often improves latency sublinearly, because narrow layers under-utilise parallel hardware. Cut width first only when the constraint is memory rather than time.
- How do you compare depth and width settings without a full training run per configuration?Run a reduced-scale proxy: a subsample of the data and a shortened schedule, everything else held fixed, with configurations matched on parameter count. You are reading the shape of the curve and where it flattens, not the absolute metric. Confirm only the two or three survivors at full scale.
- When is going wider clearly the right call over going deeper?When you are latency-bound on parallel hardware, when optimization is already fragile and you want a more forgiving objective, or when a layer would otherwise be narrower than the information it must carry. A bottleneck layer destroys structure permanently, and no amount of depth above it recovers what was lost.
saying these in an interview costs you the question
- Answers deeper is always better with no constraint named
- Assumes halving width halves inference latency
- Copies a depth-to-width ratio from an unrelated domain
- Compares configurations without matching parameter budgets
- Ignores that per-request latency scales with layer count