skip to content

A 40-1024-1024-250 MLP holds 1.35M parameters; which single width do you cut to halve it?

level: middleimportance: should knowfreq 41%

answer

  1. find the biggest matrix first
  2. products of adjacent widths, not sums
  3. the two hidden layers touch each other
  4. one width appears in two products
  5. 1024 x 1024 is 78 percent of the model

basics

~10 s

Cut the second hidden width. The 1024-to-1024 matrix holds about 1.05M of the 1.35M parameters, so taking that layer to 512 units drops the total to roughly 695,000. The 40-input layer holds almost nothing.

solid answer

~50 s

Cut the second hidden width. Dense counts are products of adjacent widths, so the single `1024 -> 1024` matrix holds `1024 * 1024 + 1024 = 1,049,600` of the 1.35M total — about 78 percent — while the `40 -> 1024` input layer is only `41,984` (3 percent) and the `1024 -> 250` output layer `256,250`. Taking the second hidden layer to 512 units shrinks both products it appears in, leaving `41,984 + 524,800 + 128,250 = 695,034`, almost exactly half. Trimming the input layer cannot help: even deleting it saves 3 percent. The output width is fixed by the 250 classes. The judgment half is what the cut costs — 512 units feeding a 250-way decision is still roughly twice the number of classes, so this is the cheapest place to buy the budget.

go deeper

for a junior

Know that a dense layer's size is the product of the two widths it joins, so the biggest matrix is the one with wide numbers on both sides. Find it before proposing any cut.

for a middle

Be able to run the arithmetic both ways: forwards from widths to a total, and backwards from a target total to the one width worth changing. Explain why a width appearing in two products shrinks both.

for a senior

Pair the arithmetic with the cost. Say what the narrowed layer still has to support — here, a 250-way decision — and what measurement you would run before promising the smaller model performs the same.

for a principal

Frame budgets so a team does not chase percentages in the wrong layers. Decide up front which widths are free variables and which are fixed by the task, and make that the shared vocabulary for size discussions.

## Parameter counts are products, so they concentrate A fully-connected layer between widths `a` and `b` holds `a * b + b` parameters. Because the dominant term is a **product** of two widths, the totals of an MLP are almost never spread evenly across its layers: the single largest adjacent product usually owns most of the model. That single fact is what makes "which width do I cut?" answerable in about fifteen seconds. ## The worked example A speaker-identification MLP takes 40 audio coefficients per frame and classifies 250 speakers, with two 1024-unit hidden layers: widths `[40, 1024, 1024, 250]`. - `40 -> 1024`: `40 * 1024 + 1024 = 41,984` — about **3 percent** - `1024 -> 1024`: `1024 * 1024 + 1024 = 1,049,600` — about **78 percent** - `1024 -> 250`: `1024 * 250 + 250 = 256,250` — about **19 percent** Total: `1,347,834`, roughly 1.35M. One matrix holds more than three quarters of the model, and it is the one whose *both* sides are wide. ## The cut Take the second hidden layer from 1024 units to 512. That width appears in two products, so both shrink: - `40 -> 1024`: unchanged at `41,984` - `1024 -> 512`: `1024 * 512 + 512 = 524,800` - `512 -> 250`: `512 * 250 + 250 = 128,250` New total: `695,034` — about 52 percent of the original. One width, one edit, budget met. ## Why the other candidates fail The input layer is capped by arithmetic: even deleting it entirely saves 41,984, or 3 percent. Squeezing a layer whose fan-in is only 40 is wasted effort no matter how aggressive you are — a very common junior instinct, because that layer *looks* like the front of the network and therefore important. The output layer is capped by the task: 250 classes is 250 output units, not a knob. Its 256,250 parameters shrink only as a side effect of narrowing the hidden layer that feeds it, which is exactly what the chosen cut does. The hidden widths are the only genuinely free variables, and among them the one that sits between two large numbers is worth the most. ## Overshooting: halving everything Halving **both** hidden widths to 512 gives `21,012 + 262,656 + 128,250 = 411,918` — under a third of the original, not a half. That is the quadratic scaling at work: multiplying every hidden width by a factor `c` scales each hidden-to-hidden matrix by about `c * c`. Uniform scaling is a blunt instrument; when a budget asks for a factor of two, a uniform halving usually buys far more than asked and gives up capacity you did not need to give up. The same arithmetic runs the other way when you are adding capacity. Bolting a third 1024-unit hidden layer onto this network adds another `1,049,600` — depth is expensive precisely *because* the network is wide. Adding a layer to a narrow network is cheap; adding one to a wide network nearly doubles it. ## What the cut costs The counting is the easy half; the judgment is knowing whether the model can afford it. A 512-unit representation feeding a 250-way classifier is still roughly twice the number of classes, so the layer is not being squeezed below the width of the decision it has to support. Had the budget demanded 128 units in front of 250 classes, the layer would have become a hard bottleneck and the honest answer would be "this budget costs accuracy; here is the measurement I would run before promising it." State that split explicitly in an interview: here is the arithmetically obvious cut, here is why the alternatives cannot deliver, and here is the empirical check I would still run. ## The wider pattern The same reasoning explains a shape that surprises people in classic image classifiers: the fully-connected head can hold more parameters than every layer beneath it. Convolutional layers reuse one small kernel at every spatial position, so their parameter count does not grow with image size, while a dense pair like `4096 -> 4096` stores an independent weight for every input-output pair — `16,777,216` weights in that one layer. Parameter mass follows dense, wide-to-wide connections, not depth and not the part of the network that does the most visible work. ## The habit to build Given any width list, write down the adjacent products, sort them, and look at the top one. That single number tells you what the model *is*, budget-wise, and it turns a vague "make it smaller" request into a specific edit you can defend.

  • What if you halve every hidden width instead of just one?
    You overshoot badly. Each hidden-to-hidden matrix scales with the product of both widths, so halving both quarters it: the total falls to `21,012 + 262,656 + 128,250 = 411,918`, under a third of the original. Uniform scaling is a blunt instrument — when a budget asks for a factor of two, it buys a factor of three and gives up capacity you did not have to give up.
  • Why does the fully-connected head of a classic image classifier often hold more parameters than every layer beneath it?
    Because convolutional layers reuse one small kernel at every spatial position, so their parameter count does not grow with image size, while a dense layer stores an independent weight per input-output pair. A single `4096 -> 4096` pair holds 16,777,216 weights. Parameter mass follows wide-to-wide dense connections, not depth and not where the visible work happens.
  • At a fixed budget, is it cheaper to add a layer or to widen an existing one?
    It depends entirely on the widths involved. Adding a layer to a narrow network is cheap; adding a third 1024-unit layer here would cost another 1,049,600 and nearly double the model, because both of its sides are wide. Widening one layer costs linearly in that width against each neighbour, so widening next to a narrow neighbour is the cheap direction.

saying these in an interview costs you the question

  • Trims the 40-input first layer to save space
  • Assumes depth drives the count more than width
  • Adds widths instead of multiplying adjacent pairs
  • Halves every width without noticing the total quarters
  • Tries to cut the 250 output classes, which the task fixes
  • Ignores whether the narrowed layer still fits the number of classes

context