skip to content

How do you count the parameters of a 784-128-10 fully-connected neural network?

level: juniorimportance: must knowfreq 70%

answer

  1. adjacent widths, not the list length
  2. multiply the pair, then add one term
  3. one extra number per output unit
  4. 784 x 128 + 128, then 128 x 10 + 10

basics

~20 s

Each fully-connected layer holds inputs times outputs weights plus one bias per output unit. For 784-128-10 that is 784 x 128 + 128 = 100,480 and 128 x 10 + 10 = 1,290, totalling 101,770 parameters.

solid answer

~50 s

A fully-connected layer stores one weight per input-output pair plus one bias per output unit, so a layer from `n_in` to `n_out` units holds `n_in * n_out + n_out` parameters. Walk the network as adjacent width pairs: for 784-128-10 the first layer is `784 * 128 + 128 = 100,480` and the second is `128 * 10 + 10 = 1,290`, giving **101,770** in total. The input itself contributes nothing — the 784 only appears as the fan-in of the first weight matrix — and each hidden width is used exactly twice, once as an output width and once as the next layer's input width. Nothing about the data changes this: batch size, dataset size and epoch count leave the total untouched, because the same weights are reused for every row. Biases are only 138 of the 101,770 here, but forgetting them reads as carelessness.

code

python · 12 lines
python
widths = [784, 128, 10]
total = 0
for n_in, n_out in zip(widths, widths[1:]):
    weights = n_in * n_out
    layer = weights + n_out
    print(n_in, "->", n_out, ":", weights, "weights +", n_out, "biases =", layer)
    total += layer
print("total:", total)

# 784 -> 128 : 100352 weights + 128 biases = 100480
# 128 -> 10 : 1280 weights + 10 biases = 1290
# total: 101770

go deeper

for a junior

Be ready to produce the number on a whiteboard in under a minute: inputs times outputs, plus one bias per output, summed over adjacent width pairs. Say the biases out loud even though they are tiny.

for a middle

Explain why the formula is what it is — every output unit owns a weighted sum over all inputs plus its own offset — and why transposing the weight matrix changes notation but not the count.

for a senior

Show that you know which numbers are not in the count: batch size, dataset size, epochs and plain activation functions. Interviewers use this to check that you separate model size from training setup.

for a principal

Own the convention. When a team quotes a parameter total across models, say exactly what is counted, so that comparisons between architectures mean the same thing to everyone reading the number.

## What counts as a parameter A **parameter** is a number that training updates: the weights and the biases stored inside the layers. It is not a hyperparameter (the layer widths, the depth, the learning rate — you choose those), and it has nothing to do with the data (the number of training rows, the batch size, the number of epochs). When someone says "this model has 101,770 parameters", they mean: this is how many numbers the optimizer is allowed to move. ## One fully-connected layer A fully-connected (dense) layer takes `n_in` numbers in and produces `n_out` numbers out. Every output unit computes a weighted sum over **all** of the inputs and then adds its own bias: ``` z_j = b_j + sum over i of ( w_ij * x_i ) ``` So one output unit owns `n_in` weights plus `1` bias — `n_in + 1` numbers. There are `n_out` such units, giving ``` params = n_in * n_out + n_out = n_out * (n_in + 1) ``` The first term is the weight matrix: one entry per (input unit, output unit) pair. The second term is the bias vector: one entry per output unit. Whether you store the matrix as `n_in` by `n_out` or transposed as `n_out` by `n_in` is a notational choice and does not change the count — only the two widths matter. ## Walking a whole MLP Write the network as a list of widths and step through **adjacent pairs**. For the classic handwritten-digit MLP with a 784-pixel input, one 128-unit hidden layer and 10 output classes, the widths are `[784, 128, 10]`, which gives two pairs: - `784 -> 128`: `784 * 128 = 100,352` weights `+ 128` biases `= 100,480` - `128 -> 10`: `128 * 10 = 1,280` weights `+ 10` biases `= 1,290` Total: **101,770 parameters**. Two things trip people up here. First, the input "layer" contributes nothing of its own — it is just the data arriving, and the 784 only appears because it is the fan-in of the first weight matrix. Second, each hidden width is used **twice**: 128 is the output width of the first layer and the input width of the second. It is counted once in each product, never twice in the same product. The pattern generalises to any depth. With widths `[w0, w1, ..., wL]` the total is the sum over consecutive pairs of `w_(k) * w_(k+1) + w_(k+1)`. ## Biases: always count them, never worry about them In the digit MLP the biases are `128 + 10 = 138` out of 101,770 — about **0.14 percent**. The gap widens as layers get wider, because weights grow with the product of two widths while biases grow with just one. An eight-layer trunk of 2048-by-2048 layers carries `8 * 2048 * 2048 = 33,554,432` weights against `8 * 2048 = 16,384` biases — under **0.05 percent**. That is the whole story of biases in a parameter count: they are the term candidates forget, and the term that never changes the answer. Forgetting them in an interview reads as carelessness rather than as a rounding decision, so say them out loud even while noting they are negligible. Dropping them is a real design choice in some layers, but it is a modelling decision, not a way to make a model smaller. ## What does not change the count - **Batch size.** The same weight matrix is applied to every row in the batch; 1 row or 1,024 rows, the parameters are identical. - **Dataset size and epochs.** Training longer or on more data moves the parameters; it does not create more of them. - **Most activation functions.** A ReLU, a sigmoid or a tanh has no learnable numbers at all, so it adds zero. Only a *parametric* activation, which learns a coefficient of its own, contributes anything — and then only a handful of values. ## Doing it fast, out loud Interviewers often want the count on a whiteboard in under a minute, so practise the shortcut: multiply the adjacent widths, keep the largest product first, then add the small change. For `[784, 128, 10]`: "about 100k in the first layer, about 1.3k in the second, so roughly 101.7k — 101,770 exactly with biases." Leading with the dominant term shows you know where the mass sits, and it makes an arithmetic slip in the small term harmless. ## Where hand counts go wrong Counting the input units as parameters; adding the widths instead of multiplying them; skipping the bias vector; counting a hidden layer's parameters once for the layer below and again for the layer above; and quoting a number that silently includes or excludes something the interviewer did not ask about. Say your convention as you go — "weights plus biases, all layers, nothing else" — and the count becomes checkable rather than a magic number.

  • What happens to that total if the layers carry no bias terms?
    It drops by 138, from 101,770 to 101,632 — about 0.14 percent. Biases grow with a single width while weights grow with a product of two, so they are always a rounding error in a wide network. Removing them is a modelling choice, never a way to make a model meaningfully smaller.
  • Does batch size or the number of training examples change the parameter count?
    No. The same weight matrix is applied to every row in a batch and to every example in the dataset, so none of those numbers appear in the count. Only the layer widths do. Batch size and dataset size change how training behaves, not how many numbers the optimizer owns.
  • Where do hand counts of an MLP most often go wrong?
    Four places: counting the input units as if they were parameters, adding the widths instead of multiplying them, skipping the bias vector, and counting a hidden width once for the layer below and again for the layer above. Saying your convention out loud — weights plus biases, all layers, nothing else — makes the number checkable.

A weight matrix is a full seating chart: one entry for every input-output pair, with no gaps. The bias vector is one spare slot per output unit, added on the side.

saying these in an interview costs you the question

  • Counts the input units themselves as parameters
  • Forgets the bias vector entirely
  • Adds the layer widths instead of multiplying adjacent pairs
  • Thinks batch size or dataset size changes the count
  • Counts a hidden layer's parameters twice, once per neighbour
  • Assumes a ReLU or sigmoid contributes parameters of its own

context