skip to content

How many parameters does an LSTM cell with 300-dimensional inputs and 256 hidden units have?

level: juniorimportance: should knowfreq 54%

answer

  1. count the blocks first, then one block
  2. each block reads input and previous hidden
  3. do not forget the bias term
  4. four times an ungated cell of the same width
  5. 4 * (d_in + d_h + 1) * d_h

basics

~20 s

Four blocks — forget, input, output and candidate — each hold an input matrix, a recurrent matrix and a bias: 4 * (300 + 256 + 1) * 256 = 570,368 parameters, four times a plain recurrent cell.

solid answer

~50 s

The cell computes four blocks from `x_t` and `h_(t-1)`, and each block has the same shape: an input matrix of `d_h` by `d_in`, a recurrent matrix of `d_h` by `d_h`, and a bias of length `d_h`. So one block is `(d_in + d_h + 1) * d_h` and the cell is `4 * (d_in + d_h + 1) * d_h`. With `d_in = 300` and `d_h = 256` that is `4 * 557 * 256 = 570,368`. The general shape of the formula is worth internalising: for large hidden sizes the `d_h * d_h` term dominates, so parameters grow roughly with `4 * d_h^2` — doubling the hidden width roughly quadruples the cell. A plain recurrent cell of identical width would need exactly one block, 142,592 parameters here, which is where the familiar four-times ratio comes from. Note the count is per step-independent: unrolling over 400 steps reuses the same weights and adds nothing.

go deeper

for a junior

Memorise the structure, not just the number: four blocks, each reading the input and the previous hidden state, each with a bias. Being able to substitute two widths into the formula on a whiteboard is the whole bar.

for a middle

Explain why there are exactly four blocks and be ready for the follow-up on which convention you counted biases under. Know that stacking layers changes the second layer's input width to the first layer's hidden width.

for a senior

Use the formula to reason about cost: hidden width is quadratic and is the knob that hurts, bidirectional doubles it, and sequence length costs activation memory rather than parameters.

for a principal

Turn the count into a budgeting argument — where the parameters buy capacity versus where they buy only latency, and whether the four-times cost of gating is justified against the simpler recurrence for the dependency lengths your data actually contains.

## Where the four comes from A gated cell computes four things at every step from the same two inputs, the current input `x_t` and the previous hidden state `h_(t-1)`: - the forget gate - the input gate - the output gate - the candidate Each is a separate affine map with its own parameters — that is exactly what makes them able to disagree with each other. So the weights come in four identically-shaped blocks. ## One block A single block maps a `d_in` vector and a `d_h` vector to a `d_h` vector: ``` block = W (d_h by d_in) + U (d_h by d_h) + b (d_h) = d_h * d_in + d_h * d_h + d_h = (d_in + d_h + 1) * d_h ``` The `+ 1` is the bias column, and forgetting it is the usual small slip. Some formulations use two biases per block, one on the input path and one on the recurrent path, which adds another `d_h` per block; if you are asked for an exact number it is fair to say which convention you are counting. ## The worked number With `d_in = 300` and `d_h = 256`: ``` per block = (300 + 256 + 1) * 256 = 557 * 256 = 142,592 cell = 4 * 142,592 = 570,368 ``` Under the two-bias convention it would be `570,368 + 4 * 256 = 571,392` — close enough that the structural reasoning, not the exact digit, is what an interviewer is grading. ## What the formula tells you **Four times a plain cell.** A non-gated recurrent cell of the same widths computes one block, so 142,592 parameters. Gating costs exactly four times the weights, and roughly four times the arithmetic per step. That is the price of the gates, and it is worth naming as a price when someone asks whether to use one. **Quadratic in hidden width.** The `d_h * d_h` recurrent term dominates once `d_h` is comparable to or larger than `d_in`. Going from 256 to 512 hidden units here takes the count from 570,368 to `4 * (300 + 512 + 1) * 512 = 1,665,024` — nearly triple, and heading toward four times as `d_h` grows. Hidden width is the expensive knob; input width is usually the cheap one. **Independent of sequence length.** Unrolling the cell over 400 timesteps does not create 400 copies of the weights; the same block is applied at every step. Sequence length costs activation memory and wall-clock time, not parameters. Candidates who multiply by the number of steps are revealing that they think of the unrolled diagram as the model. **Layers and directions multiply.** Stacking two such cells is roughly two cells' worth, except the second cell's input width is the first's hidden width. Running the sequence in both directions and concatenating is two independent cells, so double again — and it also doubles what the layer above receives. ## The answer an interviewer wants They are checking three things: that you know there are four blocks and can name them, that you know each block reads both the input and the previous hidden state, and that you remember the bias. Say `4 * (d_in + d_h + 1) * d_h`, substitute, and give the four-times comparison against an ungated cell. Adding the observation that parameters are independent of sequence length is the part that separates a memorised formula from an understood one.

  • Does unrolling the cell over 400 timesteps change the parameter count?
    No. The same four blocks are applied at every step — that weight sharing is what makes it a recurrent model at all. Sequence length costs stored activations for the backward pass and wall-clock time, both roughly linear in the number of steps, but the parameter count is fixed by the widths alone.
  • Which knob dominates the count as the cell gets bigger, the input width or the hidden width?
    The hidden width, because it appears twice — once in the recurrent matrix, which is `d_h` by `d_h`, and once as the output size of both matrices. Parameters grow roughly with the square of hidden width and only linearly with input width, so widening the hidden state is by far the more expensive change.

saying these in an interview costs you the question

  • Multiplies the count by the number of timesteps
  • Counts three blocks and forgets the candidate
  • Omits the bias term from each block
  • Forgets that each block also reads the previous hidden state
  • Says gating is free relative to a plain recurrent cell

context