skip to content

Why does a vanilla RNN reuse the same weight matrices at every time step?

level: middleimportance: must knowfreq 62%

answer

  1. one function, applied repeatedly
  2. parameter count versus sequence length
  3. step 900 would have almost no data
  4. time-invariance, like a reused filter

basics

~20 s

Reusing one input-to-hidden matrix, one hidden-to-hidden matrix and one bias makes the layer a single function applied repeatedly. The parameter count is then independent of sequence length, any length runs, and what is learned at step 1 applies at step 900.

solid answer

~50 s

The recurrent layer defines one function and applies it once per step, so the same `W_xh`, `W_hh` and bias act at step 1 and at step 900. Three things follow. First, the parameter count is `H*d + H*H + H` and does not depend on the sequence length at all — a 500-step input adds nothing. Second, the model can accept a length it never saw in training, because there is no per-step matrix that would be missing. Third, and most important statistically, every step's data trains the same weights, so a pattern is learned once and reused everywhere instead of being relearned at each position from a fraction of the data. The assumption you are buying into is time-invariance: the mapping from state and observation to next state is taken to be the same everywhere in the sequence.

code

python · 20 lines
python
import math

# vanilla RNN cell: h_t = tanh(W_xh x_t + W_hh h_(t-1) + b)
W_xh = [[0.5], [-0.3]]            # hidden 2 x input 1
W_hh = [[0.1, 0.4], [0.2, -0.5]]  # hidden 2 x hidden 2
b = [0.0, 0.1]

def step(x, h):
    return [math.tanh(sum(W_xh[i][j] * x[j] for j in range(len(x)))
                      + sum(W_hh[i][k] * h[k] for k in range(len(h)))
                      + b[i])
            for i in range(len(b))]

h = [0.0, 0.0]                     # h_0
for x in ([1.0], [0.5], [-2.0]):   # the SAME weights are reused at every step
    h = step(x, h)
    print([round(v, 4) for v in h])

n_params = 2 * 1 + 2 * 2 + 2       # W_xh + W_hh + b
print("params:", n_params, "- identical for a 3-step or a 900-step sequence")

go deeper

for a junior

Be ready to say that one set of matrices is applied at every step, and that the number of parameters therefore does not depend on how long the sequence is.

for a middle

Expect to compute the parameter count from input and hidden sizes and to explain what fails if each step had its own matrix: no weights for unseen positions and almost no data per position.

for a senior

Show you can spot when the time-invariance assumption breaks in real data and fix it by feeding positional information rather than by un-sharing weights.

for a principal

Own the inductive-bias framing: sharing across time is a prior you choose, and you should be able to argue when that prior pays and when the problem calls for an architecture that treats positions distinctly.

## What sharing means concretely A recurrent layer holds three parameter tensors: `W_xh` (H-by-d, input-to-hidden), `W_hh` (H-by-H, hidden-to-hidden) and a bias `b` (H). The update `h_t = tanh(W_xh x_t + W_hh h_(t-1) + b)` uses **those same three objects** at every t. "Unrolling" the network over time draws a deep chain of blocks, and it is easy to misread that picture as a deep feed-forward network with many layers. It is not: every block in the chain is the *same* layer, invoked again. ## Consequence 1 — parameters do not scale with length Count them: `H*d` for the input matrix, `H*H` for the recurrent matrix, `H` for the bias, giving `H*(d + H + 1)`. For d = 100 and H = 256 that is 25,600 + 65,536 + 256 = 91,392 parameters. Feed it a 3-step sequence or a 900-step sequence and the number is unchanged. (Some formulations carry two bias vectors, one per matrix, which adds H and changes nothing conceptually.) What *does* grow with length is compute and activation memory: T applications means T matrix-vector products and, during training, T stored intermediate states. Length costs time and memory, never parameters. ## Consequence 2 — variable-length inputs become trivial Because the layer only ever needs "one observation plus my previous state", it can run as many times as there are observations. A character-level surname classifier that sees a 4-letter name and a 19-letter name uses the identical weights on both; it just performs 4 versus 19 updates and reads the final state. With per-step weights there would be no matrix number 19 unless a training name had reached that length, and any longer input at inference would be unrunnable. ## Consequence 3 — the statistical argument, which is the real one Suppose you did give each step its own matrix. Formally the model is now a plain deep network of depth T with no reuse. Two failures follow immediately. 1. **Data starvation per position.** The matrix at step 900 is updated only by the examples long enough to reach step 900, and only by their step-900 observation. Late positions see a small, biased slice of the data; the parameter count explodes (T times the shared count) while the effective data per parameter collapses. Overfitting is close to guaranteed. 2. **No transfer across positions.** A motif that means the same thing wherever it occurs — a rising edge in a sensor trace, a suffix in a name — would have to be relearned independently at every position. Sharing is to time what a convolution kernel's sharing is to space: it encodes the prior that the rule is position-independent, and it multiplies the effective training signal per parameter by the number of steps. Because the same weights are used at every step, the training signal from all steps accumulates into that one set of matrices; the mechanics of how those contributions are combined during the backward pass through an unrolled graph is a separate topic, but the parameter identity is what makes it possible at all. ## What sharing assumes, and when it is wrong Sharing asserts time-invariance: the transition rule does not depend on *where* you are in the sequence. That is right for streams and text, and wrong for fixed-format records where position carries meaning (field 3 is always the country code). The fix is not to un-share the weights — that reintroduces every problem above — but to **feed position as an input**, appending a step index or a positional encoding to `x_t` so one shared function can condition on where it is. Same weights, more informative input. ## The interview answer Lead with "it is one function applied repeatedly, not T different layers". Then give the parameter count and note it is length-independent, note that variable and unseen lengths therefore work, and finish with the statistical point: each position's data trains the same weights, so the model generalizes across time instead of memorizing each position. If asked what breaks without sharing, name the two failures — no weights for unseen positions, and each position trained on a sliver of the data.

  • If sequence length does not change the parameter count, what does it change?
    Compute and memory. Length T means T applications of the update, so forward cost grows linearly in T, and training must keep the per-step intermediate states around for the backward pass, so activation memory grows linearly too. It also deepens the composed function the gradients must travel through, which is where long sequences get hard.
  • Does weight sharing assume anything about the data?
    Yes — that the transition rule is the same wherever you are in the sequence. When position genuinely carries meaning, such as a fixed-format record whose third field is always a country code, the right response is to feed a step index or positional feature as part of the input, not to give each step its own matrix.
  • How is a vanilla RNN cell's parameter count computed?
    With input dimension d and hidden size H it is H*d for the input-to-hidden matrix, H*H for the recurrent matrix and H for the bias, so H*(d + H + 1). Note the H*H term dominates once the hidden size exceeds the input size, which is why doubling the hidden state roughly quadruples the layer.

It is one recipe you follow once per ingredient, not a separate recipe written for ingredient number 900.

saying these in an interview costs you the question

  • Says the parameter count grows with sequence length
  • Thinks unrolling creates T genuinely different layers
  • Claims sharing exists only to save memory
  • Cannot say what would break for an unseen longer input
  • Confuses sharing the weights with carrying the hidden state

context