skip to content

In a vanilla RNN, how does the hidden state at step t depend on earlier inputs?

level: juniorimportance: must knowfreq 78%

answer

  1. one vector, rewritten each step
  2. previous state feeds the next
  3. unrolling reaches every earlier input
  4. fixed-size lossy summary of the prefix

basics

~20 s

A vanilla RNN computes h_t = tanh(W_xh x_t + W_hh h_(t-1) + b). Since h_(t-1) was built the same way from h_(t-2), the state at step t is a fixed-size summary of every input so far.

solid answer

~40 s

A recurrent layer keeps one vector, the hidden state, and rewrites it once per time step: `h_t = tanh(W_xh x_t + W_hh h_(t-1) + b)`. The input-to-hidden matrix reads the new observation, the hidden-to-hidden matrix reads what the layer already knew, and the nonlinearity squashes the sum. Since h_(t-1) was produced by the same rule from h_(t-2), unrolling the recursion shows h_t is a function of x_1 through x_t — the dependence on earlier steps is indirect, routed through the state rather than through any direct connection. The state is a fixed-size vector whose dimension you choose, so it is a lossy, learned compression of the prefix, not a store of the raw history. Training decides what is worth keeping in it.

go deeper

for a junior

Be ready to write the update equation from memory and say in one sentence which term brings in the new observation and which brings in the past.

for a middle

Explain the substitution that makes h_t depend on x_1, and be able to say what the hidden size controls and what it costs.

for a senior

Expect to reason about the bottleneck in practice: how you would notice that a hidden size is too small, and what the fixed-size summary implies for very long inputs.

for a principal

Own the framing choice: when a compressed running state is the right representation for the problem at all, versus a model that can address earlier positions directly.

## The one thing a recurrent layer holds A feed-forward layer maps one input vector to one output vector and remembers nothing between calls. A recurrent layer adds exactly one piece of machinery: a **hidden state**, a vector `h` of a size you choose (say 128), that persists from one time step to the next and is rewritten at every step. The standard vanilla (Elman) update is ``` h_t = tanh(W_xh x_t + W_hh h_(t-1) + b) ``` where `x_t` is the observation at step t (a d-dimensional vector), `W_xh` is an H-by-d **input-to-hidden** matrix, `W_hh` is an H-by-H **hidden-to-hidden** (recurrent) matrix, `b` is an H-dimensional bias, and `tanh` is applied element-wise. The two matrix-vector products live in the same H-dimensional space, so they can simply be added: the layer mixes "what I just saw" with "what I already knew" and squashes the result into (-1, 1). ## Why h_t reaches back further than one step Nothing in the equation connects `h_t` to `x_1` directly. The reach comes from substitution. Expand once: ``` h_t = tanh(W_xh x_t + W_hh tanh(W_xh x_(t-1) + W_hh h_(t-2) + b) + b) ``` Keep substituting and you reach `h_0`. So `h_t` is a nested composition — t applications of the same function — over `x_1 ... x_t`. This is what people mean when they say an RNN has *memory*: the dependence is real but **indirect**, carried entirely by the state vector. It is not a sliding window; there is no explicit "look back k steps" anywhere in the layer. A common beginner error is to imagine the layer receiving the last k inputs at once. It receives exactly one input per step, plus its own previous output. The practical consequence is that the influence of an early input on `h_t` has to survive being passed through the recurrence hundreds of times. Whether it survives is a training question, and repeated application of the same map is precisely what makes long-range credit assignment hard in vanilla recurrence. ## A fixed-size, lossy summary The state has H components no matter how long the sequence is. A 5-step sequence and a 900-step sequence both end with an H-dimensional vector. That is a hard information bottleneck: everything the model will ever use about the prefix must fit into H numbers. H is a capacity knob — too small and distinct histories collapse onto the same state; too large and you burn parameters (the recurrent matrix alone is H-by-H) and overfit. So the right mental model of `h_t` is a **learned running summary**, not a log. Training shapes `W_hh` to decide what gets preserved across steps and what gets overwritten, and shapes `W_xh` to decide how a new observation perturbs that summary. ## Where the state starts At the beginning of a sequence you need an `h_0`. The default is a zero vector, which says "no prior context". You can instead make `h_0` a trainable parameter, so the model learns its own preferred starting summary; this occasionally helps on short sequences where the first few steps matter a lot, and is usually a rounding error otherwise. What matters far more is that `h_0` is chosen deliberately: on data that arrives as one unbroken stream, the question of when to reset to `h_0` at all becomes a real design decision rather than a default. ## Reading something out of the state The hidden state is not itself a prediction. Downstream layers consume it — typically a linear map from H dimensions to the number of outputs. Consider a character-level classifier that predicts a surname's language of origin: it is fed one character at a time, the state is updated per character, and the classifier reads the state after the last character. Names of length 4 and length 19 are handled by the same layer, because the layer only ever sees one character plus its own state; the number of repetitions changes, the function does not. ## What to say in an interview State the update equation, say which two matrices are involved and what each reads, then make the key point explicitly: the state is a fixed-size compression of the whole prefix, updated in place, and the dependence on early inputs is transitive through the state rather than direct. If you add one sentence about the bottleneck that H imposes, you have said everything the question is testing.

  • Why call h_t a summary rather than a memory of every input?
    Because its dimension is fixed no matter how long the sequence runs. A 900-step sequence still ends in the same H numbers as a 5-step one, so distinct histories must share that space and some detail is necessarily discarded. Training decides which detail survives, which is why the hidden size is a real capacity choice rather than a formality.
  • What is h_0, and does it matter what you set it to?
    It is the state fed in before the first observation. A zero vector is the usual default and means "no prior context". You can make it a trainable parameter so the model learns a preferred starting point, which helps a little on short sequences. The bigger question is not its value but when you reset to it.
  • How does the layer handle inputs of different lengths at all?
    It applies one update per observation, so length only changes how many times the update runs. A four-character name gets four applications, a nineteen-character name nineteen, and both finish with a state of the same size. Nothing in the layer refers to a sequence length.

It is like keeping a running score in your head while someone reads out numbers: you never re-read the earlier numbers, you just update the single figure you are holding.

saying these in an interview costs you the question

  • Says the hidden state stores every past input verbatim
  • Describes the RNN as reading a fixed window of recent inputs
  • Confuses the hidden state with the model's prediction
  • Thinks the state vector grows as the sequence gets longer
  • Cannot name which two matrices feed the update

context