In a vanilla encoder-decoder RNN for translation, what does the encoder pass to the decoder?
answer
- two stacks, one wire between them
- the encoder's last time step only
- it seeds the decoder's starting state
- same size for three words or sixty
- per-step encoder states are thrown away
basics
~10 sOnly its final hidden state: one fixed-size vector, plus the cell state if the encoder is an LSTM. That single tensor initialises the decoder, and nothing else crosses between the two stacks.
solid answer
~50 sThe encoder is a recurrent stack that consumes the source tokens one at a time, updating a hidden state. When the last source token has been read, its final hidden state - for an LSTM, the final hidden and cell state of each layer - is handed to the decoder as the decoder's initial state. That handoff is the *only* channel between the halves: the decoder never sees the encoder's per-token states. From there the decoder runs its own recurrence, taking the previously emitted token's embedding at each step and producing a distribution over the target vocabulary. The context vector's size is fixed by the hidden width, typically on the order of a thousand floats, and does not grow with the source. A three-word sentence and a sixty-word sentence are both squeezed into the same number of numbers.
go deeper
Be ready to draw two boxes and name the wire between them: the encoder's final hidden state, one fixed-size vector, used as the decoder's initial state. Say out loud that the encoder's per-step states are discarded.
Explain the shapes. State that the context is hidden width times layers, times two for an LSTM's cell state, and that it is independent of input length. Describe what the decoder consumes at each output step.
Point at the consequences of a single write-once channel: the decoder cannot revisit a source position, the first output token waits for the entire source to be read, and the model has an information budget fixed before training.
Frame the design tradeoff it bought: any-length in, any-length out, trained end to end with no hand-built alignment. Be able to say what the field gave up to get that and what later designs paid to get it back.
## The two-stack picture A vanilla sequence-to-sequence model is two recurrent networks wired end to end. **The encoder** reads the source sequence left to right. At step `t` it takes the embedding of source token `x_t` and its own previous hidden state and produces a new hidden state: `h_t = f(h_(t-1), x_t)`. With an LSTM, each layer carries two tensors - a hidden state `h_t` and a cell state `c_t`. After the final source token (usually followed by an end-of-sequence marker) the encoder stops. Its per-step states `h_1 ... h_n` are discarded. **The context vector** is what survives: `h_n`, the state after the last source token, concatenated across layers and, for an LSTM, together with `c_n`. This is the fixed-length context vector the leaf is named for. **The decoder** is a second recurrent stack whose initial state is set to that context vector. At each output step it consumes the embedding of the previously emitted target token and its own running hidden state, emits a hidden state, and projects that through a linear layer plus softmax to a distribution over the target vocabulary. It repeats until it emits the end-of-sequence marker. ## Why there is exactly one channel The defining property is not that a vector is passed - it is that this is the *whole* interface. The decoder has no other view of the source. It cannot revisit source position 7 when it is about to emit the word that corresponds to it; that position exists only insofar as it left a trace in `h_n`. Everything the model will ever know about the input - which entities appeared, in what order, with what tense, in what register - has to be resident in those numbers at the moment the decoder takes its first step. This is also why the architecture was such a big deal when it appeared. Before it, mapping a variable-length input to a variable-length output of a *different* length required hand-built alignment machinery. The fixed-vector handoff removes that: the encoder absorbs any length, the decoder emits any length, and the whole thing trains end to end from source-target pairs with one loss. ## Sizes and common variants The context vector's dimensionality is a hyperparameter - the hidden width, times the number of layers, times two if the cell state travels as well. It is decided at design time and is completely independent of the input. That independence is the point of the design and, later, the source of its main weakness. A few wiring variants show up and are worth recognising: - **Initial-state only.** The context sets the decoder's state at step zero and is never referenced again. The plain form. - **Fed at every step.** The same static context vector is also concatenated to the decoder's input at every output step, so the summary cannot be gradually forgotten as the decoder's own state drifts. Note that the vector is still the same vector at every step - it is not recomputed. - **Deep stacks.** With multiple layers, the corresponding layer's final state seeds the corresponding decoder layer, so the context is really a tuple of tensors; it is still one fixed-size bundle. - **Bidirectional encoders.** Running a second encoder backwards and concatenating gives a richer final state, but it is still one final state. ## What an interviewer is checking Three things. First, that you can draw the diagram and name which tensor crosses the boundary. Second, that you know the encoder finishes reading before the decoder starts - the model cannot emit a first target word until it has consumed the entire source, which is a real latency property for long inputs. Third, that you see the consequence: because the channel is fixed-width and written once, this architecture has an information budget that does not scale with the input, and the whole later history of sequence modelling is a response to that. A quick way to feel it: imagine an abstractive summarizer asked to compress a two-thousand-word article. The encoder must finish the article and commit everything the summary might need to roughly a thousand floats before the decoder writes its first word - and it must do that without knowing which facts the summary will end up needing.
- Does the size of the context vector depend on how long the source sentence is?No. Its dimensionality is the encoder's hidden width, times the number of layers, times two if an LSTM's cell state travels alongside the hidden state. That is fixed at design time. A three-token source and a sixty-token source produce context vectors of identical shape, which is exactly why longer sources have to lose more.
- Some designs feed the context vector into every decoder step instead of only the first. What does that change?It stops the summary from being crowded out as the decoder's own state evolves over a long output, so late target words still see the source information. It does not change what is in the vector: the same static summary is presented every step, so the model still cannot bring back detail the encoder dropped.
- Can the decoder start emitting before the encoder has read the whole source?Not in this architecture. The context vector is the encoder's state after the final source token, so the first target step waits for the full source. That makes the source pass strictly sequential and adds latency proportional to input length, and it rules out streaming translation without changing the design.
It is like reading a long report, then throwing the report away and briefing a colleague from a single index card. Whatever did not fit on the card is gone before the briefing starts.
saying these in an interview costs you the question
- Says the decoder sees all of the encoder's per-step hidden states
- Thinks the context vector grows with source length
- Confuses the context vector with the last source token's embedding
- Claims the encoder and decoder share one set of weights by definition
- Cannot say which tensor crosses the encoder-decoder boundary