RNNs, LSTMs and Attention
You will learn how recurrent nets process sequences, exactly which gradient problem LSTM gates solve, and how attention removed the seq2seq bottleneck. Interviewers love 'walk me from RNN to attention' because it tests whether you understand the lineage transformers came from.
on this pageshowhide
explore
- Recurrence and BPTT19 questions
- Hidden State and Weights3 questions
- Backprop Through Time3 questions
- Vanishing Gradients Over Time3 questions
- Sequence Framings and Stacking3 questions
- Padding and Masking3 questions
- Recurrent Forecasters4 questions
- Gated Recurrent Cells8 questions
- Cell State and Gates4 questions
- GRU and Gating Limits4 questions
- Seq2Seq and Alignment12 questions
- Fixed-Length Context Bottleneck3 questions
- Teacher Forcing3 questions
- Bahdanau and Luong Scoring3 questions
- Removing the Recurrence3 questions
questions
page 2 of 2At generation time, what does a recurrent model's fixed-size state buy over a position-parallel model?
basics
~20 sConstant cost per emitted token. A recurrent cell folds all history into one hidden vector, so step 1,000 costs what step 1 costs and memory stays flat. Without recurrence, each new position is computed against every earlier one, so per-step cost grows.
A third stacked recurrent layer barely improves your tagger - what do you check?
basics
~20 sCheck whether depth is the bottleneck at all: that the stack is wired correctly, that the extra layer is actually training, and whether the remaining errors come from missing data or missing context rather than from too little capacity.
Why does a sentiment RNN learn the final clause of a review but ignore the opening sentence?
basics
~20 sGradients from early time steps decay geometrically on the way back, so the opening sentence gets almost no learning signal. Recent steps keep full-size gradients, so the model quietly becomes a short-memory model reading the ending.
Your GRU-versus-LSTM benchmark gap is inside seed-to-seed noise — which cell do you ship and how do you report it?
basics
~20 sReport it as no measurable difference at this budget, not as a win. Then decide on secondary criteria you can defend — parameter count, per-step latency, tuning cost — and state which quantity you held fixed in the comparison.
With 120 daily observations from a single ATM, would you ship a recurrent forecaster or a classical model?
basics
~20 sShip the classical model. One short series yields roughly a hundred overlapping windows and about seventeen weekly cycles, far too little to fit thousands of recurrent weights. Deep forecasters earn their keep by pooling many related series, not on one short one.
Why is a 64-unit GRU slow at streaming inference despite its tiny parameter count?
basics
~20 sBecause latency is set by the number of sequential steps, not by arithmetic. Each step needs the previous hidden state, so 1,000 frames means 1,000 tiny dependent operations, each too small to keep the hardware busy.
An LSTM loses a patient's ventilation flag long before step 400 of a vitals stream — how do you diagnose it?
basics
~20 sTrace forget-gate activations per unit across the sequence: a held fact shows forget near 1 with input near 0. If no unit holds high, memory is decaying — initialise the forget-gate bias positive and check the backward pass is not truncated.
Why group a speech corpus of 1-30 second utterances into similar-length batches?
basics
~20 sBatching similar lengths together shrinks each batch's padded width, so far less compute goes into pad steps. The price is correlated batches: examples no longer arrive in random order, so you shuffle inside buckets and randomise bucket order every epoch.
When a long-sequence RNN ignores distant context, how do you tell a gradient-reach limit from a data problem?
basics
~10 sSeparate them with controlled experiments: a synthetic task with a known dependency lag measures the model's reach independently of your data, and a short-context baseline measures whether distant context carries signal at all.
showing 31–39 of 39