skip to content

RNNs, LSTMs and Attention

You will learn how recurrent nets process sequences, exactly which gradient problem LSTM gates solve, and how attention removed the seq2seq bottleneck. Interviewers love 'walk me from RNN to attention' because it tests whether you understand the lineage transformers came from.

on this pageshow

explore

questions

page 2 of 2

At generation time, what does a recurrent model's fixed-size state buy over a position-parallel model?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Constant cost per emitted token. A recurrent cell folds all history into one hidden vector, so step 1,000 costs what step 1 costs and memory stays flat. Without recurrence, each new position is computed against every earlier one, so per-step cost grows.

open as a page

A third stacked recurrent layer barely improves your tagger - what do you check?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Check whether depth is the bottleneck at all: that the stack is wired correctly, that the extra layer is actually training, and whether the remaining errors come from missing data or missing context rather than from too little capacity.

open as a page

Why does a sentiment RNN learn the final clause of a review but ignore the opening sentence?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Gradients from early time steps decay geometrically on the way back, so the opening sentence gets almost no learning signal. Recent steps keep full-size gradients, so the model quietly becomes a short-memory model reading the ending.

open as a page

Your GRU-versus-LSTM benchmark gap is inside seed-to-seed noise — which cell do you ship and how do you report it?

level: principalimportance: should knowfreq 38%

basics

~20 s

Report it as no measurable difference at this budget, not as a win. Then decide on secondary criteria you can defend — parameter count, per-step latency, tuning cost — and state which quantity you held fixed in the comparison.

open as a page

With 120 daily observations from a single ATM, would you ship a recurrent forecaster or a classical model?

level: principalimportance: should knowfreq 44%

basics

~20 s

Ship the classical model. One short series yields roughly a hundred overlapping windows and about seventeen weekly cycles, far too little to fit thousands of recurrent weights. Deep forecasters earn their keep by pooling many related series, not on one short one.

open as a page

Why is a 64-unit GRU slow at streaming inference despite its tiny parameter count?

level: seniorimportance: nice to knowfreq 40%

basics

~20 s

Because latency is set by the number of sequential steps, not by arithmetic. Each step needs the previous hidden state, so 1,000 frames means 1,000 tiny dependent operations, each too small to keep the hardware busy.

open as a page

An LSTM loses a patient's ventilation flag long before step 400 of a vitals stream — how do you diagnose it?

level: seniorimportance: nice to knowfreq 40%

basics

~20 s

Trace forget-gate activations per unit across the sequence: a held fact shows forget near 1 with input near 0. If no unit holds high, memory is decaying — initialise the forget-gate bias positive and check the backward pass is not truncated.

open as a page

Why group a speech corpus of 1-30 second utterances into similar-length batches?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Batching similar lengths together shrinks each batch's padded width, so far less compute goes into pad steps. The price is correlated batches: examples no longer arrive in random order, so you shuffle inside buckets and randomise bucket order every epoch.

open as a page

When a long-sequence RNN ignores distant context, how do you tell a gradient-reach limit from a data problem?

level: principalimportance: nice to knowfreq 27%

basics

~10 s

Separate them with controlled experiments: a synthetic task with a known dependency lag measures the model's reach independently of your data, and a short-context baseline measures whether distant context carries signal at all.

open as a page

showing 31–39 of 39