A third stacked recurrent layer barely improves your tagger - what do you check?
answer
- Ask what the layer above consumes
- Is the extra layer even training?
- Look at which errors actually remain
- Depth costs another sequential pass
- Depth does not extend time reach
basics
~20 sCheck whether depth is the bottleneck at all: that the stack is wired correctly, that the extra layer is actually training, and whether the remaining errors come from missing data or missing context rather than from too little capacity.
solid answer
~50 sWork in three steps rather than adding layers on instinct. First, wiring: in a stack, the upper layer consumes the layer below's output **at every time step**, so the lower layer must expose its full output sequence, not just its final state - feeding only the final state is a common bug that collapses the sequence and looks like depth not helping. Second, optimisation: check the third layer's gradients and activations are alive, since deeper recurrent stacks are harder to train and often need skip connections between layers or per-layer normalisation before they pay. Third, error analysis: if the residual errors are rare words the training set never contained, or ambiguities that need context the model cannot reach, capacity was never the constraint. Then weigh cost - each layer is another full sequential pass, so latency grows roughly linearly with depth. On tagging tasks two layers usually captures most of the available gain.
go deeper
Know that recurrent layers can be stacked, that the layer above reads the layer below's output at every time step, and that two layers is a common starting point for tagging tasks.
Explain the mechanics you would verify: the lower layer emitting an output sequence rather than a final state, gradients still reaching the bottom layer, and each added layer costing another sequential pass over the whole sequence.
Show a diagnosis order - wiring, then optimisation, then error analysis on sampled failures - and be able to say when the residual errors are data or annotation limits that no amount of capacity will fix.
Own the accuracy-versus-serving-cost tradeoff: decide in advance how much latency and cost a fraction of a point is worth, and be willing to ship the shallower model and redirect the effort to data.
## What stacking actually does Stacking recurrent layers means the sequence is processed by layer 1 to produce an output at every time step, and that whole output *sequence* becomes the input sequence of layer 2, which produces its own output sequence, and so on. Each layer has its own weights and its own hidden state, and each carries state along the time axis independently. The motivation is representational hierarchy along the vertical axis, analogous to what depth buys anywhere else: the lower layer learns local, surface regularities of the input, the upper layer operates on a stream of already-abstracted features. On a part-of-speech tagger you can see the effect empirically - going from one layer to two usually gives a clear gain, two to four typically gives very little, and the curve flattens. The single most important mechanical fact, and the one interviews probe: **the upper layer reads the lower layer at every step, not just at the end.** If you wire the lower layer to expose only its final state and feed that upward, you have destroyed the sequence - the upper layer now sees a length-one input and the whole stack degenerates. This mistake produces exactly the symptom in the question: adding depth changes nothing. ## A diagnosis order ### 1. Wiring Confirm the lower layer emits a full output sequence and the upper layer consumes it step-aligned. Confirm the directions are consistent - a bidirectional layer under a forward-only layer is legal but means the upper layer's causal-looking input is not causal at all. Confirm the layer was actually added: parameter counts that did not change are the fastest tell. ### 2. Optimisation A third layer that is present but not learning looks identical from the metric. Check that gradients reaching layer 1 have not collapsed, and that the top layer's activations have not saturated. Deep recurrent stacks are harder to optimise than shallow ones because gradients now travel a long path in two directions - backwards through time *and* down through layers. Standard remedies are skip connections that let a layer's input bypass it and be added to its output, and normalisation applied per layer. If the third layer only pays once you add skip connections, that is a genuine finding, not a workaround. ### 3. Error analysis This is what separates a senior answer. Sample the remaining errors and classify them: - **Rare or unseen tokens.** No amount of depth invents evidence that is not in the data. This calls for better input representations or more or broader data. - **Genuine ambiguity** that even a human annotator resolves inconsistently. Check annotator agreement; you may already be at the ceiling of the label quality. - **Long-range dependence** where the disambiguating evidence sits far away in the sequence. Depth does **not** extend how far back a recurrent model can reach - that reach is a property of the recurrence along the time axis, not of the number of layers stacked on top. If this is the failing class, more layers is the wrong knob entirely. - **Class imbalance in the tag set,** where the errors concentrate in a handful of rare tags and the aggregate metric hides it. If the errors are dominated by the first three, capacity is not your constraint and the third layer was never going to help. ## The cost side of the ledger Even when depth does buy a little accuracy, it is not free in a way that a feed-forward stack is not. - **Latency scales with depth.** Recurrence is already sequential over time; stacking adds another full pass over all T steps. Two layers is roughly two passes. Because the time axis cannot be parallelised inside a layer, you cannot amortise this away. - **Widening the hidden state is often cheaper than deepening it** for the same parameter budget, because a larger matrix multiply per step is work that hardware parallelises well, whereas an extra layer is more sequential work. - **Memory during training** grows with depth times sequence length, since every layer's states across all steps are retained for the backward pass. So the real decision is not "does layer three help?" but "does the fraction of a point it buys survive the serving latency budget and the per-request cost target?" Frequently the answer is no, and reporting that clearly is the right outcome. ## What a strong answer sounds like "Before adding a fourth layer I would confirm the stack is wired step-wise rather than through the final state, check that the third layer's gradients are alive and try skip connections between layers, then sample two hundred errors. If they are rare words and annotation noise, depth is not my problem and I would report that the tagger is data-limited at two layers - which also happens to be the cheapest model to serve."
- What does the second layer of a stacked recurrent network receive as its input?The lower layer's output at every time step, in order - so the lower layer must expose its whole output sequence, and the upper layer runs its own recurrence over that sequence with its own weights and its own hidden state. The classic wiring bug is feeding only the lower layer's final state upward, which collapses the sequence to a single step and makes the extra layer useless.
- Why does adding recurrent depth hurt inference latency more than widening the hidden state?Because depth adds sequential work and width adds parallel work. Each extra layer is another full pass over all time steps, and the time axis inside a layer is inherently serial, so passes cannot overlap. Widening the state makes each step's matrix multiplication larger, which parallel hardware absorbs with little wall-clock change. For the same parameter budget, width is usually the cheaper knob at serving time.
- If long-range dependencies are the failing cases, is more depth the right response?No. How far back a recurrent model can effectively reach is a property of the recurrence along the time axis - the state update and how gradients survive it - not of how many layers you stack above it. Extra layers give richer per-step features over the same reach. If the deciding evidence lies outside that reach, depth adds cost and leaves the failure untouched.
saying these in an interview costs you the question
- Adds layers as the default fix whenever a metric stalls
- Believes each extra layer extends the model's time horizon
- Wires the upper layer to the lower layer's final state only
- Ignores that depth multiplies inference latency per sequence
- Never samples the remaining errors before changing the model