An LSTM loses a patient's ventilation flag long before step 400 of a vitals stream — how do you diagnose it?
answer
- measure before you change anything
- plot the gate activations over the sequence
- forget near 1, input near 0 is the latch
- 0.9 to the 400th is effectively zero
- start the forget bias positive
basics
~20 sTrace forget-gate activations per unit across the sequence: a held fact shows forget near 1 with input near 0. If no unit holds high, memory is decaying — initialise the forget-gate bias positive and check the backward pass is not truncated.
solid answer
~50 sStart by proving where the flag dies. Record the gate activations for a sequence where the flag is set early and needed late, and plot the forget gate per unit over time; a unit acting as a latch shows forget near 1 with the input gate near 0. If nothing sits high, the memory is decaying by construction — a unit at forget 0.9 retains about 5e-19 of its value after 400 steps, and even 0.99 leaves under two percent. Two fixes address the two usual causes. Initialise the forget-gate bias to a positive value such as 1, so every unit starts around 0.73 open and leans toward keeping before it has learned anything. Then check the training setup: if the backward pass is truncated shorter than the gap, or state is reset at chunk boundaries, the model has never seen a gradient linking the two events.
go deeper
Know that memory in these cells fades when the forget gate is not close to one, and that repeated multiplication by a number just under one collapses fast over hundreds of steps. Recognising the symptom is enough here.
Be ready to say what you would record and plot — the three gate vectors across a full sequence — and to explain what the carry regime looks like in that trace as opposed to steady decay.
Show the full loop: measure, quote the decay arithmetic, separate architecture causes from pipeline causes such as a truncated backward pass or a state reset, and apply the forget-bias initialisation as a starting-point fix rather than a cure.
Own the call about where the fact should live at all. Argue when a known, persistent attribute should be supplied as an input feature or the timestep coarsened, and reserve learned memory for what genuinely must be inferred from the stream.
## Frame the failure as a measurement problem A report that a model *forgets* something is a hypothesis, not a diagnosis. In a vitals stream sampled every five minutes, holding a flag such as *this patient is ventilated* across 400 readings means the fact must survive over thirty hours of input while the emitted state changes at every step. Before touching the architecture, establish which of three things is actually true: the memory is decaying, the gradient never linked the two events, or the model was never asked to. ## Step one: read the gates Instrument a forward pass on a sequence where the flag is set early and matters late, and record all three gate vectors at every step. Two patterns are worth looking for. **The carry regime.** A unit acting as a latch sits with its forget gate near 1 and its input gate near 0 for the whole stretch. Its cell value is flat while the hidden state around it churns, which is exactly the separation the output gate exists to provide. If one or two units look like this and the flag still gets lost, the memory is fine and the problem lies downstream in how that unit is read. **Collapse.** If no unit holds a high forget value over the stretch, the memory is leaking, and the arithmetic is brutal. A forget value of 0.9 retains `0.9^400`, roughly 5e-19 — the stored value is gone within a few dozen steps. Even 0.99 retains about 0.018 after 400 steps. Only values very close to one behave like a latch on this horizon, and the gap between 0.9 and 0.999 is invisible to the eye on a plot but is the whole difference between forgetting and remembering. A related pattern is a forget gate that starts high and drops sharply at a particular kind of event — a unit that has learned to clear itself on a signal that happens to co-occur with something irrelevant. That shows up as a cliff in the trace rather than a slow decay, and it points at the training data rather than the architecture. ## Step two: check the training setup, not the cell The most common real cause is not the architecture at all. **Truncated backward passes.** Long sequences are often trained by cutting the backward pass into windows. If the window is 50 steps and the dependency spans 400, the gradient linking the flag to its use never exists; the model is not failing to learn the dependency, it has never been shown it. Lengthen the window, or carry state forward across windows so the forward pass at least sees the history. **State resets.** If the state is zeroed at every chunk boundary, or sequences are shuffled and re-batched in a way that breaks continuity, the fact is erased by the data pipeline rather than by any gate. **The label does not need it.** If the target at late steps is predictable from local vitals alone, there is no gradient pressure to keep the flag at all, and the model is behaving correctly. Verify the dependency is real before demanding the model represent it. ## Step three: bias the starting point toward remembering With random initialisation every forget gate starts near the middle of its range, so a freshly initialised cell erases most of its memory at every step and long-range gradients are dead before training can teach it otherwise. The standard remedy is to initialise the forget-gate bias to a positive constant such as 1, putting the gate around 0.73 open at the start. The cell begins in a keep-leaning regime, gradients reach further back in the earliest updates, and the model can *learn* to forget from there — which is the easy direction — rather than having to discover remembering from a starting point that destroys the evidence for it. This is initialisation, not a constraint: nothing stops the model from learning small forget values where forgetting is right. ## Step four: consider changing the problem Senior judgment here often means not solving it in the cell. If a fact is known, categorical and persistent, carrying it in learned memory is an odd choice: feed it as an input feature repeated at every step, and the network no longer has to spend a unit and thirty hours of gate discipline preserving something you already know. Coarsening the timestep so the dependency spans forty steps rather than four hundred is the same move from the other side. Reserve learned memory for facts the model must infer. ## What a strong answer sounds like Measure first, and say what you would measure. Quote the decay arithmetic to show why *near* one is not one. Separate architecture causes from pipeline causes, and name the truncation check explicitly. Offer the forget-bias initialisation as a starting-point fix rather than a cure. Finish with the engineering option of supplying the fact directly. That sequence demonstrates having actually debugged one of these rather than having read about them.
- Why initialise the forget-gate bias to a positive value rather than zero?At zero the gate starts around half open, so a freshly initialised cell loses roughly half of every stored value at every step and long-range gradient is dead before training can shape the gates. A positive bias such as 1 starts each unit around 0.73 open, in a keep-leaning regime, so early gradients reach far back. Learning to forget from there is easy; discovering remembering from a memory-destroying start is not.
- The gate trace shows one unit with a forget value pinned near 1 for the whole sequence, yet predictions still ignore the flag. What now?Memory is not the problem — storage is working. Look at the read side: the output gate on that unit may be closed at the steps where the flag matters, so the fact never reaches the layer above, or the head above may simply not be using that dimension. Check the emitted values at the decision steps, and check whether the loss actually rewards using the flag.
- How would you tell a decaying memory apart from a memory that is being deliberately cleared?Look at the shape of the cell value over time. Decay is smooth and geometric — the value slides toward zero at a steady rate with the forget gate sitting at some middling constant. Deliberate clearing is a cliff: the forget gate drops sharply at one step, usually triggered by a specific input pattern. Smooth decay points at initialisation and training length; a cliff points at what is in the data at that step.
- When would you not fix this in the model at all?When the fact is already known outside the model. A ventilation status is recorded, categorical and persistent, so feeding it as an input feature at every step removes the need for any unit to preserve it across thirty hours of readings. Learned memory should be spent on things that must be inferred from the stream, not on relaying a value you could simply supply.
saying these in an interview costs you the question
- Jumps to a bigger hidden size without measuring anything
- Says a forget value of 0.9 is close enough to one
- Never checks whether the backward pass was truncated
- Treats forget-bias initialisation as a hard constraint
- Ignores that the label may not depend on the far-back fact