In one LSTM step, what does each of the forget, input and output gates compute?
answer
- three gates, one candidate
- sigmoid gates, tanh candidate
- keep term plus write term
- products are elementwise, per dimension
- c_t = f*c_(t-1) + i*g, h_t = o*tanh(c_t)
basics
~20 sThe forget gate scales the previous cell state, the input gate scales a tanh candidate, and their sum is the new cell state: c_t = fc_(t-1) + ig. The output gate scales tanh(c_t) into the emitted hidden state.
solid answer
~50 sAt each step the cell reads the current input `x_t` and the previous hidden state `h_(t-1)` and computes four blocks from them: three sigmoid gates and one tanh candidate. The forget gate `f_t` decides, per dimension, what fraction of the old cell state to keep; the input gate `i_t` decides how much of the candidate `g_t` to write; the new memory is the sum `c_t = f_t * c_(t-1) + i_t * g_t`, every product elementwise. The output gate `o_t` then decides what to expose: `h_t = o_t * tanh(c_t)`. So two vectors go forward, not one — `c_t`, the private memory that is only ever scaled and added to, and `h_t`, the filtered view that leaves the cell and feeds the next step's gates. Gates are sigmoids because a value in (0, 1) works as a soft, differentiable switch; the candidate is tanh so a write can push a dimension up or down.
code
python · 19 linesimport math
def sigmoid(z):
return 1 / (1 + math.exp(-z))
# One dimension of a gated cell, one step.
z_f, z_i, z_o, z_g = 2.5, -1.5, 0.4, 0.9 # the four pre-activations
c_prev = 3.0 # a value already stored here
f = sigmoid(z_f) # forget gate: fraction of c_prev kept
i = sigmoid(z_i) # input gate: fraction of the candidate written
o = sigmoid(z_o) # output gate: fraction of memory exposed
g = math.tanh(z_g) # candidate: signed value to write
c = f * c_prev + i * g # additive cell-state update
h = o * math.tanh(c) # emitted hidden state
print(round(f, 3), round(i, 3), round(c, 3), round(h, 3))
# 0.924 0.182 2.903 0.595go deeper
Be able to name the three gates and say the cell keeps a memory vector that is scaled by the forget gate and added to by the input gate. Knowing which gate does which is the bar here.
Write both update lines from memory, say the products are elementwise, and explain why gates are sigmoids while the candidate is a tanh. Expect to be asked what leaves the cell versus what stays inside.
Talk about the regimes: what forget near 1 with input near 0 means in a trained model, when a dimension is acting as a latch versus an accumulator, and what you would look at in a gate trace to tell those apart.
Own the argument for whether gating is worth its cost on a given problem at all — four blocks per step and a strictly sequential unroll — versus a simpler recurrence or an architecture that reads the whole sequence at once.
## The shape of one step A gated recurrent cell keeps a state of width `d_h` and consumes inputs of width `d_in`. It carries **two** vectors from step to step: - `c_t` — the **cell state**, an internal memory that never passes through a weight matrix on its way forward. It is only scaled and added to. - `h_t` — the **hidden state**, the vector the cell actually emits: to the layer above, to the loss, and back into its own gates at the next step. Everything the cell decides at step `t` is a function of just `x_t` and `h_(t-1)`. ## The four blocks From `x_t` and `h_(t-1)` the cell computes four affine blocks, each with its own input matrix, recurrent matrix and bias: ``` f_t = sigmoid(W_f x_t + U_f h_(t-1) + b_f) forget gate i_t = sigmoid(W_i x_t + U_i h_(t-1) + b_i) input gate o_t = sigmoid(W_o x_t + U_o h_(t-1) + b_o) output gate g_t = tanh( W_g x_t + U_g h_(t-1) + b_g) candidate ``` Three of these are **gates** and one is a **candidate**, and the difference matters. A gate is squashed into (0, 1), so multiplying by it can only attenuate: 0 means block completely, 1 means pass completely, and anything in between is a partial, differentiable open-ness. The candidate is squashed into (-1, 1) instead, because a write must be able to push a memory dimension up or down, not only up. ## The update ``` c_t = f_t * c_(t-1) + i_t * g_t h_t = o_t * tanh(c_t) ``` All products are **elementwise**, which is the part candidates most often miss. There is no single scalar deciding to remember or forget; each of the `d_h` dimensions has its own forget value, its own input value and its own output value at every step. One dimension can sit at `f = 0.99, i = 0.01` and hold a flag steady for hundreds of steps while the dimension next to it is overwritten every step. Read the two terms as an answer to two independent questions. *How much of what I already knew survives?* — that is `f_t * c_(t-1)`. *How much of what I just saw gets written in?* — that is `i_t * g_t`. Because they are separate gates, the cell can keep and write at the same time, or keep without writing, or overwrite by opening the input gate while closing the forget gate. ## Why the output gate exists If the emitted state were just the memory, the cell would have no way to hold something it does not currently need to expose. The output gate separates *storing* from *reporting*: a dimension can carry a fact for a long stretch with the output gate closed, and open it only on the steps where downstream layers need it. The `tanh` around `c_t` bounds what is emitted, because the cell state is an accumulator and can drift well outside (-1, 1) while the emitted hidden state stays bounded. ## The regimes worth naming - **Carry unchanged**: `f` near 1, `i` near 0. The dimension is a latch; its value survives. - **Overwrite**: `f` near 0, `i` near 1. The dimension becomes the candidate; history is discarded. - **Accumulate**: `f` near 1, `i` moderate. The dimension integrates over time, which is how a cell builds a running count or a running average. - **Erase**: `f` near 0, `i` near 0. The dimension is zeroed and stays empty until something writes to it. ## Common misreadings The forget gate does not decide what to *output*, and the output gate does not touch memory — `o_t` appears nowhere in the `c_t` equation. The gates are not a softmax and do not sum to one across the three of them; they are three independent sigmoids and can all be near 1 at once. And nothing here is a hard switch: at training time every gate is a smooth function, which is exactly what makes the whole cell differentiable end to end. ## Saying it in an interview Write the two update lines, say the products are elementwise and per-dimension, name which gate multiplies which term, and finish with the distinction between the private cell state and the emitted hidden state. That answer, in that order, covers almost everything an interviewer is checking.
- Why is the candidate a tanh rather than a sigmoid like the gates?Because a candidate is a value, not a valve. A gate multiplies, so it only needs to live in (0, 1) — zero blocks, one passes. A write has to be able to move a memory dimension in either direction, so it needs a signed range; tanh gives (-1, 1). A sigmoid candidate could only ever push cell values upward, which would make a dimension monotonically drift.
- Are the three gates competing, so that opening one closes another?No. They are three independent sigmoids computed from the same `x_t` and `h_(t-1)` but with separate weights, so they can all be near 1 at once. Nothing normalises them against each other. Keeping and writing are genuinely separate decisions: a dimension can hold its old value and add to it in the same step, which is how a cell accumulates.
- What actually leaves the cell and goes into the layer above?Only `h_t`. The cell state `c_t` is passed sideways to the next timestep of the same cell and is never handed to the next layer or to the loss directly. That is why `h_t = o_t * tanh(c_t)` matters: a fact can sit in memory for hundreds of steps with the output gate mostly closed, and only surface on the step where it is needed.
The cell state is a row of dimmer switches on a shelf of stored values: the forget gate dims what is already there, the input gate dims how much new value gets added in, and the output gate dims how much of the shelf is visible from outside.
saying these in an interview costs you the question
- Says the output gate decides what stays in memory
- Treats the three gates as a softmax that sums to one
- Describes gating as one scalar rather than per-dimension
- Confuses the tanh candidate with the input gate
- Thinks the cell state is what feeds the next layer