Why does a training forward pass keep its intermediate tensors when an inference-only forward pass can discard them?
answer
- what does inference never run?
- no reverse sweep, nothing to replay
- the graph is a training-time record
- gradient rules read forward values
- an op keeps what its own backward reads
basics
~20 sTraining runs a backward pass, and each op's local gradient is a function of the tensors that op actually saw. Inference never runs backward, so an intermediate can be released the moment the next op has consumed it.
solid answer
~50 sDuring training the forward pass has two jobs: produce the output, and leave behind a record the backward pass can replay. That record is the graph of ops plus, for each op, the specific tensor its own gradient rule reads back. Because those gradient rules are evaluated at the point the forward pass visited, the values have to still be there when the reverse sweep arrives. In inference-only mode there is no reverse sweep, so nothing needs to survive past its last consumer: an accelerometer anomaly detector scoring one window can overwrite each hidden activation as soon as the next layer has read it, and only the final score leaves the network. That is why the same architecture, the same weights and the same input behave differently in the two modes even though the arithmetic of the forward pass is identical.
go deeper
Be ready to say plainly that training needs a backward pass and inference does not, and that the backward pass reads values the forward pass produced. Knowing that inference-only mode records no graph is the expected answer.
Explain the mechanism: recording appends a record per op and holds a reference to the tensor that op's gradient rule reads, which is what keeps intermediates alive. Be able to say why a gradient rule is evaluated at the forward values.
Show you know where this bites in practice — a validation or logging pass accidentally left recording on, or a handle kept on a graph-attached tensor past the end of a step. Distinguish freezing parameters from disabling recording without being prompted.
Own the policy view: which code paths in a training system should record and which must not, how that is enforced rather than remembered, and what it costs the team when an evaluation path silently records. Treat mode discipline as an interface decision, not a per-call habit.
## Two different jobs for the same forward pass Running a network forward means applying a chain of ops to an input until you get an output. If that is *all* you want — a score, a class, a prediction — then every value produced along the way is scaffolding. Once layer 3 has read layer 2's output, layer 2's output is dead: nothing will ever look at it again, and the storage holding it can be reused immediately. Training is different, because training is a forward pass *followed by* a backward pass. The backward pass computes, for every parameter, the derivative of the loss with respect to that parameter, and it does so by walking the chain in reverse and applying each op's local gradient rule. The crucial point is that a local gradient rule is not a constant — it is a function evaluated **at the point the forward pass visited**. ## Local gradients are evaluated at the forward values Take `h = tanh(a)`. The derivative of `tanh` at `a` is `1 - tanh(a)^2`, which depends on where `a` was. If you throw away both `a` and `h`, the backward pass has no way to know which point on the curve to evaluate — the answer is different for `a = 0.1` and `a = 3.0`. The same is true of a linear layer: the gradient with respect to its weights is built from the layer's *input*, so that input has to still exist when the sweep reaches that layer. So the rule is not "training keeps everything because training is expensive". The rule is: **an op keeps whichever tensor its own backward step will read**, and it keeps it until the reverse sweep arrives at that op. Everything else can go. ## What "recording" actually means When recording is on, running an op does three things instead of one: 1. computes the output tensor; 2. appends a record to the graph saying which op ran, on which inputs, producing which output; 3. holds a reference to the tensor(s) that op's gradient rule will need, which keeps them alive. Step 3 is the reason intermediates stay reachable. Not because someone decided to hoard them, but because there is a live reference from the graph record to them. Ordinary garbage collection cannot reclaim a tensor that the graph is still pointing at. ## Inference-only mode When recording is turned off, step 2 and step 3 simply do not happen. There is no graph, no saved tensors, and consequently no backward pass is possible — asking for gradients afterwards is not "slow", it is undefined, because the information they would be computed from was never written down. Consider a small MLP that scores accelerometer windows for anomalies. In deployment it evaluates the same matrix multiplies and nonlinearities as during training, but each hidden activation is transient: produced, consumed by the next layer, released. Only the anomaly score survives the call. This asymmetry is worth internalising because it explains several everyday observations at once: - The same batch that runs fine at serving time may not run the same way while training, even with identical weights. - A forward pass done purely to *look at* something — logging an activation, computing a validation metric, generating a sample — has no reason to record, and recording it is pure waste. - Conversely, if you accidentally run a step with recording off and then ask for gradients, nothing is there. ## The distinction that trips people up Inference-only recording is not the same thing as freezing parameters. Freezing means a parameter will not be updated; the graph is still recorded, gradients still flow *through* the frozen layer to reach whatever sits behind it. Turning recording off is a statement about the graph itself: no record, no saved tensors, no backward. A network can be fully trainable and still be run in non-recording mode for a validation batch, and a network can have every parameter frozen while still recording, because something downstream needs the gradient path. ## Mental model Think of the forward pass under training as writing a receipt for every op: what ran, and the one value needed to reverse it. The backward pass reads the receipts in reverse order. Inference throws the receipts away as it goes, because nobody is going to ask for a refund.
- If gradients are only needed for the parameters, why keep activations at all rather than just the weights?Because a parameter's gradient is not a function of the parameter alone. A linear layer's weight gradient is built from the tensor that entered that layer, so the activation flowing in is exactly what the rule reads. Weights persist anyway — they are long-lived state. It is the transient per-batch activations that the graph has to deliberately hold onto until the sweep reaches them.
- Is running with recording disabled the same as freezing a layer's parameters?No. Freezing says a parameter will not be updated, but the graph is still built and gradients still pass through that layer to reach anything behind it. Disabling recording says no graph is built at all, so no backward pass is possible anywhere in that region. You can freeze every parameter and still need recording, and you can run a fully trainable network with recording off for a validation batch.
- You want to log a hidden activation during training. Does that change what has to be retained?Reading the value for logging does not, by itself, change the graph — the tensor was already being held for the backward pass. What matters is whether you keep your own reference to it after the step finishes: holding a handle to an activation, or to anything that still points at the graph record, keeps that whole record alive past the point the sweep would have released it. Log the plain number, not the graph-attached tensor.
Cooking a dish for dinner, you rinse each bowl the moment it is empty. Cooking it while writing the recipe, you keep every measured quantity on the counter until you have written the step down.
saying these in an interview costs you the question
- Says training keeps activations just because training is slower anyway
- Thinks gradients can be recomputed from the weights alone
- Believes the backward pass re-runs the forward pass from the input
- Confuses disabling recording with freezing parameters
- Claims inference and training forward passes differ in their arithmetic