Narrate one training step for a batch of 64 — which tensors change and which do not?
answer
- four phases, always in order
- one scalar, then one update
- data is read, never written
- gradients add unless you clear them
- logged loss predates the update
basics
~20 sThe forward pass turns the 64 inputs into activations and one loss number, the backward pass fills a gradient for every parameter, and the update writes new parameter values. Only parameters, optimizer state and gradient buffers change; the batch, targets and architecture do not.
solid answer
~50 sFour things happen in order. **Forward**: the batch of 64 inputs flows through the current parameters, producing activations at every layer and finally predictions. **Loss**: the predictions and the 64 targets are reduced to a single scalar, normally the mean of the per-example losses. **Backward**: a gradient of that scalar is computed with respect to every parameter and written into a gradient buffer. **Update**: the optimizer reads those gradients and writes new parameter values. What changed: the parameters, the optimizer's internal state, the gradient buffers, and the step counter. What did not: the input batch and its targets, the architecture and every tensor shape, the loss function, and the parameter count. Activations are transient scratch space. Crucially, the loss you logged was computed from the parameters *before* the update — it describes the model that no longer exists.
go deeper
Memorise the order: forward, loss, backward, update. Know that only the update changes the weights and that the data is read-only.
Be able to name every piece of state a step touches and every piece it leaves alone, and explain why gradient buffers must be cleared between steps.
Show the diagnostic instinct: when a run looks stuck, check whether a parameter actually moved between two steps before touching any hyperparameter.
Own the loop contract for the team — where state is mutated, what is checkpointed, and what invariants a smoke test asserts so that a silent no-op loop can never reach a long run.
## The four phases A single training step on a batch of 64 examples is always these four phases in this order. **1. Forward pass.** The 64 inputs are pushed through the network using the parameter values as they currently stand. Each layer produces an *activation* — an intermediate tensor whose first dimension is 64, one row per example — and the final layer produces 64 predictions. Nothing is learned here; the forward pass is pure evaluation of the current function on this batch. **2. Loss computation.** The 64 predictions and the 64 targets are compared, producing 64 per-example losses, which are then *reduced* to a single scalar — normally the mean. The reduction matters: it is why the size of the batch does not by itself change the magnitude of the loss, and it is what makes "the gradient" a single well-defined object rather than 64 of them. **3. Backward pass.** For that one scalar, a partial derivative is produced with respect to every trainable parameter and written into that parameter's gradient buffer. The forward pass's stored activations are consumed here — that is why they were kept. **4. Update.** The optimizer reads the gradient buffers and writes new values into the parameters. This is the only phase that changes what the model computes. ## What changes, precisely - **Parameters (weights and biases).** New values, in place. This is the whole point of the step. - **Optimizer internal state.** Optimizers that carry running quantities across steps update them on every step. Even a stateless update advances a step counter. - **Gradient buffers.** Refilled with this batch's gradients — but read the next section, because *how* they are refilled is a classic bug. - **Any buffer a layer maintains from the data it sees.** Some layers keep stored statistics that the forward pass refreshes. These are not trainable parameters and no gradient touches them, but they are part of the model's state and they do move during a training step. - **Counters and logs.** The step number, the running loss average, the wall clock. ## What does not change - **The input batch and its targets.** The update writes to parameters, never back into the data. If your data changed, something is wrong. - **The architecture.** Layer count, layer types, every tensor shape and the total parameter count are fixed for the whole run. A step changes the *values* in a fixed-shape container. - **The loss function** and the reduction it uses. - **The activations, in any lasting sense.** They exist only between the forward and backward of this one step and are recomputed from scratch for the next batch. They are not state. ## The gradient-buffer trap In reverse-mode systems, the backward pass *adds into* each parameter's gradient buffer rather than overwriting it. A correct loop therefore clears the buffers before (or immediately after) each backward pass. Forget it, and step *k* applies the sum of the gradients from batches 1 through *k*: an effectively enormous and increasingly stale update direction dominated by data the model has already moved past. The symptom is a loss that behaves as if the learning rate were far too high and grows worse the longer the epoch runs. It is a loop-structure bug, not a hyperparameter problem, and it will not be fixed by turning any dial down. ## The mirror-image trap: no update at all The opposite defect is a loop that computes the loss, logs it, and never applies the update inside the batch loop — the update statement sits one indentation level out, so it fires once per epoch instead of once per batch, or is missing entirely. The loss trace then looks *almost* plausible: it jitters batch to batch because different batches have genuinely different difficulty, and it does not fall. Candidates misread this as a rate problem for hours. The decisive check takes seconds: read one parameter value before and after a step and see whether it moved. If it did not, no dial will help. ## The subtlety worth saying out loud The loss you log for step *k* was produced in the forward pass, using the parameters *as they were before* step *k*'s update. It is a measurement of the previous model, not of the one you now hold. This is harmless for monitoring but it explains why the logged training loss lags the model's true current quality by one step, and why a per-step training loss is a running measurement of a model that is changing underneath it rather than a clean evaluation of any single parameter set. Being able to narrate all of this crisply — forward, loss, backward, update; parameters and optimizer state change, data and shapes do not — is one of the most reliable middle-tier signals an interviewer has that you have actually written a training loop rather than only configured one.
- Why must the gradient buffers be cleared between steps?Because the backward pass adds into them rather than overwriting them. Without clearing, step k applies the sum of every earlier batch's gradients — a huge, stale direction that gets worse as the epoch goes on. It looks exactly like a learning rate set far too high, but it is a loop-structure bug and no hyperparameter change fixes it.
- A loop logs a loss every batch but the update statement sits outside the batch loop. What do you see?A loss trace that jitters with batch difficulty and never trends down, because the model receives one update per epoch instead of one per batch. It reads like a stuck run. Read a single parameter value before and after a step: if it is unchanged, the update is not firing where you think it is.
- Does the loss logged for a step reflect the parameters after that step?No. It was computed in the forward pass from the pre-update parameters, so it measures the model as it stood at the start of the step. The updated model's loss on that same batch is never computed unless you deliberately run a second forward pass.
- Are activations part of the model's state?No. They are scratch tensors produced by the forward pass for this batch, kept only long enough for the backward pass to consume them, then discarded. The next batch recomputes its own. Only parameters, optimizer state and layer buffers persist across steps.
saying these in an interview costs you the question
- Thinks the backward pass changes the weights
- Believes gradient buffers reset themselves each step
- Says the logged loss reflects the updated parameters
- Thinks the update writes back into the input batch
- Cannot separate parameters from activations
- Calls a missing update a learning-rate problem