skip to content

Backpropagation and Autodiff

You will derive backprop as reverse-mode autodiff over a computational graph and explain why gradients vanish, explode or get clipped. A whiteboard derivation is a classic senior DL screen.

on this pageshow

explore

questions

30

Why does a training forward pass keep its intermediate tensors when an inference-only forward pass can discard them?

level: juniorimportance: must knowfreq 68%

answer

  1. what does inference never run?
  2. no reverse sweep, nothing to replay
  3. the graph is a training-time record
  4. gradient rules read forward values
  5. an op keeps what its own backward reads

basics

~20 s

Training runs a backward pass, and each op's local gradient is a function of the tensors that op actually saw. Inference never runs backward, so an intermediate can be released the moment the next op has consumed it.

solid answer

~50 s

During training the forward pass has two jobs: produce the output, and leave behind a record the backward pass can replay. That record is the graph of ops plus, for each op, the specific tensor its own gradient rule reads back. Because those gradient rules are evaluated at the point the forward pass visited, the values have to still be there when the reverse sweep arrives. In inference-only mode there is no reverse sweep, so nothing needs to survive past its last consumer: an accelerometer anomaly detector scoring one window can overwrite each hidden activation as soon as the next layer has read it, and only the final score leaves the network. That is why the same architecture, the same weights and the same input behave differently in the two modes even though the arithmetic of the forward pass is identical.

go deeper

for a junior

Be ready to say plainly that training needs a backward pass and inference does not, and that the backward pass reads values the forward pass produced. Knowing that inference-only mode records no graph is the expected answer.

for a middle

Explain the mechanism: recording appends a record per op and holds a reference to the tensor that op's gradient rule reads, which is what keeps intermediates alive. Be able to say why a gradient rule is evaluated at the forward values.

for a senior

Show you know where this bites in practice — a validation or logging pass accidentally left recording on, or a handle kept on a graph-attached tensor past the end of a step. Distinguish freezing parameters from disabling recording without being prompted.

for a principal

Own the policy view: which code paths in a training system should record and which must not, how that is enforced rather than remembered, and what it costs the team when an evaluation path silently records. Treat mode discipline as an interface decision, not a per-call habit.

## Two different jobs for the same forward pass Running a network forward means applying a chain of ops to an input until you get an output. If that is *all* you want — a score, a class, a prediction — then every value produced along the way is scaffolding. Once layer 3 has read layer 2's output, layer 2's output is dead: nothing will ever look at it again, and the storage holding it can be reused immediately. Training is different, because training is a forward pass *followed by* a backward pass. The backward pass computes, for every parameter, the derivative of the loss with respect to that parameter, and it does so by walking the chain in reverse and applying each op's local gradient rule. The crucial point is that a local gradient rule is not a constant — it is a function evaluated **at the point the forward pass visited**. ## Local gradients are evaluated at the forward values Take `h = tanh(a)`. The derivative of `tanh` at `a` is `1 - tanh(a)^2`, which depends on where `a` was. If you throw away both `a` and `h`, the backward pass has no way to know which point on the curve to evaluate — the answer is different for `a = 0.1` and `a = 3.0`. The same is true of a linear layer: the gradient with respect to its weights is built from the layer's *input*, so that input has to still exist when the sweep reaches that layer. So the rule is not "training keeps everything because training is expensive". The rule is: **an op keeps whichever tensor its own backward step will read**, and it keeps it until the reverse sweep arrives at that op. Everything else can go. ## What "recording" actually means When recording is on, running an op does three things instead of one: 1. computes the output tensor; 2. appends a record to the graph saying which op ran, on which inputs, producing which output; 3. holds a reference to the tensor(s) that op's gradient rule will need, which keeps them alive. Step 3 is the reason intermediates stay reachable. Not because someone decided to hoard them, but because there is a live reference from the graph record to them. Ordinary garbage collection cannot reclaim a tensor that the graph is still pointing at. ## Inference-only mode When recording is turned off, step 2 and step 3 simply do not happen. There is no graph, no saved tensors, and consequently no backward pass is possible — asking for gradients afterwards is not "slow", it is undefined, because the information they would be computed from was never written down. Consider a small MLP that scores accelerometer windows for anomalies. In deployment it evaluates the same matrix multiplies and nonlinearities as during training, but each hidden activation is transient: produced, consumed by the next layer, released. Only the anomaly score survives the call. This asymmetry is worth internalising because it explains several everyday observations at once: - The same batch that runs fine at serving time may not run the same way while training, even with identical weights. - A forward pass done purely to *look at* something — logging an activation, computing a validation metric, generating a sample — has no reason to record, and recording it is pure waste. - Conversely, if you accidentally run a step with recording off and then ask for gradients, nothing is there. ## The distinction that trips people up Inference-only recording is not the same thing as freezing parameters. Freezing means a parameter will not be updated; the graph is still recorded, gradients still flow *through* the frozen layer to reach whatever sits behind it. Turning recording off is a statement about the graph itself: no record, no saved tensors, no backward. A network can be fully trainable and still be run in non-recording mode for a validation batch, and a network can have every parameter frozen while still recording, because something downstream needs the gradient path. ## Mental model Think of the forward pass under training as writing a receipt for every op: what ran, and the one value needed to reverse it. The backward pass reads the receipts in reverse order. Inference throws the receipts away as it goes, because nobody is going to ask for a refund.

  • If gradients are only needed for the parameters, why keep activations at all rather than just the weights?
    Because a parameter's gradient is not a function of the parameter alone. A linear layer's weight gradient is built from the tensor that entered that layer, so the activation flowing in is exactly what the rule reads. Weights persist anyway — they are long-lived state. It is the transient per-batch activations that the graph has to deliberately hold onto until the sweep reaches them.
  • Is running with recording disabled the same as freezing a layer's parameters?
    No. Freezing says a parameter will not be updated, but the graph is still built and gradients still pass through that layer to reach anything behind it. Disabling recording says no graph is built at all, so no backward pass is possible anywhere in that region. You can freeze every parameter and still need recording, and you can run a fully trainable network with recording off for a validation batch.
  • You want to log a hidden activation during training. Does that change what has to be retained?
    Reading the value for logging does not, by itself, change the graph — the tensor was already being held for the backward pass. What matters is whether you keep your own reference to it after the step finishes: holding a handle to an activation, or to anything that still points at the graph record, keeps that whole record alive past the point the sweep would have released it. Log the plain number, not the graph-attached tensor.

Cooking a dish for dinner, you rinse each bowl the moment it is empty. Cooking it while writing the recipe, you keep every measured quantity on the counter until you have written the step down.

saying these in an interview costs you the question

  • Says training keeps activations just because training is slower anyway
  • Thinks gradients can be recomputed from the weights alone
  • Believes the backward pass re-runs the forward pass from the input
  • Confuses disabling recording with freezing parameters
  • Claims inference and training forward passes differ in their arithmetic

context

open as a page

How does an elementwise max such as ReLU route gradient in backpropagation?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A max passes the incoming gradient straight to whichever input won and gives exactly zero to the loser. For ReLU that is a 0/1 mask: positive inputs pass the upstream gradient through unchanged, non-positive ones cut it to zero.

open as a page

In the backward pass, how does a max-pooling layer route gradients compared with average pooling?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Max pooling sends a window's entire upstream gradient to the one input that was the maximum and exactly zero to the others. Average pooling spreads it evenly, so each input in a k-by-k window receives one k-squared-th of it.

open as a page

In reverse-mode autodiff, why is a reused tensor's gradient the sum of its consumers' gradients?

level: middleimportance: must knowfreq 62%

basics

~20 s

A reused tensor appears in several terms of the chain rule, so its gradient is the sum of one contribution per consumer. The reverse sweep adds each contribution into a single buffer as it reaches that consumer.

open as a page

What is the computational graph a forward pass records, and in what order does the backward sweep replay it?

level: middleimportance: must knowfreq 72%

basics

~20 s

The forward pass records a directed acyclic graph whose nodes are ops and whose edges are the tensors between them. The backward sweep replays it in reverse topological order: a node is visited only after every op that consumed its output.

open as a page

How does a central finite-difference check verify a hand-derived analytic gradient?

level: middleimportance: must knowfreq 55%

basics

~20 s

Perturb one parameter by plus and minus a small h, recompute the scalar loss each time, and compare (L_plus - L_minus)/(2h) with the analytic gradient for that coordinate. Score the agreement by relative error, not by raw difference.

open as a page

How does clipping gradients by global norm differ from clipping each gradient element by value?

level: middleimportance: must knowfreq 62%

basics

~20 s

Global-norm clipping rescales the whole gradient by a single factor once its norm passes a threshold, so the update direction is unchanged. Elementwise value clipping clamps each component on its own, which rotates the update toward the diagonal.

open as a page

In a 3-layer MLP, how do you derive dL/dW2 and the gradient handed back to the hidden layer?

level: middleimportance: must knowfreq 72%

basics

~10 s

Start at layer 2's pre-activation with delta2 = dL/dz2. Then dL/dW2 is the outer product delta2 h1^T, dL/db2 is delta2 itself, and the hidden layer receives W2^T delta2, which you multiply elementwise by f'(z1).

open as a page

In backprop through a matrix product C = A B, what are the gradients with respect to A and B?

level: middleimportance: must knowfreq 70%

basics

~20 s

With G as the upstream gradient of the loss with respect to C, dL/dA = G B^T and dL/dB = A^T G. Matmul has two separate local rules, one per operand, and each carries exactly one transpose.

open as a page

In a convolution layer, what are the gradients with respect to its input and its kernel weights?

level: middleimportance: must knowfreq 62%

basics

~20 s

A convolution's input gradient is itself a convolution: the upstream gradient against the kernel rotated 180 degrees. The kernel gradient correlates the layer input with the upstream gradient, summed over every position the kernel visited.

open as a page

Why does training a 100-million-parameter network use one reverse sweep instead of one forward sweep per parameter?

level: middleimportance: must knowfreq 68%

basics

~20 s

Reverse mode costs one sweep per output; forward mode costs one per input. Training has one scalar loss and a hundred million inputs, so a single reverse sweep gets every gradient, while forward mode would need a hundred million.

open as a page

Why can gradients vanish or explode as they backpropagate through 50 layers?

level: middleimportance: must knowfreq 82%

basics

~20 s

Backpropagation multiplies one Jacobian per layer, so the gradient reaching layer 1 is a product of 50 factors. If those factors average below one the gradient decays geometrically with depth; above one it grows the same way.

open as a page

With a batch of 64 records, what shape is a 128-unit layer's bias gradient, and why?

level: juniorimportance: should knowfreq 58%

basics

~20 s

The bias gradient is a 128-length vector, the same shape as the bias itself. All 64 records contribute their own delta to the same shared bias, so the batch axis is summed away rather than kept.

open as a page

How does automatic differentiation differ from symbolic and numerical differentiation?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Automatic differentiation applies exact per-operation derivative rules to numeric values as the program runs, giving machine-precision derivatives at one input point. Symbolic differentiation manipulates formulas and can blow up in size; numerical differentiation perturbs inputs and carries approximation error.

open as a page

What does a training loss that spikes and then becomes NaN say about gradients?

level: juniorimportance: should knowfreq 60%

basics

~20 s

It is the classic signature of exploding gradients: one update was large enough to push weights into a range where activations overflow to infinity, and infinity arithmetic then produces NaN. Once NaN reaches the weights, the run never recovers.

open as a page

Why does shrinking h in a finite-difference gradient check eventually make it worse?

level: middleimportance: should knowfreq 43%

basics

~20 s

Two errors pull in opposite directions. A central difference's truncation error falls like h squared, but subtracting two nearly identical losses and dividing by 2h amplifies round-off like 1/h. In double precision the total is smallest near h = 1e-5.

open as a page

Why does a bias vector broadcast across a batch get a gradient summed over the batch axis?

level: middleimportance: should knowfreq 58%

basics

~20 s

Broadcasting copies the same bias into every row of the batch, so every row contributes a separate term to the same parameter. The backward of a copy is a sum, so those contributions add up into one bias-shaped gradient.

open as a page

When one embedding id appears many times in a batch, what gradient does its row receive?

level: middleimportance: should knowfreq 44%

basics

~20 s

The sum of the upstream gradient vectors from every occurrence. The lookup gathers rows, so the backward scatter-adds into them. Overwriting instead of accumulating would apply only the last occurrence and silently discard the rest.

open as a page

Why does a hard argmax in the forward pass hand back a zero gradient to everything upstream?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A hard argmax is flat between its jumps, so its derivative is zero almost everywhere and undefined at ties. Reverse mode multiplies by that zero, and any parameter whose only route to the loss crosses it gets an exactly-zero gradient.

open as a page

For a linear -> tanh -> linear chain, which tensor must each op stash for its own backward step?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Each op stashes only what its own backward step reads: a linear op keeps its input, tanh keeps its own output, ReLU needs nothing more than the sign of its input, and a plain addition keeps no tensor at all.

open as a page

Why does a gradient check fail on ReLU units whose pre-activation sits near zero?

level: seniorimportance: should knowfreq 34%

basics

~20 s

The loss is piecewise linear in that parameter with a corner between the two probe points, so the numerical secant blends both slopes while the backward pass reports one. The check fails though the code is correct.

open as a page

How do you pick a gradient-clipping threshold instead of inheriting a default of 1.0?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Measure before you choose. Record the global gradient norm for the first several hundred steps with the clip effectively off, look at the distribution, and set the threshold near its upper tail — around the 90th percentile — so ordinary steps pass and only outliers get rescaled.

open as a page

Which hand-derived backprop mistakes survive a shape check and still break training?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Shape checks only prove conformability. A transpose flipped between equal dimensions, an activation derivative dropped or taken at the wrong layer, and a sum where the loss averaged all leave every shape valid while corrupting the gradient.

open as a page

How does switching a loss from a batch mean to a plain sum change every gradient in the graph?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Every gradient is multiplied by the batch size N. A mean reduction hands each element a local gradient of 1/N; a plain sum hands it 1. The descent direction is identical, the step is N times larger.

open as a page

A 40-layer network has gradient norms of 1e-2 at the head and 1e-9 at the input — what do you conclude?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Seven orders of magnitude across the depth means the gradient is decaying geometrically layer by layer. Only the top few layers learn; the rest still hold their initial weights, acting as a fixed random feature map. Effective depth is about four.

open as a page

When is gradient clipping the wrong fix for a run whose gradients keep exploding?

level: principalimportance: should knowfreq 37%

basics

~20 s

Whenever the huge gradient has a findable cause. Clipping bounds the step but leaves the cause running, and because the run now survives, nobody investigates. It is a seatbelt for rare tail events, not a substitute for finding what produces them.

open as a page

Why does BatchNorm's backward pass make one sample's input gradient depend on the whole batch?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Because the batch mean and variance are functions of every sample, not constants. Differentiating through them puts batch-wide sums into each sample's gradient, so the same input gets a different gradient depending on its batch-mates.

open as a page

Which automatic differentiation mode builds the Jacobian of thousands of residuals against 6 robot-arm parameters in fewer sweeps?

level: seniorimportance: nice to knowfreq 30%

basics

~10 s

Forward mode. One forward sweep produces one Jacobian column, so six sweeps cover the six parameters. Reverse mode produces one row per sweep and would need thousands, one per residual.

open as a page

Where do you place stop-gradients in a model's graph, and what does each cut change?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

A stop-gradient passes its input forward unchanged and returns zero gradient, so the optimizer treats that value as a constant. Put one wherever a model-computed quantity should act as data, not as something the loss can lower by changing it.

open as a page

When does a finite-difference gradient check earn a permanent place in a test suite?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Wherever a derivative was written by a human rather than composed from already-trusted primitives: a custom layer, a custom loss, a hand-optimised operation. Keep the test tiny, deterministic and in double precision so it runs in seconds and never flakes.

open as a page