skip to content

Computational Graphs

The forward pass builds a directed graph of ops over tensors and records what the reverse sweep needs, which is what lets you derive a small network's gradients by hand on a whiteboard.

on this pageshow

explore

questions

14

Why does a training forward pass keep its intermediate tensors when an inference-only forward pass can discard them?

level: juniorimportance: must knowfreq 68%

answer

  1. what does inference never run?
  2. no reverse sweep, nothing to replay
  3. the graph is a training-time record
  4. gradient rules read forward values
  5. an op keeps what its own backward reads

basics

~20 s

Training runs a backward pass, and each op's local gradient is a function of the tensors that op actually saw. Inference never runs backward, so an intermediate can be released the moment the next op has consumed it.

solid answer

~50 s

During training the forward pass has two jobs: produce the output, and leave behind a record the backward pass can replay. That record is the graph of ops plus, for each op, the specific tensor its own gradient rule reads back. Because those gradient rules are evaluated at the point the forward pass visited, the values have to still be there when the reverse sweep arrives. In inference-only mode there is no reverse sweep, so nothing needs to survive past its last consumer: an accelerometer anomaly detector scoring one window can overwrite each hidden activation as soon as the next layer has read it, and only the final score leaves the network. That is why the same architecture, the same weights and the same input behave differently in the two modes even though the arithmetic of the forward pass is identical.

go deeper

for a junior

Be ready to say plainly that training needs a backward pass and inference does not, and that the backward pass reads values the forward pass produced. Knowing that inference-only mode records no graph is the expected answer.

for a middle

Explain the mechanism: recording appends a record per op and holds a reference to the tensor that op's gradient rule reads, which is what keeps intermediates alive. Be able to say why a gradient rule is evaluated at the forward values.

for a senior

Show you know where this bites in practice — a validation or logging pass accidentally left recording on, or a handle kept on a graph-attached tensor past the end of a step. Distinguish freezing parameters from disabling recording without being prompted.

for a principal

Own the policy view: which code paths in a training system should record and which must not, how that is enforced rather than remembered, and what it costs the team when an evaluation path silently records. Treat mode discipline as an interface decision, not a per-call habit.

## Two different jobs for the same forward pass Running a network forward means applying a chain of ops to an input until you get an output. If that is *all* you want — a score, a class, a prediction — then every value produced along the way is scaffolding. Once layer 3 has read layer 2's output, layer 2's output is dead: nothing will ever look at it again, and the storage holding it can be reused immediately. Training is different, because training is a forward pass *followed by* a backward pass. The backward pass computes, for every parameter, the derivative of the loss with respect to that parameter, and it does so by walking the chain in reverse and applying each op's local gradient rule. The crucial point is that a local gradient rule is not a constant — it is a function evaluated **at the point the forward pass visited**. ## Local gradients are evaluated at the forward values Take `h = tanh(a)`. The derivative of `tanh` at `a` is `1 - tanh(a)^2`, which depends on where `a` was. If you throw away both `a` and `h`, the backward pass has no way to know which point on the curve to evaluate — the answer is different for `a = 0.1` and `a = 3.0`. The same is true of a linear layer: the gradient with respect to its weights is built from the layer's *input*, so that input has to still exist when the sweep reaches that layer. So the rule is not "training keeps everything because training is expensive". The rule is: **an op keeps whichever tensor its own backward step will read**, and it keeps it until the reverse sweep arrives at that op. Everything else can go. ## What "recording" actually means When recording is on, running an op does three things instead of one: 1. computes the output tensor; 2. appends a record to the graph saying which op ran, on which inputs, producing which output; 3. holds a reference to the tensor(s) that op's gradient rule will need, which keeps them alive. Step 3 is the reason intermediates stay reachable. Not because someone decided to hoard them, but because there is a live reference from the graph record to them. Ordinary garbage collection cannot reclaim a tensor that the graph is still pointing at. ## Inference-only mode When recording is turned off, step 2 and step 3 simply do not happen. There is no graph, no saved tensors, and consequently no backward pass is possible — asking for gradients afterwards is not "slow", it is undefined, because the information they would be computed from was never written down. Consider a small MLP that scores accelerometer windows for anomalies. In deployment it evaluates the same matrix multiplies and nonlinearities as during training, but each hidden activation is transient: produced, consumed by the next layer, released. Only the anomaly score survives the call. This asymmetry is worth internalising because it explains several everyday observations at once: - The same batch that runs fine at serving time may not run the same way while training, even with identical weights. - A forward pass done purely to *look at* something — logging an activation, computing a validation metric, generating a sample — has no reason to record, and recording it is pure waste. - Conversely, if you accidentally run a step with recording off and then ask for gradients, nothing is there. ## The distinction that trips people up Inference-only recording is not the same thing as freezing parameters. Freezing means a parameter will not be updated; the graph is still recorded, gradients still flow *through* the frozen layer to reach whatever sits behind it. Turning recording off is a statement about the graph itself: no record, no saved tensors, no backward. A network can be fully trainable and still be run in non-recording mode for a validation batch, and a network can have every parameter frozen while still recording, because something downstream needs the gradient path. ## Mental model Think of the forward pass under training as writing a receipt for every op: what ran, and the one value needed to reverse it. The backward pass reads the receipts in reverse order. Inference throws the receipts away as it goes, because nobody is going to ask for a refund.

  • If gradients are only needed for the parameters, why keep activations at all rather than just the weights?
    Because a parameter's gradient is not a function of the parameter alone. A linear layer's weight gradient is built from the tensor that entered that layer, so the activation flowing in is exactly what the rule reads. Weights persist anyway — they are long-lived state. It is the transient per-batch activations that the graph has to deliberately hold onto until the sweep reaches them.
  • Is running with recording disabled the same as freezing a layer's parameters?
    No. Freezing says a parameter will not be updated, but the graph is still built and gradients still pass through that layer to reach anything behind it. Disabling recording says no graph is built at all, so no backward pass is possible anywhere in that region. You can freeze every parameter and still need recording, and you can run a fully trainable network with recording off for a validation batch.
  • You want to log a hidden activation during training. Does that change what has to be retained?
    Reading the value for logging does not, by itself, change the graph — the tensor was already being held for the backward pass. What matters is whether you keep your own reference to it after the step finishes: holding a handle to an activation, or to anything that still points at the graph record, keeps that whole record alive past the point the sweep would have released it. Log the plain number, not the graph-attached tensor.

Cooking a dish for dinner, you rinse each bowl the moment it is empty. Cooking it while writing the recipe, you keep every measured quantity on the counter until you have written the step down.

saying these in an interview costs you the question

  • Says training keeps activations just because training is slower anyway
  • Thinks gradients can be recomputed from the weights alone
  • Believes the backward pass re-runs the forward pass from the input
  • Confuses disabling recording with freezing parameters
  • Claims inference and training forward passes differ in their arithmetic

context

open as a page

How does an elementwise max such as ReLU route gradient in backpropagation?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A max passes the incoming gradient straight to whichever input won and gives exactly zero to the loser. For ReLU that is a 0/1 mask: positive inputs pass the upstream gradient through unchanged, non-positive ones cut it to zero.

open as a page

In the backward pass, how does a max-pooling layer route gradients compared with average pooling?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Max pooling sends a window's entire upstream gradient to the one input that was the maximum and exactly zero to the others. Average pooling spreads it evenly, so each input in a k-by-k window receives one k-squared-th of it.

open as a page

What is the computational graph a forward pass records, and in what order does the backward sweep replay it?

level: middleimportance: must knowfreq 72%

basics

~20 s

The forward pass records a directed acyclic graph whose nodes are ops and whose edges are the tensors between them. The backward sweep replays it in reverse topological order: a node is visited only after every op that consumed its output.

open as a page

In a 3-layer MLP, how do you derive dL/dW2 and the gradient handed back to the hidden layer?

level: middleimportance: must knowfreq 72%

basics

~10 s

Start at layer 2's pre-activation with delta2 = dL/dz2. Then dL/dW2 is the outer product delta2 h1^T, dL/db2 is delta2 itself, and the hidden layer receives W2^T delta2, which you multiply elementwise by f'(z1).

open as a page

In backprop through a matrix product C = A B, what are the gradients with respect to A and B?

level: middleimportance: must knowfreq 70%

basics

~20 s

With G as the upstream gradient of the loss with respect to C, dL/dA = G B^T and dL/dB = A^T G. Matmul has two separate local rules, one per operand, and each carries exactly one transpose.

open as a page

In a convolution layer, what are the gradients with respect to its input and its kernel weights?

level: middleimportance: must knowfreq 62%

basics

~20 s

A convolution's input gradient is itself a convolution: the upstream gradient against the kernel rotated 180 degrees. The kernel gradient correlates the layer input with the upstream gradient, summed over every position the kernel visited.

open as a page

With a batch of 64 records, what shape is a 128-unit layer's bias gradient, and why?

level: juniorimportance: should knowfreq 58%

basics

~20 s

The bias gradient is a 128-length vector, the same shape as the bias itself. All 64 records contribute their own delta to the same shared bias, so the batch axis is summed away rather than kept.

open as a page

Why does a bias vector broadcast across a batch get a gradient summed over the batch axis?

level: middleimportance: should knowfreq 58%

basics

~20 s

Broadcasting copies the same bias into every row of the batch, so every row contributes a separate term to the same parameter. The backward of a copy is a sum, so those contributions add up into one bias-shaped gradient.

open as a page

When one embedding id appears many times in a batch, what gradient does its row receive?

level: middleimportance: should knowfreq 44%

basics

~20 s

The sum of the upstream gradient vectors from every occurrence. The lookup gathers rows, so the backward scatter-adds into them. Overwriting instead of accumulating would apply only the last occurrence and silently discard the rest.

open as a page

For a linear -> tanh -> linear chain, which tensor must each op stash for its own backward step?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Each op stashes only what its own backward step reads: a linear op keeps its input, tanh keeps its own output, ReLU needs nothing more than the sign of its input, and a plain addition keeps no tensor at all.

open as a page

Which hand-derived backprop mistakes survive a shape check and still break training?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Shape checks only prove conformability. A transpose flipped between equal dimensions, an activation derivative dropped or taken at the wrong layer, and a sum where the loss averaged all leave every shape valid while corrupting the gradient.

open as a page

How does switching a loss from a batch mean to a plain sum change every gradient in the graph?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Every gradient is multiplied by the batch size N. A mean reduction hands each element a local gradient of 1/N; a plain sum hands it 1. The descent direction is identical, the step is N times larger.

open as a page

Why does BatchNorm's backward pass make one sample's input gradient depend on the whole batch?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Because the batch mean and variance are functions of every sample, not constants. Differentiating through them puts batch-wide sums into each sample's gradient, so the same input gets a different gradient depending on its batch-mates.

open as a page