skip to content

Deriving Backprop By Hand

Walking a two-layer network back from the loss: the delta at the output, the same delta pushed through the transposed weight matrix, and each weight gradient falling out as an outer product.

on this pageshow

questions

3

In a 3-layer MLP, how do you derive dL/dW2 and the gradient handed back to the hidden layer?

level: middleimportance: must knowfreq 72%

answer

  1. define the intermediate at the pre-activation
  2. weight gradient pairs an output delta with an input value
  3. outer product, not a matmul of two matrices
  4. the sum over the output index becomes a transpose
  5. activation derivative belongs to the layer below

basics

~10 s

Start at layer 2's pre-activation with delta2 = dL/dz2. Then dL/dW2 is the outer product delta2 h1^T, dL/db2 is delta2 itself, and the hidden layer receives W2^T delta2, which you multiply elementwise by f'(z1).

solid answer

~40 s

Take a triage net with three weight matrices: W1 is 64x12, W2 is 32x64, W3 is 3x32, and each hidden layer applies an elementwise nonlinearity f. Define delta at every pre-activation: `delta_k = dL/dz_k`. Layer 2's delta comes from the layer above as `delta2 = (W3^T delta3) * f'(z2)`, elementwise. Because `z2 = W2 h1 + b2`, the loss touches W2 only through z2, so `dL/dW2_ij = delta2_i * h1_j` — that is the outer product `delta2 h1^T`, shaped 32x64 like W2. The bias sits under a coefficient of one, so `dL/db2 = delta2`. Going further back, each h1_j feeds every z2_i, so you sum over the output index: `dL/dh1 = W2^T delta2`, shaped 64x1. Multiply that elementwise by `f'(z1)` to get delta1, and `dL/dW1 = delta1 x^T`.

go deeper

for a junior

Be ready to say what the backward pass produces: one gradient per parameter, each the same shape as the parameter it updates. Knowing that the weight gradient combines a signal from above with the layer's own input is enough at this level.

for a middle

You are expected to produce the three lines at the whiteboard without notes: weight gradient as the outer product, bias gradient as the delta, and W transposed times delta as what travels further back, followed by the elementwise activation derivative.

for a senior

Show the discipline, not just the formulas. Write the shape ledger first, name where each intermediate came from in the forward pass, and be able to say which slip your derivation would produce if a transpose or an activation index were wrong.

for a principal

Own the framing question: what is worth deriving by hand at all when the tooling differentiates for you. Argue that the value is in reading gradient-flow problems and custom layers, and set the expectation of which derivations your team should be able to do cold.

## The network you are deriving on Use a concrete one so nothing is ambiguous: a 3-layer MLP over 12-feature clinical triage records, predicting one of 3 acuity classes. ``` z1 = W1 x + b1 W1: 64x12 z1, h1: 64x1 h1 = f(z1) z2 = W2 h1 + b2 W2: 32x64 z2, h2: 32x1 h2 = f(z2) z3 = W3 h2 + b3 W3: 3x32 z3: 3x1 L = loss(z3, y) ``` Here `x` is one record as a column vector, `f` is an elementwise nonlinearity (any of them — the derivation does not care which), and "3-layer" means three weight matrices. ## Cut the network at the pre-activations The single decision that makes hand derivation easy is *where you define your intermediate variable*. Define it at the **pre-activation**, not at the weights and not after the nonlinearity: ``` delta_k = dL/dz_k (same shape as z_k) ``` This is the quantity people say out loud as "the delta of layer k". Everything else on the layer is one short step away from it, and the recursion between layers is a single line. The backward sweep starts at the top: the loss hands you `delta3 = dL/dz3`, whatever form that takes for your particular output head. From there each layer does the same three-part move. ## Step 1 — the weight gradient is an outer product Write `z2` entrywise: ``` z2_i = sum_j W2_ij * h1_j + b2_i ``` A single weight `W2_ij` appears in exactly one output entry, `z2_i`, with coefficient `h1_j`. So `dz2_i/dW2_ij = h1_j`, and multiplying by the delta that already sits on `z2_i`: ``` dL/dW2_ij = delta2_i * h1_j ``` Every entry of the weight gradient is (an output-side delta) times (an input-side activation). Stacked over all i and j, that is the **outer product**: ``` dL/dW2 = delta2 h1^T (32x1)(1x64) = 32x64 ``` It is an outer product and not an ordinary matmul of two big matrices because for one example, one index runs over the layer's outputs and the other over the layer's inputs, and nothing is being summed at all. The bias falls out of the same expansion: `b2_i` sits in `z2_i` with a coefficient of one, so ``` dL/db2 = delta2 (32x1) ``` ## Step 2 — push the delta back through the transpose Now ask what the *hidden activation* h1 receives. Unlike a single weight, `h1_j` feeds **every** entry of `z2`, so the chain rule sums over the output index: ``` dL/dh1_j = sum_i delta2_i * W2_ij ``` That sum contracts the first index of W2, which is exactly what the transpose does: ``` dL/dh1 = W2^T delta2 (64x32)(32x1) = 64x1 ``` This is the whole reason the transpose appears in backprop. Forward, `W2` maps a 64-vector to a 32-vector; backward, gradients travel the opposite way, from a 32-vector to a 64-vector, and `W2^T` is the only orientation with those dimensions. ## Step 3 — cross the nonlinearity elementwise `h1 = f(z1)` acts entry by entry, so its derivative is a diagonal Jacobian and the product degenerates to an elementwise multiply: ``` delta1 = (W2^T delta2) * f'(z1) elementwise, 64x1 ``` Two details are worth saying out loud, because both are common interview slips. First, `f'` is evaluated at the **pre-activation** `z1`, not at `h1` — even when a convenient identity lets you rewrite it in terms of the output value, the argument is still z1 conceptually. Second, the derivative belongs to the layer you are *entering*, layer 1, not layer 2; an off-by-one here is the classic hand-derivation bug. Then the recursion repeats: `dL/dW1 = delta1 x^T` (64x12), `dL/db1 = delta1`. ## The shape ledger Before writing any arithmetic, write the shapes. Two invariants hold at every layer and catch most mistakes instantly: - a parameter gradient has the **same shape as the parameter** — `dL/dW2` is 32x64 like W2, `dL/db2` is 32x1 like b2; - an activation gradient has the **same shape as the activation** — `dL/dh1` is 64x1 like h1. So when you are unsure whether to write `W2^T delta2` or `delta2^T W2`, you do not need to re-derive: only one of them produces a 64x1 result. ## The five-minute whiteboard version Asked to derive the second-to-last layer's update under time pressure, say this in order: define delta at the pre-activation; the weight gradient is delta times the layer's input transposed; the bias gradient is the delta; the gradient sent further back is W transposed times delta; then multiply by the earlier layer's activation derivative. State each shape as you go. That sequence is the whole of dense backprop, and every layer repeats it unchanged.

  • Why is dL/dW2 an outer product rather than a product of two full matrices?
    For a single example, a given weight W2_ij enters exactly one output entry, z2_i, with coefficient h1_j, so nothing is summed. The gradient entry is just delta2_i times h1_j, and stacking all i and j gives delta2 h1^T. A contraction only appears once you accumulate over a batch, where the same weight is reused by every example.
  • Why is the activation derivative evaluated at z1 rather than at h1?
    The nonlinearity is a function of the pre-activation: h1 = f(z1), so its derivative is f'(z1). Some activations admit an identity that re-expresses f'(z1) in terms of h1, which is a convenience, not a change of argument. If you carelessly plug h1 into f' where no such identity holds, you get a wrong delta that is still perfectly conformable and so passes every shape check.
  • If two linear layers sit back to back with no nonlinearity between them, what do the two transposes do?
    Backward through the second gives W2^T delta, and backward through the first gives W1^T W2^T delta, which equals (W2 W1)^T delta. The backward map through the composite is the transpose of the forward composite, in reversed order — the same reversal you see forward, read the other way. It is a useful sanity check that your transposes are placed correctly.

Forward, the weight matrix is a set of pipes carrying values up; backward, the same pipes carry blame down, and reading the same pipe network in the opposite direction is exactly what taking the transpose means.

saying these in an interview costs you the question

  • Multiplies by W2 instead of W2^T when going backwards
  • Writes dL/dW2 as h1 delta2^T, flipping the outer product
  • Applies the activation derivative of layer 2 to layer 1's delta
  • Evaluates f' at the post-activation h1 with no supporting identity
  • Thinks delta means the loss value rather than dL/dz
  • Derives entrywise but never states a single shape

context

open as a page

With a batch of 64 records, what shape is a 128-unit layer's bias gradient, and why?

level: juniorimportance: should knowfreq 58%

basics

~20 s

The bias gradient is a 128-length vector, the same shape as the bias itself. All 64 records contribute their own delta to the same shared bias, so the batch axis is summed away rather than kept.

open as a page

Which hand-derived backprop mistakes survive a shape check and still break training?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Shape checks only prove conformability. A transpose flipped between equal dimensions, an activation derivative dropped or taken at the wrong layer, and a sum where the loss averaged all leave every shape valid while corrupting the gradient.

open as a page