skip to content

With a batch of 64 records, what shape is a 128-unit layer's bias gradient, and why?

level: juniorimportance: should knowfreq 58%

answer

  1. parameters are shared across the batch
  2. the derivative of a sum is a sum of derivatives
  3. the update must subtract elementwise
  4. activations keep the batch axis, parameters do not
  5. the matmul that contracts the batch index is the sum

basics

~20 s

The bias gradient is a 128-length vector, the same shape as the bias itself. All 64 records contribute their own delta to the same shared bias, so the batch axis is summed away rather than kept.

solid answer

~40 s

Parameters are shared across the batch, so the batch axis never survives into a parameter gradient. Take a per-store weekly demand-forecast head with a 128-unit hidden layer and a batch of 64 store-weeks: the layer's pre-activations form a 64x128 block, and its delta block `Delta1 = dL/dZ1` is also 64x128. Because the same b1 is added to every row, `dL/db1` is the column-wise sum of Delta1 over the batch axis, giving shape (128,) — one number per unit, not one per unit per example. The weight gradient behaves identically: it is the sum of the 64 per-example outer products, which in a row-batched layout is written `X^T Delta1`, and its shape stays that of W. The rule to state: a parameter gradient always has the parameter's shape; an activation gradient keeps the batch axis.

code

python · 16 lines
python
W = [[0.5, -1.0, 2.0], [1.5, 0.0, -0.5]]                    # 2 outputs x 3 inputs
H = [[1.0, 2.0, 3.0], [0.0, 1.0, -1.0], [2.0, -1.0, 0.5]]   # 3 examples x 3 inputs
D = [[0.1, -0.2], [0.3, 0.5], [-0.4, 0.0]]                  # dL/dz: 3 examples x 2

dW = [[0.0] * 3 for _ in range(2)]
db = [0.0] * 2
for h, d in zip(H, D):                 # accumulate over the batch axis
    for i in range(2):
        db[i] += d[i]                  # bias: plain sum of deltas
        for j in range(3):
            dW[i][j] += d[i] * h[j]    # weights: outer product per example

dH = [[sum(d[i] * W[i][j] for i in range(2)) for j in range(3)] for d in D]

print(len(dW), len(dW[0]), len(db))    # 2 3 2  -> dW matches W, db matches b
print(len(dH), len(dH[0]))             # 3 3    -> one input gradient per example

go deeper

for a junior

Recall the two rules and you can answer any shape question on this leaf: parameter gradients match the parameter, activation gradients match the activation. Say why — the same bias is reused by every example in the batch.

for a middle

Explain the mechanics: the matrix product that forms the weight gradient is exactly the sum of per-example outer products, and the batch index is the one it contracts. Be able to write the full ledger for a two-layer head from memory.

for a senior

Demonstrate that you use the ledger as a debugging tool. Show how choosing distinct layer widths and writing every shape before any arithmetic turns a whole class of derivation mistakes into an immediate conformability failure.

for a principal

Own the convention itself. Decide whether losses in your codebase sum or average over the batch, state it once, and make sure step sizes and gradient-scale expectations are quoted against that convention so results stay comparable across batch sizes.

## The invariant There are exactly two shape rules in a hand-derived backward pass, and almost every batch confusion dissolves once you say them out loud: 1. **A gradient with respect to a parameter has the parameter's shape.** No batch axis, ever. 2. **A gradient with respect to an activation has that activation's shape.** The batch axis is present, because each example has its own activations. The reason for the first rule is that the parameter is *shared*. Every one of the 64 records in the batch is pushed through the same bias vector, so the loss depends on `b1` through 64 different routes, and the derivative of a sum is the sum of the derivatives. The 64 contributions are added together, not stacked. ## A concrete ledger Take the demand-forecast head: 64 store-weeks in the batch, a hidden layer of 128 units, an output head of 10 values (a 10-week horizon, one value per week). Write the forward pass in the row-batched layout, where each row of a block is one example: ``` X : 64 x 32 one row per store-week Z1 = X W1 + b1 W1: 32 x 128, b1: (128,) -> Z1: 64 x 128 H1 = f(Z1) H1: 64 x 128 Z2 = H1 W2 + b2 W2: 128 x 10, b2: (10,) -> Z2: 64 x 10 ``` The bias is *broadcast* down the batch: the same 128 numbers are added to all 64 rows. Now the backward ledger, with `Delta_k = dL/dZ_k`: ``` Delta2 = dL/dZ2 64 x 10 (activation gradient: batch axis kept) dL/dW2 = H1^T Delta2 128 x 10 (parameter gradient: batch axis contracted) dL/db2 = column sum of Delta2 (10,) (parameter gradient: batch axis summed) dL/dH1 = Delta2 W2^T 64 x 128 (activation gradient: batch axis kept) Delta1 = dL/dH1 * f'(Z1) 64 x 128 dL/db1 = column sum of Delta1 (128,) <- the answer to the question dL/dW1 = X^T Delta1 32 x 128 ``` Notice how `dL/dW2 = H1^T Delta2` is doing double duty. Per example, the weight gradient is an outer product; the matrix product `H1^T Delta2` is precisely the sum of those 64 outer products, with the shared index — the batch index — contracted away by the matmul. The batch sum is not an extra step bolted on; it is what that transpose-then-multiply *is*. ## Which orientation of the transpose is conformable People memorise "the backward pass uses W transposed" and then cannot recall whether to write `W^T Delta` or `Delta W^T`. Do not memorise it — read it off the layout. - In the **per-example column layout** (`z = W h + b`, W is 10x128, delta is 10x1), the input gradient is `W^T delta`: (128x10)(10x1) = 128x1. - In the **row-batched layout** (`Z = H W + b`, W is 128x10, Delta is 64x10), the input gradient is `Delta W^T`: (64x10)(10x128) = 64x128. They are the same mathematics under a transposed layout. In each case only one arrangement is conformable and only one produces the shape of the thing you are differentiating with respect to, so the shape ledger decides the question for you. ## Sum or mean — where the 1/64 lives The batch axis is summed in the gradient, but that does not settle whether the *loss* summed or averaged over the batch. If you define `L = (1/64) * sum_i L_i`, the factor `1/64` is already baked into `Delta2` at the top of the backward pass and rides along through every step; no later stage needs a second division. If you define L as a plain sum, every parameter gradient is 64 times larger, and training with an unchanged step size behaves as if the learning rate were 64 times bigger. Both conventions are internally consistent — mixing them is what hurts, because the effective step size then silently tracks the batch size. ## The mistake this question is testing for The wrong answer is "64 x 128": one bias gradient per example. It sounds plausible because every other block in the backward pass carries a batch axis. But a parameter update subtracts the gradient from the parameter elementwise, so a 64x128 object could not be applied to a 128-length bias at all. Per-example parameter gradients are computable and are occasionally wanted for specialised purposes, but they are not what an ordinary update consumes; the ordinary backward pass produces the summed one directly.

  • In a row-batched layout with Delta at 64x10 and W at 128x10, which orientation gives the input gradient?
    Delta W^T: (64x10)(10x128) gives 64x128, which is the shape of the layer's input block. The other arrangement is not conformable at all. In the per-example column layout, where W is 10x128 and delta is 10x1, the same gradient is written W^T delta. Read the orientation off the shapes rather than memorising a formula.
  • If the loss averages over the batch instead of summing, where does the 1/64 appear in the backward pass?
    At the very top. The factor lands in the delta at the output pre-activation and is carried unchanged through every transpose, outer product and batch sum below it, so no later step divides again. Sum-based and mean-based conventions are both correct; what breaks is mixing them, because the effective step size then scales with batch size.
  • Why can't you keep a per-example weight gradient and apply it directly?
    You can compute one — it is a stack of 64 outer products — but an ordinary update subtracts a single matrix from W elementwise, so the stack has to be reduced before it is usable. The batch sum is that reduction, and doing it inside the backward pass costs nothing extra because the matmul that computes the weight gradient already contracts the batch index.

saying these in an interview costs you the question

  • Says the bias gradient is 64 by 128, one row per example
  • Thinks a bigger batch produces a bigger-shaped parameter gradient
  • Averages the weight gradient but sums the bias gradient
  • Confuses summing over the batch with summing over the layer's units
  • Guesses the transpose orientation instead of reading the shapes

context