skip to content

Why does a bias vector broadcast across a batch get a gradient summed over the batch axis?

level: middleimportance: should knowfreq 58%

answer

  1. the bias is used in every row
  2. one variable, many uses in the graph
  3. the adjoint of copying
  4. reduce over exactly the expanded axes

basics

~20 s

Broadcasting copies the same bias into every row of the batch, so every row contributes a separate term to the same parameter. The backward of a copy is a sum, so those contributions add up into one bias-shaped gradient.

solid answer

~50 s

Adding a bias of shape (D,) to activations of shape (N, D) is a broadcast: the one bias vector is reused in all N rows. Forward, that costs nothing and changes no shape logic. Backward, each row's upstream gradient is a genuine derivative contribution to the *same* D numbers, and the chain rule over a shared variable adds contributions, so `dL/db` is the upstream gradient summed over the batch axis, giving shape (D,). The gradient with respect to the activations is the upstream gradient unchanged, since addition passes gradient through untouched. The general rule is that a broadcast in the forward pass becomes a sum over the broadcast axes in the backward pass. Forget the sum and you hand a (256, D) gradient to a (D,) parameter -- either a loud shape error or, worse, a silent re-broadcast that corrupts every update.

go deeper

for a junior

Remember that a gradient always has the same shape as the thing it belongs to. If a parameter is smaller than the activations it was added to, something in the backward pass has to shrink the gradient back down.

for a middle

Be ready to derive the summation from the chain rule over a variable that appears in many places, and to state the general pairing: forward broadcast means backward sum over the expanded axes.

for a senior

Show the failure modes. Explain how a missing reduction can pass silently because the update itself broadcasts, and why asserting gradient shapes against parameter shapes belongs in any hand-written backward or custom op.

for a principal

Frame it as the general accumulation rule for shared parameters across batch, time, or position, and use it to set team conventions for custom ops -- gradient shape assertions and a finite-difference check on every hand-written backward before it goes near a training run.

## What broadcasting actually is Suppose a layer produces activations `Z` of shape (256, D) -- a batch of 256 rows, each of width D -- and a bias vector `b` of shape (D,) is added: `Y = Z + b`. The shapes do not match, so the smaller operand is *broadcast*: conceptually `b` is replicated into 256 identical rows, and then the addition is elementwise. Implementations do not physically copy anything, but the mathematics is exactly the copy. That word *copy* is the whole answer. Broadcasting is a `copy` op wearing an `add` op's clothes, and the local gradient of a copy is what makes the bias gradient look different from the activation gradient. ## The chain rule over a reused variable Call the upstream gradient `G = dL/dY`, of shape (256, D). Entry-by-entry, the forward is ``` Y[i][j] = Z[i][j] + b[j] ``` The variable `b[j]` appears in row 0, row 1, ..., row 255 -- 256 separate places in the graph. The multivariate chain rule says the derivative with respect to a variable used in several places is the *sum* of the derivatives through each use: ``` dL/db[j] = sum over i of G[i][j] * (dY[i][j] / db[j]) = sum over i of G[i][j] ``` because the local derivative of an addition with respect to either operand is 1. So `dL/db` is `G` summed down the batch axis, of shape (D,). And ``` dL/dZ[i][j] = G[i][j] ``` unchanged, because `Z` is used in exactly one place and addition scales nothing. Notice the asymmetry: the same op hands one operand the gradient untouched and the other a reduction of it. That is entirely down to which operand was broadcast. ## The rule to remember **A broadcast in the forward pass is a sum over the broadcast axes in the backward pass, and a sum in the forward pass is a broadcast in the backward pass.** The two operations are adjoints of each other. Once you hold that pairing, you never have to re-derive a broadcast gradient; you just ask which axes were expanded and sum over exactly those. The procedure generalises to any broadcast, not just a bias: 1. Find every axis where the operand was expanded -- an axis it did not have at all, or one it had with length 1 while the other operand had length n. 2. Sum the upstream gradient over those axes. 3. For axes that existed with length 1, keep them as length 1 so the result matches the operand's original shape; for axes the operand never had, drop them entirely. A (1, D) bias and a (D,) bias behave identically in the forward pass and produce gradients of different shapes in the backward pass -- (1, D) versus (D,) -- which is a genuine source of confusion when hand-writing a backward. ## What breaks when you forget The failure mode is worth rehearsing, because it is the classic hand-written-backward bug. - **The loud version.** You compute `dL/db = G`, hand a (256, D) gradient to a (D,) parameter, and the update step raises a shape mismatch. Annoying, but self-diagnosing. - **The quiet version.** The update `b <- b - lr * grad` itself broadcasts. Now `b` -- shape (D,) -- minus a (256, D) gradient produces a (256, D) *parameter*. Nothing errors. The bias silently stops being a bias, memory grows, and the model's behaviour becomes batch-order dependent. This is why shape assertions on gradients are worth writing even when the framework would eventually complain. - **The subtle version.** You sum over the wrong axis -- down the feature axis instead of the batch axis -- and get a (256,) gradient. If D happens to equal the batch size, even the shape check passes, and you get a model that trains, badly, for reasons nothing will report. ## Scale, not just shape One more property follows from the sum: the raw magnitude of `dL/db` grows with the number of rows summed, because it accumulates N terms. Whether that is a problem depends on how the loss aggregates over the batch -- a loss that averages over the batch already divides those N contributions back down, while one that adds them does not. The bias gradient is simply the most visible place where that batch-aggregation choice shows up, because it is where the summation over the batch is explicit rather than hidden. ## Interview framing If you are asked to hand-derive a small network's backward pass, the bias line is where interviewers watch for the sum. Saying "the bias gradient is the upstream gradient" is the giveaway that a candidate has memorised a per-example derivation and never thought about batching; saying "summed over the batch, because the bias is shared across rows" shows you understand that shared parameters accumulate.

  • What is the backward of a sum reduction, and how does it relate to the broadcast rule?
    They are adjoints of each other. A forward broadcast becomes a backward sum over the expanded axes, and a forward sum becomes a backward broadcast: every element that contributed to the sum receives the same upstream gradient, copied back into the original shape. Holding the pair together means you only have to remember one direction and can always recover the other.
  • If the bias gradient is not summed, what actually goes wrong at update time?
    Best case the update raises a shape mismatch, since a (D,) parameter cannot take an (N, D) gradient. Worst case the subtraction broadcasts instead of failing, and the bias silently becomes an (N, D) array -- no error, growing memory, and behaviour that depends on batch size. Asserting that each gradient matches its parameter's shape catches both.
  • Does the same summation rule apply to any other parameter shared across the batch?
    Yes -- the rule is about sharing, not about biases. Any parameter reused across positions, time steps, or batch elements accumulates one gradient contribution per use, and its total gradient is the sum over all of them. A bias is just the easiest case to see, because the sharing is a plain broadcast along a single axis.

saying these in an interview costs you the question

  • Says the bias gradient equals the upstream gradient unchanged
  • Treats broadcasting as a forward-only convenience with no backward cost
  • Sums over the feature axis instead of the batch axis
  • Thinks addition scales the gradient it passes to each operand
  • Assumes a shape error would always be raised if the sum is skipped

context