A layer's (8, 256) output adds a bias stored as (8, 1) instead of (256,) - what goes wrong?
answer
- broadcasting aligns from the right
- a length-256 vector versus an 8-by-1 column
- per-unit offset versus per-example offset
- change the batch size and it crashes
basics
~20 sBroadcasting aligns from the right, so a (256,) bias adds one value per unit to every row, while an (8, 1) bias adds one value per row across all units. Both run; only the first is a real bias.
solid answer
~50 sBroadcasting matches shapes from the trailing axis backwards, stretching any axis of extent 1. A `(256,)` bias is treated as `(1, 256)` and stretched down the 8 rows, so every example receives the same per-unit offset - that is what a bias means. An `(8, 1)` bias is stretched along the 256 columns instead, so example `i` gets a single scalar added to all of its units. The result still has shape `(8, 256)`, no error is raised, and training proceeds - but the offsets are now per example rather than per unit, which is not a function of the input at all, so the layer's units are all shifted together. The tell is that the bug is batch-size dependent: at batch 16 the leading axes 16 and 8 neither match nor are 1, so the addition finally raises an error.
go deeper
Know that a layer's bias has one entry per output unit, not one per example, and that its length must equal the layer's output width.
Be able to walk the broadcasting alignment aloud, trailing axis first, and say exactly which element ends up added to output position [i][j] under each bias shape.
Demonstrate the diagnosis: a shape bug that appears only at certain batch sizes points at a parameter carrying a batch-shaped axis, and a batch-independence test confirms it in one run.
Argue for the invariant, not the fix. Shape contracts asserted at parameter construction turn a class of silent correctness bugs into build-time failures across every model a team ships.
## What broadcasting actually does Elementwise addition between arrays of different shapes is resolved by a fixed rule: align the shapes at their **trailing** axis and walk leftwards. At each position the two extents must either be equal, or one of them must be 1, in which case that operand is virtually repeated along the axis. Missing leading axes are treated as 1. Nothing is copied in memory; the smaller operand is simply read repeatedly. ## The correct bias A fully-connected layer with 256 output units holds a bias of shape `(256,)` - one learned offset per unit. Adding it to a `(8, 256)` activation aligns as: (8, 256) (256,) -> treated as (1, 256) -------- (8, 256) The trailing axes match at 256; the leading axis is 1 against 8, so the bias row is reused for every example. Output element `[i][j]` gets `b[j]`. That is exactly the intent: the offset depends on **which unit** you are looking at, never on which example. ## The wrong bias Store the same numbers as `(8, 1)` - the mistake happens when a bias is built as a column, or squeezed out of a wrongly oriented parameter - and the alignment becomes: (8, 256) (8, 1) -------- (8, 256) The trailing axis is 1 against 256, so the single column is stretched across all 256 units; the leading axis matches at 8. Output element `[i][j]` now gets `b[i]`. The offset depends on **which example** you are looking at and is identical across every unit of that example. ## Why this is worse than it looks Three things make it nasty. First, the output shape is `(8, 256)` in both cases, so nothing downstream complains and the whole forward pass, loss and update run to completion. Second, a per-example offset is not a function the layer can meaningfully learn: the parameters have no way to know which row of a batch a given example landed in, so the same example gets a different offset depending on where it sits in the batch, and shuffling changes the model's output for an unchanged input. Third, it silently removes the real bias - the per-unit degrees of freedom the layer was supposed to have - so the layer becomes a plain linear map with a nuisance term glued on, and the symptom is a mildly worse model rather than a broken one. ## The batch-size tell The accident only conforms because the batch happened to equal 8, the same as the leading extent of the bias. Run the identical code at batch 16 and the alignment is 16 against 8 on the leading axis - neither equal nor 1 - and the addition raises a shape error. This produces the classic confusing report: it works, then it crashes on the last partial batch of the epoch, or after someone changes one number in a configuration file. If a shape bug appears and disappears with batch size, suspect a parameter that carries a batch-shaped axis it should not have. ## How to prevent it Assert the invariants where the parameters are created: a bias is rank 1, and its length equals the layer's output width. That single assertion rules out every orientation of this bug and costs nothing at run time. A second cheap check is a batch-independence test - run the same example alone and as row 3 of a batch of 8, and require the outputs to match to numerical tolerance. Any parameter that has picked up a batch-shaped axis fails that test immediately, whatever its rank happens to be.
- Why is a bias shaped (outputs,) rather than (batch, outputs)?Because a bias is a learned parameter of the layer, one offset per unit, shared by every example forever. A batch-shaped parameter would have no meaning at inference on a single input, would change with shuffling, and could not be saved and reloaded independently of the batch size used to train it.
- What one assertion would have caught this at construction time?Require the bias to be rank 1 with length equal to the layer's output width. That rejects (8, 1), (1, 256) and (256, 1) alike, runs once at build time, and is far cheaper than discovering the problem from a slightly worse validation curve weeks later.
- How would you detect it in a model you did not write and cannot re-run cheaply?Feed one example alone, then the same example placed at different rows of a padded batch, and compare outputs. A correct layer gives identical results; a per-example offset makes the output depend on the row index, which no legitimate feed-forward layer ever does.
The right bias is a stamp pressed identically onto every row. The wrong one is the same stamp rotated ninety degrees and pressed along each row instead - same ink, wrong direction.
saying these in an interview costs you the question
- Thinks any bias shape works because addition is elementwise
- Says (8, 1) and (256,) broadcast to the same thing
- Assumes a shape error would have caught it
- Believes each layer has one scalar bias overall
- Cannot say which axis broadcasting aligns first