skip to content

Which hand-derived backprop mistakes survive a shape check and still break training?

level: seniorimportance: should knowfreq 40%

answer

  1. conformability is not correctness
  2. square layers hide half the mistakes
  3. an elementwise multiply changes no shape
  4. pick ugly, all-different layer widths on purpose
  5. wrong scale still trains, wrong direction does not

basics

~20 s

Shape checks only prove conformability. A transpose flipped between equal dimensions, an activation derivative dropped or taken at the wrong layer, and a sum where the loss averaged all leave every shape valid while corrupting the gradient.

solid answer

~50 s

Conformability is a weak test, and it gets weaker the more equal your dimensions are. The classic shape-silent bugs are: flipping a transpose in a layer whose input and output widths match, so both orientations are legal; writing the weight gradient as `h delta^T` instead of `delta h^T` on that same square layer; forgetting the elementwise `f'(z)` factor entirely, since dropping an elementwise multiply changes no shape; evaluating `f'` at the wrong layer's pre-activation, an off-by-one that is also shape-clean; and summing over the batch when the loss averaged, which scales every gradient by the batch size. The practical defence is to make the bug loud: derive and test on a network whose layer widths are all different, so any misplaced transpose becomes an immediate conformability failure, and to separate scale errors from direction errors by their symptoms.

go deeper

for a junior

Take away the headline: shapes lining up does not mean the gradient is right. Learn the two examples that show it — an elementwise activation factor left out, and a transpose flipped on a layer whose widths happen to match.

for a middle

Be able to list the shape-silent bugs and explain what each does to training: a flipped transpose sends gradients through the forward map, a missing batch factor rescales every step, a misindexed activation derivative corrupts every layer below it.

for a senior

Show a method, not a list. Distinct layer widths to restore the shape check's power, a batch-size comparison to test the batch axis, the scale-versus-direction symptom split, and the top-layer-versus-recursion split to localise the fault.

for a principal

Frame it as a defect-class question. Decide what your team standardises on so these bugs cannot survive — required conventions for custom gradients, debug configurations with deliberately unequal shapes, and a single documented batch-reduction convention across all losses.

## Why the shape check is not enough Writing the shape ledger before the arithmetic is the right habit, and it catches a lot. But a shape check answers only one question — *is this product legal, and does the result match the thing being updated* — and there are whole families of derivation errors it cannot see. Knowing which ones they are is what separates someone who has actually hand-derived a backward pass in anger from someone who has only read the formulas. ## The shape-silent family **1. A flipped transpose between equal dimensions.** The transpose in `W^T delta` is forced by the shapes only when the layer's input and output widths differ. A layer of 128 units feeding another 128 units makes `W` square, and then `W delta` is just as conformable as `W^T delta` while being the wrong gradient. Gradients now flow through the forward map instead of its adjoint. This is the single most common shape-silent bug, and the reason it hides is that debugging networks are so often built with uniform widths. **2. The outer product written backwards.** `dL/dW = delta h^T` becomes `h delta^T` — the transpose of the correct gradient. On a square layer both are legal shapes, so the update applies happily and moves the weights in a systematically wrong direction. **3. A dropped elementwise activation derivative.** Multiplying elementwise by `f'(z)` does not change any shape, so omitting it changes nothing a shape check can see. What it does is remove the nonlinearity's contribution from every delta below that point, so all lower layers train on a gradient that ignores which units were in a saturated or inactive regime. **4. The activation derivative at the wrong layer.** Using `f'(z2)` where `f'(z1)` belongs is an off-by-one in the recursion. When the two layers have the same width, the shapes line up perfectly, and every lower layer is corrupted while the top layer looks fine. **5. Sum where the loss took a mean.** If the loss is defined as an average over the batch but the derivation accumulates a plain sum, every gradient is larger by the batch factor. Shapes are untouched. This one does not point the update in a wrong direction — it just rescales it — which is why it is the most insidious of the five: training still works, just at an effective step size that silently tracks the batch size. **6. Pairing the delta with the wrong activation.** `dL/dW2 = delta2 h1^T` uses the *input* of layer 2. Reaching for `h2` instead — the output of the same layer — is a natural slip, and when the widths coincide it is shape-clean. ## Making the bugs loud The strongest defence is structural rather than diagnostic: **derive and test on a network with all-distinct layer widths**. Take input 7, hidden 13, hidden 5, output 3. Now a square matrix does not exist anywhere in the network, so every flipped transpose, every reversed outer product and every mispaired activation becomes an immediate conformability failure instead of a slow training mystery. Uniform widths are convenient for real models and terrible for debugging derivations; use ugly widths deliberately for the derivation, then move to real ones. A second structural move: run the backward pass with batch size 1 and again with batch size 2 on duplicated rows. Everything that treats the batch axis correctly gives the same parameter gradients when the loss averages, or exactly doubled ones when it sums. A batch-axis bug shows up as neither. ## Separating scale errors from direction errors When training is bad, the first split to make is *is the gradient pointing the wrong way, or merely scaled wrong*: - A **scale** error (the mean/sum factor) keeps the direction correct. The loss still falls, the run is unusually sensitive to the step size, and rescaling the step size by the batch factor makes the pathology disappear entirely. Nothing else about the run looks strange. - A **direction** error does not respond to step-size tuning in that way. Shrinking the step size just slows down whatever wrong thing is happening; the loss stalls, wanders, or diverges across a wide range of step sizes. ## Localising by layer One more piece of leverage: the layers are not symmetric in how much of the derivation they exercise. The top layer's parameter gradient uses only the delta the loss handed back and that layer's input — it does not use the recursion at all. Every layer below it additionally uses the transpose step and the elementwise activation step. So if the top layer's gradients behave and everything below is broken, the fault is almost certainly in the recursion — the transpose orientation or the activation-derivative factor — and not in the loss or in the outer-product rule. If even the top layer is wrong, look at the delta the loss produces or at the outer product itself. That single split usually cuts the search space in half in one step. ## What to say in the interview Name two or three concrete shape-silent bugs, then immediately give the structural defence: distinct layer widths so the shape check regains its teeth, batch-size-1-versus-2 to expose batch-axis handling, and the scale-versus-direction symptom split to classify what you are looking at. The interviewer is checking whether you know that "the shapes match" is a statement about conformability and not about correctness.

  • How would you design a small debug network so that a flipped transpose cannot stay silent?
    Give every layer a different width — say 7, 13, 5, 3. With no square weight matrix anywhere, W delta and W^T delta can never both be conformable, and neither can a reversed outer product or a delta paired with the wrong activation. The shape check then catches structurally what it would otherwise miss, before you spend any time on training curves.
  • Which symptom distinguishes a missing 1/batch factor from a genuinely wrong gradient direction?
    A missing factor is a pure scale error: the direction is right, the loss still falls, and rescaling the step size by the batch factor removes the problem completely. A wrong direction does not clean up that way — reducing the step size only slows the wrong behaviour down, and the loss stalls or diverges across a wide range of step sizes.
  • The top layer's gradients look right while every layer below is wrong. What does that localise?
    The recursion. The top layer's weight gradient uses only the delta from the loss and that layer's input, so it exercises neither the transpose step nor the activation-derivative step. Everything below uses both. That points squarely at a flipped transpose or a dropped or misindexed f'(z), and exonerates the loss and the outer-product rule.

A shape check is like verifying that a wiring diagram's plugs fit the sockets. Every connector seating correctly tells you nothing about whether two wires are swapped, and identical connectors everywhere is what makes the swap possible.

saying these in an interview costs you the question

  • Treats matching shapes as proof the derivation is correct
  • Debugs on a network whose layers all have the same width
  • Blames the learning rate before checking the derivation
  • Ignores whether the loss summed or averaged over the batch
  • Assumes a wrong gradient always makes the loss explode
  • Cannot say which layer a symptom localises the bug to

context