In the chain rule for f(g(h(x))), why is the total Jacobian the product J_f J_g J_h in that order?
answer
- shapes must be conformable
- innermost stage sits furthest right
- composition of linear maps, right to left
- associative but not commutative
- three factors, three evaluation points
basics
~20 sEach factor must consume the output of the stage before it, so shapes conform only outermost-first: (m by q)(q by p)(p by n) yields the m by n total derivative. Matrix products are not commutative, so no other order works.
solid answer
~40 sWrite the stages as `h: R^n -> R^p`, `g: R^p -> R^q`, `f: R^q -> R^m`. Their Jacobians have shapes p by n, q by p and m by q, and the only conformable arrangement is `J_f J_g J_h`, which is m by n — exactly the shape the total derivative of the composed map must have. The deeper reason is that the derivative of a composition is the composition of the linear approximations, and composing linear maps corresponds to multiplying their matrices outermost-first. Matrix multiplication is not commutative, so reordering is not a harmless rearrangement. Just as important, each factor is evaluated at a different point: `J_h` at x, `J_g` at `h(x)`, and `J_f` at `g(h(x))`. Dropping the evaluation points is the error that makes a correct-looking product numerically wrong.
go deeper
Be ready to state that the chain rule composes derivatives stage by stage and that in the vector case the composition becomes a product of matrices rather than of numbers.
An interviewer expects you to derive the ordering from the shapes, give the correct dimensions of the result, and name the point at which each factor is evaluated.
Show that you check composed derivatives by dimension before computing them, and that you know intermediate values must be retained because the factors depend on them.
Own the framing that linearization respects composition, and be able to explain why grouping is free but ordering is not, including the cost consequences of the grouping choice.
## The statement Let `h: R^n -> R^p`, `g: R^p -> R^q` and `f: R^q -> R^m` all be differentiable, and let `F = f o g o h`, so `F: R^n -> R^m`. The chain rule says `J_F(x) = J_f(g(h(x))) * J_g(h(x)) * J_h(x)` with shapes (m by q)(q by p)(p by n) = m by n. ## Why the order is forced, argument one: shapes A matrix product `A B` is defined only when the number of columns of A equals the number of rows of B. Here `J_h` is p by n, `J_g` is q by p, `J_f` is m by q. Try the reversed order: `J_h J_g` would need n columns of the left factor to meet q rows of the right, which happens only by coincidence, and even when the numbers accidentally line up the resulting matrix does not have the meaning you want. The dimensions themselves tell you which side each Jacobian goes on: the innermost stage, the one that touches the original input, must sit at the far right, because it is the factor that will multiply the input displacement first. ## Why the order is forced, argument two: composition of linear maps The derivative is not fundamentally a matrix, it is a linear map: the best linear approximation to the function near a point. If `A` is the linear map approximating f and `B` the one approximating g, then the approximation to `f o g` is the composed map `A o B`, meaning apply B first and then A. In matrix language, composition of linear maps is matrix multiplication written in that same right-to-left order. So the chain rule is not a separate fact about derivatives; it is the statement that linearization respects composition. You can see it operationally too. Nudge the input by a small d. The first stage produces a change of about `J_h d`. The second stage sees that change as its own input and produces about `J_g (J_h d)`. The third produces about `J_f (J_g J_h d)`. Peeling the parentheses off gives the product in exactly the stated order. ## Non-commutativity is not a technicality Even when all the stages map R^k to R^k and every Jacobian is square, so that every ordering is conformable, the answer still depends on the order, because matrix products generally differ when swapped. There is no version of the rule where you may reorder the factors for convenience. This is the precise sense in which the single-variable mnemonic misleads: in one dimension every Jacobian is a 1 by 1 matrix, ordinary numbers commute, and so the du/dx-style cancellation appears to work. It is a coincidence of dimension one. ## Evaluation points Each Jacobian is a function of where you evaluate it, and the three factors are evaluated at three different places: - `J_h` at the original input x, - `J_g` at the intermediate value `h(x)`, - `J_f` at the second intermediate value `g(h(x))`. A derivation that writes `J_f J_g J_h` without saying this is only half correct, and a numerical implementation that evaluates all three at x will produce plausible-looking garbage. In practice this means a forward pass through the stages must record the intermediate values, because the derivative factors need them. ## A worked shape check Suppose `h: R^5 -> R^3`, `g: R^3 -> R^2`, `f: R^2 -> R^4`. Then `J_h` is 3 by 5, `J_g` is 2 by 3, `J_f` is 4 by 2, and `J_F = J_f J_g J_h` has shape (4 by 2)(2 by 3)(3 by 5) = 4 by 5, matching a map from R^5 to R^4. Every adjacent pair agrees, and the outer dimensions reproduce the overall input and output sizes. That two-line check is the fastest way to validate any composed derivative before touching numbers. ## Associativity, and why it is useful Matrix multiplication is associative even though it is not commutative, so `(J_f J_g) J_h` and `J_f (J_g J_h)` give the same result. The grouping is free; the order is not. The grouping does change the arithmetic cost — for a scalar final output, multiplying from the left keeps every intermediate a single row, whereas multiplying from the right builds full matrices — which is why the associativity remark is worth knowing even when only correctness is being asked about. ## Answering well Give the shape argument first because it is short and checkable, then the linear-map argument because it explains rather than verifies, then name the evaluation points. Adding that associativity holds while commutativity does not is the detail that shows real command of the rule.
- Where is each of the three Jacobians evaluated?`J_h` at the original input x, `J_g` at the intermediate `h(x)`, and `J_f` at `g(h(x))`. Each factor is a matrix-valued function of its own stage's input, so the intermediate values must be computed and kept. Evaluating all three at x is a silent numerical error that leaves shapes correct and answers wrong.
- Why does the single-variable chain rule seem to be order-free?Because in one dimension every Jacobian is a 1 by 1 matrix, that is, an ordinary number, and numbers commute. The apparent cancellation of differentials in the Leibniz notation is an artefact of dimension one. As soon as an intermediate quantity is a vector, the factors are matrices and their order is fixed by conformability and by non-commutativity.
- If all three maps go from R^k to R^k, may the factors then be reordered?No. Square shapes make every ordering conformable, but matrix multiplication still does not commute, so a reordered product is generally a different matrix. Conformability is a necessary condition, never a licence to rearrange. Associativity does hold, so you may regroup with parentheses freely, which affects cost but not the result.
It is an assembly line read backwards: the last station's response depends on what the middle station handed it, which depends on what the first station made.
saying these in an interview costs you the question
- Writes the product innermost-first without checking shapes
- Claims matrix factors may be reordered when they are square
- Evaluates every Jacobian at the original input point
- Treats the differentials as cancelling in the vector case
- Confuses associativity with commutativity