skip to content

Chain Rule & Curvature

How derivatives compose through nested functions, and what second derivatives say about a surface's shape: Jacobians, Hessians, Taylor expansion. Curvature explains why training crawls or diverges.

on this pageshow

explore

questions

9

For a map f from R^n to R^m, what shape is the Jacobian and what is entry (i, j)?

level: middleimportance: must knowfreq 68%

answer

  1. outputs by inputs
  2. reconstruct it from J times d
  3. rows are component gradients
  4. entry is one partial derivative
  5. it is a function of the point

basics

~20 s

The Jacobian is m by n: one row per output component, one column per input variable. Entry (i, j) is the partial derivative of output component f_i with respect to input x_j, evaluated at a specific point.

solid answer

~50 s

For `f: R^n -> R^m` the Jacobian `J` is an m by n matrix whose entry (i, j) is `d f_i / d x_j`, the partial derivative of the i-th output with respect to the j-th input. Row i is the gradient of the scalar component `f_i` written as a row. The shape is forced by what the Jacobian is for: it is the best linear approximation of f near a point, `f(x + d) ~= f(x) + J d`, and for `J d` to make sense with d in R^n and the result in R^m, J must be m by n. Two consequences matter in practice: J depends on the point, so it must always be written as `J(x)`, and when the output is a single scalar (m = 1) the Jacobian is a 1 by n row vector whose transpose is the gradient column.

go deeper

for a junior

Be ready to state the shape as outputs-by-inputs and to write entry (i, j) as the partial derivative of output i with respect to input j for a small explicit map.

for a middle

An interviewer expects you to justify the shape from the linear approximation f(x + d) is about f(x) + J d, rather than reciting it, and to compute a 2 by 2 Jacobian on the spot.

for a senior

Demonstrate shape checking as a working habit: catching dropped transposes and mismatched layouts by dimension alone, before evaluating anything numerically.

for a principal

Own the convention call. Decide numerator or denominator layout for a codebase or a team's notation, state it once, and explain what class of bug the inconsistency otherwise produces.

## Definition Let `f: R^n -> R^m` be differentiable, with input `x = (x_1, ..., x_n)` and output components `f_1(x), ..., f_m(x)`. The Jacobian matrix of f at a point x is `J(x)[i][j] = d f_i / d x_j` evaluated at x and it has m rows and n columns. Say it as outputs-by-inputs: the row index runs over the things being produced, the column index over the things being varied. ## Why m by n and not n by m The defining property of the derivative is linear approximation. Near x, `f(x + d) ~= f(x) + J(x) d` where d is a small displacement in the input space, so d is a column in R^n, and the left-hand side lives in R^m. A matrix that eats an n-vector and returns an m-vector has m rows and n columns. That is the whole argument, and it is the one to give in an interview: the shape is not a convention to memorise, it is whatever makes the product conformable. If you ever cannot remember the orientation, reconstruct it from `J d` in half a second. Row i of J is `(d f_i / d x_1, ..., d f_i / d x_n)` — the gradient of the scalar-valued component f_i laid out horizontally. Column j is the vector of how every output responds to nudging the single input x_j. ## A worked example Take `f: R^2 -> R^2` with `f(x_1, x_2) = (x_1^2 * x_2, x_1 + sin(x_2))`. Then `J = [[2 * x_1 * x_2, x_1^2], [1, cos(x_2)]]` Row 1 differentiates `x_1^2 * x_2` with respect to x_1 and then x_2; row 2 does the same for `x_1 + sin(x_2)`. Here n = m = 2, so the matrix is square — a special case, not the general one. A second example fixes the shape rule in memory: `f: R^5 -> R^3` given by `f(x) = A x + b` with A of shape 3 by 5. Every partial derivative is a constant entry of A, so `J(x) = A` everywhere. An affine map is its own linearization, and its Jacobian is exactly the matrix of the map — 3 by 5, m by n, as required. ## The scalar-output special case When m = 1, f is a scalar field such as a loss, and the Jacobian is a single row of length n. The gradient is conventionally a column vector, so `gradient = J^T`. Keeping these straight avoids a lot of transposition pain: Jacobians are wide (outputs by inputs), gradients of scalars are tall. Many disagreements about a stray transpose are really disagreements about which of these two objects someone wrote down. ## It is a function of the point The Jacobian is not one matrix attached to f; it is a matrix-valued function of x. In the worked example above, J at (1, 0) is `[[0, 1], [1, 1]]` while J at (0, 1) is `[[0, 0], [1, cos(1)]]`. Any statement about a Jacobian that does not say where it is evaluated is incomplete, and forgetting the evaluation point is the error that ruins multi-stage chain-rule computations, where each stage is evaluated at a different intermediate value. ## Layout conventions What is described here is the numerator layout: outputs index rows. Some texts use the transposed denominator layout, in which the same object appears as n by m. Neither is wrong, but mixing them inside one derivation produces transposes that appear from nowhere. The practical discipline is to declare the convention once, then verify every result by shape: if a quantity is a derivative of something m-dimensional with respect to something n-dimensional, and your expression is not m by n, something is wrong before any number is computed. ## Why it matters downstream Every composed derivative is built from these blocks. Once the shape rule is automatic, checking a long derivation reduces to checking that adjacent matrices are conformable and that the ends match the overall input and output dimensions. That single habit catches most sign-free algebra errors without evaluating anything. ## Answering well A strong answer states the shape, gives the entry formula, justifies the shape from the linear approximation rather than from memory, notes that rows are component gradients, and adds that the Jacobian is evaluated at a point. Mentioning the m = 1 case and the transpose relationship to the gradient shows the fluency an interviewer is listening for.

  • What is the Jacobian of the affine map f(x) = A x + b, where A is 3 by 5?
    It is A itself, at every point, shape 3 by 5. Each output is a fixed linear combination of the inputs, so every partial derivative is a constant entry of A, and the constant offset b differentiates away. An affine map is already its own linear approximation, which is why the Jacobian does not vary with x.
  • How does the Jacobian relate to the gradient when the output is a single scalar loss?
    With m = 1 the Jacobian is a 1 by n row vector of partial derivatives, and the gradient is its transpose, an n by 1 column. Same numbers, different orientation. Fixing that convention up front is what prevents stray transposes appearing halfway through a derivation.
  • You wrote a derivative expression and it is n by m instead of m by n. What does that tell you?
    Either a transpose was dropped, or you have silently switched to the denominator layout convention partway through. Check the shape against the rule outputs-by-inputs and against any product the expression must fit into. Shape checking catches this class of error before you evaluate a single number.

Think of a mixing desk: one row per output channel, one column per input fader, and each knob setting says how much that fader moves that channel.

saying these in an interview costs you the question

  • Says the Jacobian is n by m for a map from R^n to R^m
  • Calls the Jacobian a fixed matrix rather than a function of x
  • Confuses a column of the Jacobian with a component gradient
  • Cannot state entry (i, j) as one partial derivative
  • Treats the scalar-output Jacobian and the gradient as identical in shape

context

open as a page

What is the Hessian matrix of a scalar function, and when is it symmetric?

level: middleimportance: must knowfreq 72%

basics

~20 s

The Hessian is the square matrix of all second partial derivatives of a scalar function, with entry (i, j) equal to d2f/dxi dxj. It is symmetric wherever those second partials are continuous, by Clairaut's theorem.

open as a page

How do you derive the gradient of the loss L = ||Wx - y||^2 with respect to the matrix W?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Set the residual r = Wx - y, so L = r transpose r. Entry W_ij affects only r_i, and does so through x_j, so the gradient is 2 (Wx - y) x transpose, a matrix with the same shape as W.

open as a page

What is the second-order Taylor expansion of log(1 + x) about x = 0?

level: juniorimportance: should knowfreq 46%

basics

~10 s

log(1 + x) is approximately x - x^2/2 near x = 0. The dropped remainder starts at the cubic term x^3/3, so the error shrinks roughly eightfold each time x is halved.

open as a page

In the chain rule for f(g(h(x))), why is the total Jacobian the product J_f J_g J_h in that order?

level: middleimportance: should knowfreq 52%

basics

~20 s

Each factor must consume the output of the stage before it, so shapes conform only outermost-first: (m by q)(q by p)(p by n) yields the m by n total derivative. Matrix products are not commutative, so no other order works.

open as a page

When is the local quadratic model f(x0) + g^T d + 0.5 d^T H d a trustworthy stand-in for a loss surface?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Only near the base point, where the loss is smooth and curvature changes slowly. The model's error grows like the cube of the displacement, so it is exact only when the loss is itself quadratic.

open as a page

What is the Jacobian determinant of the polar map (r, theta) -> (r cos theta, r sin theta)?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

The determinant is r. The Jacobian is [[cos theta, minus r sin theta], [sin theta, r cos theta]], whose determinant is r cos^2 theta plus r sin^2 theta, which equals r. That factor is the map's local area scaling.

open as a page

How do you read the curvature of a surface along a direction v from its Hessian H?

level: middleimportance: nice to knowfreq 33%

basics

~10 s

Compute the scalar v^T H v for a unit vector v. That number is the second derivative of the function along the line in direction v: positive bends upward, near zero is locally flat.

open as a page