For a map f from R^n to R^m, what shape is the Jacobian and what is entry (i, j)?
answer
- outputs by inputs
- reconstruct it from J times d
- rows are component gradients
- entry is one partial derivative
- it is a function of the point
basics
~20 sThe Jacobian is m by n: one row per output component, one column per input variable. Entry (i, j) is the partial derivative of output component f_i with respect to input x_j, evaluated at a specific point.
solid answer
~50 sFor `f: R^n -> R^m` the Jacobian `J` is an m by n matrix whose entry (i, j) is `d f_i / d x_j`, the partial derivative of the i-th output with respect to the j-th input. Row i is the gradient of the scalar component `f_i` written as a row. The shape is forced by what the Jacobian is for: it is the best linear approximation of f near a point, `f(x + d) ~= f(x) + J d`, and for `J d` to make sense with d in R^n and the result in R^m, J must be m by n. Two consequences matter in practice: J depends on the point, so it must always be written as `J(x)`, and when the output is a single scalar (m = 1) the Jacobian is a 1 by n row vector whose transpose is the gradient column.
go deeper
Be ready to state the shape as outputs-by-inputs and to write entry (i, j) as the partial derivative of output i with respect to input j for a small explicit map.
An interviewer expects you to justify the shape from the linear approximation f(x + d) is about f(x) + J d, rather than reciting it, and to compute a 2 by 2 Jacobian on the spot.
Demonstrate shape checking as a working habit: catching dropped transposes and mismatched layouts by dimension alone, before evaluating anything numerically.
Own the convention call. Decide numerator or denominator layout for a codebase or a team's notation, state it once, and explain what class of bug the inconsistency otherwise produces.
## Definition Let `f: R^n -> R^m` be differentiable, with input `x = (x_1, ..., x_n)` and output components `f_1(x), ..., f_m(x)`. The Jacobian matrix of f at a point x is `J(x)[i][j] = d f_i / d x_j` evaluated at x and it has m rows and n columns. Say it as outputs-by-inputs: the row index runs over the things being produced, the column index over the things being varied. ## Why m by n and not n by m The defining property of the derivative is linear approximation. Near x, `f(x + d) ~= f(x) + J(x) d` where d is a small displacement in the input space, so d is a column in R^n, and the left-hand side lives in R^m. A matrix that eats an n-vector and returns an m-vector has m rows and n columns. That is the whole argument, and it is the one to give in an interview: the shape is not a convention to memorise, it is whatever makes the product conformable. If you ever cannot remember the orientation, reconstruct it from `J d` in half a second. Row i of J is `(d f_i / d x_1, ..., d f_i / d x_n)` — the gradient of the scalar-valued component f_i laid out horizontally. Column j is the vector of how every output responds to nudging the single input x_j. ## A worked example Take `f: R^2 -> R^2` with `f(x_1, x_2) = (x_1^2 * x_2, x_1 + sin(x_2))`. Then `J = [[2 * x_1 * x_2, x_1^2], [1, cos(x_2)]]` Row 1 differentiates `x_1^2 * x_2` with respect to x_1 and then x_2; row 2 does the same for `x_1 + sin(x_2)`. Here n = m = 2, so the matrix is square — a special case, not the general one. A second example fixes the shape rule in memory: `f: R^5 -> R^3` given by `f(x) = A x + b` with A of shape 3 by 5. Every partial derivative is a constant entry of A, so `J(x) = A` everywhere. An affine map is its own linearization, and its Jacobian is exactly the matrix of the map — 3 by 5, m by n, as required. ## The scalar-output special case When m = 1, f is a scalar field such as a loss, and the Jacobian is a single row of length n. The gradient is conventionally a column vector, so `gradient = J^T`. Keeping these straight avoids a lot of transposition pain: Jacobians are wide (outputs by inputs), gradients of scalars are tall. Many disagreements about a stray transpose are really disagreements about which of these two objects someone wrote down. ## It is a function of the point The Jacobian is not one matrix attached to f; it is a matrix-valued function of x. In the worked example above, J at (1, 0) is `[[0, 1], [1, 1]]` while J at (0, 1) is `[[0, 0], [1, cos(1)]]`. Any statement about a Jacobian that does not say where it is evaluated is incomplete, and forgetting the evaluation point is the error that ruins multi-stage chain-rule computations, where each stage is evaluated at a different intermediate value. ## Layout conventions What is described here is the numerator layout: outputs index rows. Some texts use the transposed denominator layout, in which the same object appears as n by m. Neither is wrong, but mixing them inside one derivation produces transposes that appear from nowhere. The practical discipline is to declare the convention once, then verify every result by shape: if a quantity is a derivative of something m-dimensional with respect to something n-dimensional, and your expression is not m by n, something is wrong before any number is computed. ## Why it matters downstream Every composed derivative is built from these blocks. Once the shape rule is automatic, checking a long derivation reduces to checking that adjacent matrices are conformable and that the ends match the overall input and output dimensions. That single habit catches most sign-free algebra errors without evaluating anything. ## Answering well A strong answer states the shape, gives the entry formula, justifies the shape from the linear approximation rather than from memory, notes that rows are component gradients, and adds that the Jacobian is evaluated at a point. Mentioning the m = 1 case and the transpose relationship to the gradient shows the fluency an interviewer is listening for.
- What is the Jacobian of the affine map f(x) = A x + b, where A is 3 by 5?It is A itself, at every point, shape 3 by 5. Each output is a fixed linear combination of the inputs, so every partial derivative is a constant entry of A, and the constant offset b differentiates away. An affine map is already its own linear approximation, which is why the Jacobian does not vary with x.
- How does the Jacobian relate to the gradient when the output is a single scalar loss?With m = 1 the Jacobian is a 1 by n row vector of partial derivatives, and the gradient is its transpose, an n by 1 column. Same numbers, different orientation. Fixing that convention up front is what prevents stray transposes appearing halfway through a derivation.
- You wrote a derivative expression and it is n by m instead of m by n. What does that tell you?Either a transpose was dropped, or you have silently switched to the denominator layout convention partway through. Check the shape against the rule outputs-by-inputs and against any product the expression must fit into. Shape checking catches this class of error before you evaluate a single number.
Think of a mixing desk: one row per output channel, one column per input fader, and each knob setting says how much that fader moves that channel.
saying these in an interview costs you the question
- Says the Jacobian is n by m for a map from R^n to R^m
- Calls the Jacobian a fixed matrix rather than a function of x
- Confuses a column of the Jacobian with a component gradient
- Cannot state entry (i, j) as one partial derivative
- Treats the scalar-output Jacobian and the gradient as identical in shape