skip to content

What is the Hessian matrix of a scalar function, and when is it symmetric?

level: middleimportance: must knowfreq 72%

answer

  1. second derivatives, all pairs
  2. square grid, one entry per variable pair
  3. diagonal is per-axis curvature
  4. mixed partials commute when continuous
  5. Clairaut / Schwarz condition

basics

~20 s

The Hessian is the square matrix of all second partial derivatives of a scalar function, with entry (i, j) equal to d2f/dxi dxj. It is symmetric wherever those second partials are continuous, by Clairaut's theorem.

solid answer

~40 s

For a scalar function `f` of `n` variables, the Hessian `H` is the `n x n` matrix whose `(i, j)` entry is the second partial derivative `d2f/dxi dxj`. The diagonal holds the pure second derivatives `d2f/dxi^2` — curvature along each coordinate axis — and the off-diagonal entries hold the mixed partials, describing how the slope in one variable changes as another moves. Clairaut's (Schwarz's) theorem says the mixed partials commute, `d2f/dx dy = d2f/dy dx`, whenever the second partials are continuous near the point, which covers essentially every loss met in practice, so `H` is symmetric there. For `f(x, y) = x^2 * y^3`: `f_x = 2*x*y^3` gives `f_xy = 6*x*y^2`, and `f_y = 3*x^2*y^2` gives `f_yx = 6*x*y^2` — the same. Symmetry means only about half the entries are independent.

go deeper

for a junior

Be ready to state that the Hessian collects all second partial derivatives of a scalar function in a square matrix, and to compute a 2x2 one for a simple polynomial without dropping a factor.

for a middle

Explain the layout precisely: diagonal entries are per-axis second derivatives, off-diagonals are mixed partials, and Clairaut's theorem makes the matrix symmetric when the second partials are continuous.

for a senior

Show you know symmetry is a hypothesis, not a gift: name the continuity requirement, and explain the practical payoff — half the entries to compute or store, and a matrix that plays well with the standard tooling for symmetric matrices.

for a principal

Own the cost conversation: an n x n Hessian is quadratic in parameter count to form and store, so decide deliberately when full curvature information is worth it versus a diagonal or low-rank stand-in that keeps the second-order intuition at linear cost.

## What the Hessian is Take a scalar-valued function of several variables, `f(x1, x2, ..., xn)` — a loss as a function of parameters, say. Its **gradient** is the vector of first partial derivatives, `g = (df/dx1, ..., df/dxn)`, and it describes the local slope. Differentiating each of those first derivatives again with respect to each variable produces `n * n` second partial derivatives, and arranging them in a grid gives the **Hessian matrix**: ``` H[i][j] = d2f / (dxi dxj) ``` So `H` is `n x n`. The **diagonal** entries `H[i][i] = d2f/dxi^2` are pure second derivatives: how fast the slope along axis `i` is itself changing — the curvature of the 1-D slice you get by moving along that axis alone. The **off-diagonal** entries `H[i][j]` with `i != j` are **mixed partials**: how the slope in direction `i` changes as you move in direction `j`. A non-zero mixed partial means the two variables interact — the surface is tilted relative to the coordinate axes rather than being a clean sum of one-variable pieces. ## A worked example Let `f(x, y) = x^2 * y^3`. First derivatives: - `f_x = 2*x*y^3` - `f_y = 3*x^2*y^2` Second derivatives: - `f_xx = 2*y^3` - `f_yy = 6*x^2*y` - `f_xy = d/dy (2*x*y^3) = 6*x*y^2` - `f_yx = d/dx (3*x^2*y^2) = 6*x*y^2` So ``` H(x, y) = [[ 2*y^3 , 6*x*y^2 ], [ 6*x*y^2 , 6*x^2*y ]] ``` Two things are worth noticing. First, `f_xy` and `f_yx` came out identical even though the differentiations were done in opposite orders. Second, the Hessian is a **function of the point**, not a single fixed matrix: at `(1, 1)` it is `[[2, 6], [6, 6]]`, at `(1, 2)` it is `[[16, 24], [24, 12]]`. Curvature is local, exactly as slope is. ## Why symmetry holds The result that `d2f/dx dy = d2f/dy dx` is **Clairaut's theorem** (also called Schwarz's theorem). Its hypothesis matters: the mixed second partials must exist and be **continuous** in a neighbourhood of the point. Under that condition the order of differentiation is irrelevant and `H` is a symmetric matrix, `H = H^T`. It is not a triviality that can be assumed for free. There is a classic counterexample, `f(x, y) = x*y*(x^2 - y^2) / (x^2 + y^2)` with `f(0, 0) = 0`, whose two mixed partials at the origin exist but differ in sign; the reason is precisely that they fail to be continuous there. In practice, the losses and models you differentiate for optimisation are smooth enough on the region of interest that symmetry holds, and every treatment of curvature quietly relies on it. The honest interview answer is "symmetric when the second partials are continuous, which is the normal case", not "always symmetric". A useful intuition for *why* it should hold: the mixed partial measures the second difference of `f` across a small rectangle with corners `(x, y)`, `(x+h, y)`, `(x, y+k)`, `(x+h, y+k)` — namely `f(x+h, y+k) - f(x+h, y) - f(x, y+k) + f(x, y)`, divided by `h*k`. That expression is manifestly unchanged if you swap the roles of the two variables, since it just re-groups the same four corner values. Taking limits in either order lands on the same number when things are well behaved. ## Why practitioners care The Hessian is the object that turns a first-order picture of a surface into a second-order one. The gradient tells you which way is downhill and how steeply; the Hessian tells you how that steepness is changing, i.e. whether the surface is a gentle trough, a tight ravine, or nearly flat. It is the coefficient of the quadratic term in a Taylor expansion, so it is the smallest amount of extra information beyond the gradient that lets you approximate a surface by a bowl rather than a plane. Symmetry has practical consequences too. An `n x n` Hessian has `n^2` entries but only `n*(n+1)/2` distinct ones — for `n = 3` that is 6 numbers, not 9 — which halves both the storage and the number of second derivatives you actually have to compute or estimate. Symmetry also means the matrix behaves well under the standard machinery for symmetric matrices, which is what the linear-algebra treatment of curvature builds on. ## Dimensions and common slips Keep the shapes straight: for `f: R^n -> R`, the gradient is `n x 1` and the Hessian is `n x n`. The Hessian is *not* defined for a vector-valued function in this simple form — stacking second derivatives of a map `R^n -> R^m` gives a three-index object, not a matrix. And the Hessian is not the square of the gradient, nor the derivative of the function twice along one direction only; every pair of variables contributes an entry.

  • How many distinct numbers does an n-variable Hessian actually contain?
    n*(n+1)/2. The matrix has n^2 entries, but symmetry pairs H[i][j] with H[j][i], so you only need the diagonal (n entries) plus the upper triangle (n*(n-1)/2). For n = 3 that is 6 numbers rather than 9 — a real saving when each entry costs a second-derivative evaluation.
  • Is a Hessian a fixed matrix for a given function?
    No — it is a matrix-valued function of the point. For f(x, y) = x^2 * y^3 the Hessian at (1, 1) is [[2, 6], [6, 6]] and at (1, 2) it is [[16, 24], [24, 12]]. Curvature is local, just like slope, so any statement about a Hessian must name the point it is evaluated at.
  • What does a zero off-diagonal entry in a Hessian mean?
    It means the two variables do not interact to second order at that point: the slope along one is locally insensitive to moving along the other. If all off-diagonals vanish, the Hessian is diagonal and the surface is locally a sum of independent one-variable curvatures, one per coordinate axis.

The gradient is a topographic map's slope arrow at your feet; the Hessian is the description of how the ground bends around you — flat, bowl-shaped, or ridged — and its symmetry says the bend you feel walking north-then-east matches the one you feel walking east-then-north.

saying these in an interview costs you the question

  • Says the Hessian is always symmetric with no continuity condition
  • Confuses the Hessian with the gradient squared or its outer product
  • Treats the Hessian as a constant matrix rather than point-dependent
  • Thinks off-diagonal entries measure correlation between variables in data
  • Claims a Hessian exists in matrix form for vector-valued functions

context