How do you read the curvature of a surface along a direction v from its Hessian H?
answer
- restrict the surface to a line
- one-variable function of a step size
- chain rule applied twice
- a scalar quadratic form, not a vector
- normalise the direction first
basics
~10 sCompute the scalar v^T H v for a unit vector v. That number is the second derivative of the function along the line in direction v: positive bends upward, near zero is locally flat.
solid answer
~50 sRestrict the function to a line through the base point: `phi(t) = f(x0 + t*v)`. Differentiating twice with the chain rule gives `phi''(0) = v^T H v`, where `H` is the Hessian at `x0`. So the quadratic form `v^T H v` is the curvature of the 1-D slice you get by walking in direction `v`, provided `v` is normalised — otherwise it is scaled by `||v||^2`, since the form is quadratic in `v`. Take `f(x, y) = x^2 + 100*y^2`, whose Hessian is the constant diagonal matrix `diag(2, 200)`. Along `v = (1, 0)` the curvature is `2`; along `v = (0, 1)` it is `200`. Same surface, same point, curvature differing by a factor of 100 — the bowl is gentle in one direction and steep-walled in the other. Along `v = (1, 1)/sqrt(2)` you get `0.5*2 + 0.5*200 = 101`.
go deeper
Know that curvature in several variables depends on the direction you move, and that for a diagonal Hessian the diagonal entries give the curvature along each coordinate axis.
Derive v^T H v by restricting the function to the line x0 + t*v and differentiating twice, and state the shape argument that makes the result a scalar.
Show the operational angle: probe curvature along a chosen direction using only the product of the Hessian with that vector, so you get second-order information without ever forming an n x n matrix.
Decide how much curvature information a system should carry: a handful of directional probes, a diagonal approximation, or the full matrix — each trades fidelity against cost, and the anisotropy of the actual surface is what should settle it.
## Curvature is a direction-dependent quantity In one variable, curvature at a point is a single number: `f''(x)`. In several variables it is not — how sharply the surface bends depends on which way you walk. The Hessian is the compact object that answers the question for every direction at once, and the extraction rule is a quadratic form. ## The slice construction Fix a base point `x0` and a direction vector `v`. Walking along that direction traces the 1-D function ``` phi(t) = f(x0 + t*v) ``` This is an ordinary single-variable function of `t`, so its second derivative is an ordinary curvature. Differentiating once by the chain rule gives `phi'(t) = g(x0 + t*v)^T v`, where `g` is the gradient. Differentiating again gives ``` phi''(t) = v^T H(x0 + t*v) v ``` and at `t = 0`, ``` phi''(0) = v^T H v ``` with `H` the Hessian at `x0`. So **`v^T H v` is the second derivative of `f` along `v`**. It is a single number, not a vector or matrix: `v` is `n x 1`, `H` is `n x n`, and `v^T H v` collapses to `1 x 1`. ## Normalisation matters The form is quadratic in `v`: replacing `v` by `2*v` multiplies `v^T H v` by 4. That is correct behaviour — walking twice as fast along the same line makes `phi` change four times as fast to second order — but it means the number is only a **curvature** in the geometric sense when `v` is a unit vector. Always say "for unit `v`" when quoting `v^T H v` as curvature, or divide by `||v||^2`. ## A concrete anisotropic bowl Let `f(x, y) = x^2 + 100*y^2`. The gradient is `(2*x, 200*y)` and the Hessian is the constant matrix ``` H = [[2, 0], [0, 200]] ``` (constant because the function is exactly quadratic; in general `H` varies with position). Reading curvature off it: - Along `v = (1, 0)`: `v^T H v = 2`. - Along `v = (0, 1)`: `v^T H v = 200`. - Along `v = (1, 1)/sqrt(2)`: `v^T H v = (1/2)*2 + (1/2)*200 = 101`. - Along `v = (cos a, sin a)`: `2*cos^2(a) + 200*sin^2(a)`, which sweeps smoothly between 2 and 200 as the angle turns. The picture is an elongated bowl: cross-sections are ellipses stretched along `x`, because it takes a hundred times more movement in `x` than in `y` to gain the same height. This is what *anisotropic curvature* means — the surface has no single curvature scale, and any statement like "the surface is steep here" is incomplete without naming a direction. ## Reading a diagonal Hessian, and reading a non-diagonal one When `H` is diagonal, the diagonal entries *are* the curvatures along the coordinate axes: `H[i][i] = e_i^T H e_i`, since picking the `i`-th standard basis vector selects that entry. That is the easy case, and it is the reason a diagonal Hessian is so readable — the variables do not interact, and each axis carries its own independent curvature. When `H` has non-zero off-diagonals, the coordinate axes are no longer the natural directions of the surface, and the diagonal entries are still the axis curvatures but they no longer bracket the range of curvature over all directions. For example, `H = [[1, 3], [3, 1]]` has axis curvatures of `1` and `1`, yet along `(1, 1)/sqrt(2)` the curvature is `(1/2)*(1 + 3 + 3 + 1) = 4` and along `(1, -1)/sqrt(2)` it is `-2`. The extreme directions of a general symmetric matrix are found by matrix decomposition, which is a separate body of machinery; the point to hold here is simply that the diagonal alone can badly understate the spread. ## Sign and what it does and does not tell you A positive `v^T H v` means the slice along `v` bends upward — move either way along that line and, to second order, the function rises relative to the tangent line. Negative means it bends downward. Near zero means the slice is locally straight to second order, and the third-order behaviour decides what actually happens; this is the flat-valley case where second-order information is nearly uninformative. One discipline point: `v^T H v` is a statement about a **single direction at a single point**, nothing more. It says nothing on its own about where you are on the surface, whether the gradient vanishes, or what kind of point this is — those are separate questions with their own machinery. Keeping curvature-along-a-direction distinct from those questions is what stops the sign of one quadratic form from being over-interpreted. ## Practical use The quadratic form is how you probe curvature cheaply. Evaluating `v^T H v` needs only the effect of `H` on one vector, not the whole matrix, so for a large parameter space you can ask "how curved is the loss along *this* direction?" — for example along the current descent direction — without ever forming an `n x n` object. That makes directional curvature the workable version of second-order information at scale.
- What happens to v^T H v if you double the length of v?It quadruples. The expression is quadratic in v, so scaling v by c scales the value by c^2. That is why the quantity is only a geometric curvature when v is a unit vector; otherwise divide by ||v||^2 to compare directions fairly.
- Why is the shape v^T H v a scalar rather than a vector?Dimensions: v^T is 1 x n, H is n x n, and v is n x 1, so the product is 1 x 1. Intuitively it must be a scalar, because it is the second derivative of the one-variable function you get by restricting f to a line — and a one-variable second derivative is a single number.
- Can you get directional curvature without forming the full Hessian?Yes. v^T H v only needs the product H*v, which is the directional derivative of the gradient along v — obtainable by differentiating the gradient once more in that one direction. For a large parameter vector this avoids ever materialising an n x n matrix, so probing curvature along a few chosen directions stays affordable.
saying these in an interview costs you the question
- Reports H*v instead of v^T H v as the curvature
- Quotes v^T H v as curvature without normalising v
- Assumes diagonal entries bracket curvature in every direction
- Thinks curvature is one number for a multivariable surface
- Treats a constant Hessian as the general case