When is the local quadratic model f(x0) + g^T d + 0.5 d^T H d a trustworthy stand-in for a loss surface?
answer
- value, slope and bend all matched
- one step beyond a linear model
- cubic remainder governs the radius
- exact only for a quadratic loss
- compare predicted against actual reduction
basics
~20 sOnly near the base point, where the loss is smooth and curvature changes slowly. The model's error grows like the cube of the displacement, so it is exact only when the loss is itself quadratic.
solid answer
~50 sThe model is the second-order Taylor expansion of the loss at `x0`: `f(x0 + d) ~ f(x0) + g^T d + 0.5 * d^T H d`, with `g` the gradient and `H` the Hessian there. It matches the true surface in value, slope and curvature at `x0`, and its error for a smooth loss grows like `||d||^3`. So it holds inside a **trust region** — a displacement small enough that the cubic remainder is negligible against the improvement predicted. It breaks in four ways: the step is too big; the loss has kinks, so no curvature exists there; curvature varies fast, so `H` at `x0` misdescribes the surface nearby; or `g` and `H` come from a noisy sample. The discipline is to propose a step, compare actual loss reduction with predicted, and shrink or grow the region accordingly.
go deeper
Recall the three pieces of the model — the loss value, the gradient term and the half-quadratic curvature term — and that it only describes the surface near the point it was expanded around.
Explain the error scaling: a smooth loss leaves a remainder of order the cube of the displacement, so the model is exact for a quadratic loss and degrades fast as steps grow.
Demonstrate the operating loop: propose a displacement, compare predicted with actual loss change, and shrink or grow the trust region on that evidence rather than assuming a fixed safe radius.
Own the modelling call: decide when second-order structure earns its cost at all, given non-smooth loss components, noisy curvature estimates from sampled data, and the engineering weight of maintaining curvature machinery.
## What the model is Write the loss near a base point `x0` as a function of the displacement `d`. The second-order Taylor expansion is ``` m(d) = f(x0) + g^T d + 0.5 * d^T H d ``` where `g` is the gradient at `x0` (an `n`-vector) and `H` is the Hessian at `x0` (an `n x n` symmetric matrix). Three terms, three roles: the constant pins the height, the linear term `g^T d` matches the slope in every direction, and the quadratic term `0.5 * d^T H d` matches the curvature in every direction. It is the simplest model that can represent a bowl rather than a plane, and it is the reason the Hessian is worth computing at all. Note the `0.5`. It is the `1/2!` from the Taylor formula, and omitting it is the single most common algebra slip here — it doubles all the curvature the model believes in. ## The error term is the whole story For a loss with continuous third derivatives, the difference between `m(d)` and the true `f(x0 + d)` is of order `||d||^3`. Two consequences follow directly. **Locality.** Halve the displacement and the model error falls roughly eightfold. Double it and the error grows roughly eightfold. Accuracy is a steep function of how far you move, which is why every use of a quadratic model comes bundled with an explicit or implicit radius. **Exactness in the pure case.** If the loss is itself exactly quadratic, all third and higher derivatives vanish and the model is not an approximation at all — it reproduces the surface everywhere. Ordinary least-squares loss is the canonical case. This is the extreme end of the spectrum, and it is why quadratic reasoning is so much more reliable for linear-model losses than for deeply nonlinear ones. ## The four ways it goes wrong **1. The displacement is too large.** The cubic remainder is not a rounding error once `||d||` leaves the neighbourhood where third-order effects are small. The model may then predict a decrease where the surface actually rises. **2. The loss is not twice differentiable there.** Piecewise-linear components, hard clipping, absolute-value penalties and hinge-shaped losses have kinks; at a kink the second derivative does not exist, and a model built from a one-sided curvature can be badly wrong across the kink even for a tiny step. Regions containing a kink are not a smooth-model regime at all. **3. Curvature varies quickly.** The model uses `H` at `x0` for the whole neighbourhood. If curvature changes sharply over the distance you intend to move — a narrow ravine, a rapidly tightening bowl — then `H(x0)` is a poor summary of the surface even a short way off, and the third-derivative magnitude, which controls exactly this, is large. **4. `g` and `H` are estimated, not exact.** When the loss is an average over a sample and you evaluate on a subsample, both the gradient and the curvature carry sampling noise. Curvature estimates are typically noisier than gradient estimates because second differences amplify noise, so a model built from a small batch can confidently describe a bowl that is largely an artefact of which examples were drawn. ## The operational fix: a trust region Rather than trying to prove in advance how far the model is valid, the practical discipline measures it. Pick a radius, choose a displacement `d` within it, and compare: - the **predicted** change `m(d) - f(x0)`, which is `g^T d + 0.5 * d^T H d`; - the **actual** change `f(x0 + d) - f(x0)`, one extra loss evaluation. Their ratio is a direct measurement of how well the model is describing the surface at this scale. A ratio near 1 says the quadratic picture is accurate over that displacement — accept the step and consider allowing a larger radius. A ratio far below 1, or negative, says the model overpromised — reject or shrink. This loop needs no assumption about third derivatives; it observes the discrepancy the third derivative causes. It is also self-correcting as the surface changes character during a long run. ## Sanity checks worth naming - At `d = 0` the model returns `f(x0)` exactly; if your implementation does not, the constant term or the sign convention is wrong. - The predicted change is a scalar; a shape error in `d^T H d` usually shows up as a vector or matrix appearing where a number belongs. - The model is defined relative to a base point. Quoting a quadratic model without saying where it was expanded is meaningless, since both `g` and `H` are point-dependent. - The model is only as smooth as the loss. If you have deliberately introduced non-smooth structure, expect second-order reasoning to be locally invalid rather than merely imprecise. ## Why interviewers ask it The question separates candidates who have memorised the Taylor formula from those who know what it buys. The formula is a line of algebra; the judgment is knowing that its validity radius is an empirical quantity, that the cubic error term is the thing you are trading against, and that measuring the predicted-versus-actual ratio is cheaper and more honest than reasoning about third derivatives you cannot compute.
- For which loss is this quadratic model exact rather than approximate?An exactly quadratic loss, such as ordinary least squares in the parameters. All third and higher derivatives vanish, so the remainder is identically zero and the model reproduces the surface at any displacement, not just nearby. The Hessian is then a constant matrix, independent of where you expand.
- How would you empirically check whether the model is valid at your current step size?Compute the predicted change g^T d + 0.5 * d^T H d, take the step, and measure the actual change in loss. The ratio of actual to predicted is a direct readout: near 1 means the quadratic picture holds at that scale, well below 1 or negative means the displacement has outrun the model and the radius should shrink.
- Why can a minibatch-estimated Hessian make the model misleading even at small displacements?Because the model is then describing the sampled subset's curvature, not the population loss's. Second-order estimates are noisier than first-order ones, so a small batch can suggest a well-defined bowl that disappears on the next batch. Averaging curvature information over more data, or over time, is the usual mitigation.
- What does the 0.5 factor in the quadratic term come from?The 1/2! in the Taylor formula: the k-th order term carries 1/k!. Dropping it doubles every curvature the model represents, which makes the surface look twice as tightly bowled as it is and systematically distorts any quantity derived from the model.
saying these in an interview costs you the question
- Treats the quadratic model as globally valid
- Drops the 0.5 factor on the quadratic term
- Ignores that the loss must be twice differentiable
- Never checks predicted reduction against actual reduction
- Forgets the model is tied to a specific base point