Why is a zero gradient necessary but not sufficient for a local minimum?
answer
- first-order information only
- maxima have horizontal tangents too
- necessary, not sufficient
- think about x cubed at zero
basics
~20 sAt a smooth interior local minimum the gradient must vanish, so a zero gradient is necessary. But maxima, saddle points and flat inflection points also have a zero gradient, so vanishing slope on its own certifies nothing.
solid answer
~40 sIf `f` is differentiable and `c` is an interior point of the domain, then `c` being a local minimum forces `grad f(c) = 0`: if any partial derivative were nonzero you could take a tiny step in the downhill direction and lower `f`, contradicting minimality. That is the first-order necessary condition, and points satisfying it are called stationary (or critical) points. It is not sufficient, because the same condition is satisfied by local maxima, by saddle points, and by flat inflection points such as `x = 0` on `f(x) = x^3`, where the tangent is horizontal but the function keeps increasing through the point. So a zero gradient only narrows the search to a candidate list; classifying each candidate needs second-order (curvature) information or a direct comparison of nearby values.
go deeper
Be ready to state the condition and immediately give a counterexample showing it is not sufficient. Knowing that maxima and the flat point of a cubic both have zero derivative is enough to pass this screen.
Explain why the condition holds: a nonzero slope gives you a downhill direction, which contradicts minimality. Also name the assumptions it needs, namely differentiability and an interior point.
Show how this changes what you do with a solver that has stopped. Argue that a converged search has found a stationary point and nothing more, and describe how you would check whether it is worth trusting.
Frame the cost of the gap. Decide when it is worth spending compute to certify or improve a stationary point versus accepting it, and make that policy explicit rather than letting each team invent its own stopping rule.
## The first-order necessary condition Let `f` be a real-valued function that is differentiable at a point `c` lying strictly inside its domain. The derivative `f'(c)` in one variable, or the gradient `grad f(c)` — the vector of all first partial derivatives — measures how fast `f` changes as you move away from `c` in each coordinate direction. The claim is: **if `c` is a local minimum, then `grad f(c) = 0`.** The argument is short. Suppose some component of the gradient is nonzero, say the rate of change along direction `d` is `g != 0`. Then for small `t`, `f(c + t*d)` is approximately `f(c) + t*g`. Choose the sign of `t` so that `t*g < 0` and the value strictly drops. That contradicts `c` being a point that no nearby point beats. The same reasoning run with the inequality reversed shows a local maximum also forces a zero gradient. Points where the gradient vanishes are called **stationary points**; the broader term **critical point** is often used interchangeably, and in some texts additionally covers points where the derivative fails to exist. ## Why it is only necessary "Necessary" means every interior smooth minimum is on the list. It does not mean everything on the list is a minimum. A zero gradient says the tangent line (or tangent plane) is horizontal, and horizontal tangents occur in several distinct geometric situations: - **A local minimum.** The surface curves upward away from the point in every direction. - **A local maximum.** It curves downward in every direction. - **A saddle point.** It curves up along some directions and down along others, so the point is a minimum along one slice and a maximum along another. - **A flat inflection point.** The canonical one-variable example is `f(x) = x^3` at `x = 0`. Here `f'(x) = 3x^2`, so `f'(0) = 0`; but `f` is strictly increasing on the whole real line, so `x = 0` is neither a maximum nor a minimum. The curve just momentarily flattens as it passes through. Because all four cases satisfy the same first-order equation, solving `grad f = 0` produces a **candidate list**, not an answer. Deciding among the cases is exactly what second-order (curvature) tests are for, and when curvature also vanishes even those tests can come up empty. ## The conditions on the statement matter Two qualifiers do real work. **Differentiable.** If `f` has a kink, the gradient need not exist at the optimum. `f(x) = |x|` has its global minimum at `x = 0`, and no derivative there at all. Solving `f'(x) = 0` would find nothing, because the minimum is not a stationary point in the ordinary sense. **Interior.** On a restricted domain the optimum can sit on the boundary, where you cannot step in both directions. Minimising `f(x) = x` on the closed interval `[0, 1]` gives the answer `x = 0`, yet `f'(0) = 1`. The first-order condition simply does not apply at a boundary point, and any search that only solves `grad f = 0` will miss it. Constrained problems replace the condition with a different stationarity statement of their own. ## What this means in practice When you set derivatives to zero and solve, you are enumerating suspects. The workflow is: find every stationary point, add the boundary and the non-differentiable points to the list, then classify or compare. A candidate answer that stops at "the derivative is zero here, so this is the minimum" is incomplete, and an interviewer will usually probe it with `x^3`. The same logic explains a familiar diagnostic behaviour: a downhill search that has stopped moving has, at best, reached a point where the gradient is (numerically) zero. That alone tells you it is stationary. It does not tell you whether it is a minimum, and it certainly does not tell you it is the best minimum available.
- Can a minimum occur at a point where the gradient is not zero?Yes, in two situations. At a boundary point of the feasible region: `f(x) = x` on `[0, 1]` is minimised at `x = 0` where `f'(0) = 1`. And at a point where `f` is not differentiable: `f(x) = |x|` is minimised at `x = 0`, where no derivative exists. The first-order condition assumes an interior, differentiable point, and both of these violate that assumption.
- What is the difference between a stationary point and a critical point?For a function that is differentiable everywhere the two terms are used interchangeably: both mean the gradient is zero. `Critical point` is the broader word — many texts also count points where the derivative fails to exist, such as the kink of `|x|` at the origin. Using the broader definition is safer when you are enumerating optimum candidates, because it captures kinks as well.
- Given a stationary point, what information do you need next to classify it?Curvature. You need to know how the function bends away from the point: upward in every direction means a local minimum, downward in every direction a local maximum, and up in some directions and down in others a saddle. That is second-order information. Failing that, you can simply evaluate `f` at nearby points on both sides and compare directly.
A flat spot on a hiking trail could be a valley floor, a summit, a mountain pass, or just a level stretch on a continuous climb. Feeling level underfoot does not tell you which.
saying these in an interview costs you the question
- Claims a zero derivative proves the point is a minimum
- Forgets that local maxima also have zero gradient
- Never mentions saddle points or flat inflection points
- Ignores boundary optima and non-differentiable kinks
- Jumps from local minimum straight to global minimum