Why is the origin a saddle point of f(x, y) = x^2 - y^2 rather than a minimum?
answer
- the gradient does vanish there
- check each direction separately
- up along one axis, down the other
- a minimum in x, a maximum in y
basics
~20 sThe gradient of x^2 - y^2 vanishes at the origin, but the surface curves up along the x-axis and down along the y-axis. A point that is a minimum one way and a maximum another is a saddle.
solid answer
~50 sThe gradient is `(2x, -2y)`, which is zero at the origin, so the origin is stationary and passes the first-order test. Now restrict the function to each axis. Along `y = 0` it becomes `f(x, 0) = x^2`, which has a strict minimum at the origin. Along `x = 0` it becomes `f(0, y) = -y^2`, which has a strict maximum there. So arbitrarily close to the origin there are points with larger values and points with smaller values, which rules out both a local minimum and a local maximum. That mixed curvature — up along some directions, down along others — is the definition of a **saddle point**. Saddles are the reason a stationary point in many variables is usually not an optimum: a minimum needs upward curvature in *every* direction, and one bad direction is enough to break it.
go deeper
Be able to compute the two partial derivatives, show they vanish at the origin, and then say what happens along each axis. Producing the slices x^2 and -y^2 is the whole answer at this level.
Explain what a saddle is in general terms, not just for this function: a stationary point with increase along one direction and decrease along another. Expect a follow-up asking why one bad direction is enough.
Demonstrate awareness that stalled high-dimensional searches usually sit near saddles rather than minima, and describe how you would probe for a descent direction when the gradient itself has gone quiet.
Own the decision of how much effort is worth spending to detect and escape saddles versus restarting from elsewhere, and be able to justify that call in terms of compute budget and the shape of the problem.
## Checking the first-order condition For `f(x, y) = x^2 - y^2` the partial derivatives are `df/dx = 2x` and `df/dy = -2y`. Both vanish exactly when `x = 0` and `y = 0`, so the origin is the unique stationary point of this function. Every first-order test is satisfied there, and if a zero gradient certified optimality the story would end. It does not. ## Restricting to lines through the point The fastest way to see what is happening is to walk away from the origin along two different directions and watch the value. Along the `x`-axis, set `y = 0`. The function becomes `f(x, 0) = x^2`, which equals 0 at the origin and is strictly positive everywhere else on that line. Restricted to this slice, the origin is a **strict minimum**. Along the `y`-axis, set `x = 0`. The function becomes `f(0, y) = -y^2`, which equals 0 at the origin and is strictly negative everywhere else on that line. Restricted to this slice, the origin is a **strict maximum**. Put those together. In every neighbourhood of the origin, however small, there are points where `f > 0` and points where `f < 0`, while `f(0, 0) = 0`. A local minimum requires no nearby point to be lower; the `y`-axis violates that. A local maximum requires no nearby point to be higher; the `x`-axis violates that. So the origin is neither, and the standard name for a stationary point with this mixed behaviour is a **saddle point** — the surface really does look like a riding saddle or a mountain pass, rising in one pair of opposite directions and falling in the perpendicular pair. ## The general definition A saddle point is a stationary point that is not a local extremum because the function increases along at least one direction leaving the point and decreases along at least one other. Note that the two directions do not have to be coordinate axes; they are just the easiest slices to compute here. The second-order test in several variables makes the same statement quantitatively: strictly positive curvature along every direction gives a strict local minimum, strictly negative along every direction gives a strict local maximum, and both signs present gives a saddle. ## Why saddles dominate in high dimensions This is the part with practical bite. To be a local minimum, a stationary point in `d` variables must curve upward along **all** `d` independent directions. To be a saddle it needs only one direction curving down. As `d` grows, the first requirement gets harder to satisfy while the second stays trivially easy, so among the stationary points of a complicated high-dimensional objective, saddles vastly outnumber local minima. A stalled search in a large parameter space is therefore far more likely to have found a saddle than a genuine minimum — and unlike a minimum, a saddle has an escape route: there exists a direction along which the value strictly decreases, even though the gradient at the point itself is zero. ## When even mixed curvature is not the story Curvature can also vanish along a direction, and then no second-order test can classify the point at all. The **monkey saddle** `f(x, y) = x^3 - 3xy^2` is the standard illustration. Its partials are `df/dx = 3x^2 - 3y^2` and `df/dy = -6xy`, both zero at the origin, so the origin is stationary. But every second partial derivative — `6x`, `-6x` and `-6y` — also vanishes at the origin, so the second-order information is entirely empty and the test is inconclusive. Looking at the function directly settles it: along `y = 0` it reduces to `x^3`, which increases through the origin, so the origin is not an extremum. The surface has three regions going up and three going down rather than the usual two and two, which is where the name comes from. ## What an interviewer is listening for A complete answer does three things: verifies the gradient is zero, exhibits one direction of increase and one of decrease with explicit restricted functions, and then names the conclusion. Candidates who assert "it is a saddle because the curvatures have different signs" without ever producing the two slices are reciting a rule; candidates who produce `x^2` and `-y^2` have actually demonstrated it.
- Why do saddle points become more common than local minima as the number of variables grows?A local minimum needs upward curvature along every one of the `d` directions, while a saddle needs only one direction curving downward. Satisfying `d` simultaneous conditions gets rapidly harder as `d` grows, whereas breaking one gets easier. So among the stationary points of a high-dimensional objective, saddles greatly outnumber minima.
- How does the monkey saddle f(x, y) = x^3 - 3xy^2 defeat the second-order test at the origin?Its first partials `3x^2 - 3y^2` and `-6xy` vanish at the origin, so it is stationary. But all of its second partials, `6x`, `-6x` and `-6y`, also vanish there, so second-order information is completely empty and no curvature-based test can classify it. Restricting to `y = 0` gives `x^3`, which increases through the origin, so it is not an extremum.
- If a search stops at a saddle point, is there still a direction that lowers the objective?Yes. By definition a saddle has at least one direction along which the function strictly decreases as you move away from it. The gradient is zero at the point itself, so first-order information gives no hint about which direction that is, but the descent direction genuinely exists — which is why a saddle is an escapable stopping place rather than a dead end.
A mountain pass is the lowest point on the ridge you are crossing and the highest point on the trail between the two valleys. Standing there, you can go down in one direction and up in another, so it is neither a summit nor a valley floor.
saying these in an interview costs you the question
- Calls the origin a minimum because the gradient is zero
- Says a saddle is just a very flat minimum
- Only checks one direction before concluding
- Assumes saddles are rare in high-dimensional problems
- Claims a saddle offers no descent direction at all