skip to content

How do you distinguish a long flat plateau from a true stationary point of an objective?

level: seniorimportance: nice to knowfreq 33%

answer

  1. exactly zero versus merely tiny
  2. watch the gradient over a window
  3. is there a consistent direction left
  4. take one very long step and look
  5. tolerances must be scale-relative

basics

~20 s

At a true stationary point the gradient is exactly zero. On a plateau it is merely small and still consistently signed, so the objective keeps creeping down and a long enough step leaves the flat region.

solid answer

~50 s

Geometrically the two are different: a stationary point has `grad f = 0`, while a plateau or ridge is a region where the gradient norm is small but nonzero. Numerically you can never observe exactly zero, so the practical test is not the magnitude alone but the *structure* of what remains. On a plateau the surviving gradient components keep the same signs across many evaluations and the objective decreases slowly but monotonically, so moving a long distance along that persistent direction produces a real drop. Near a genuine stationary point the residual gradient is noise-like: its direction wanders, and a longer move gains nothing. Two further checks help — probing curvature to see whether the surface bends at all, and making any tolerance scale-relative, since multiplying the objective by a constant multiplies every gradient by the same constant without changing the geometry.

go deeper

for a junior

Know the definition: stationary means the gradient is exactly zero, while a plateau just has a small one. Being able to say that a nearly flat region is not the same as an optimum is enough here.

for a middle

Explain why finite precision turns the exact condition into a tolerance test, and describe at least one check beyond the gradient norm, such as whether the objective is still falling over a window.

for a senior

Demonstrate a real diagnostic procedure: sign persistence over a window, a long probe step, curvature sampling, and scale-relative tolerances. Interviewers here want evidence you have been fooled by a plateau before.

for a principal

Own the stopping policy across a team: which criteria are combined, how tolerances are made scale-invariant, what gets logged so a stop can be audited afterwards, and how much budget is allowed before a run is cut off.

## The two geometries **A stationary point** is a single point where the gradient is exactly zero. **A plateau** is a whole region where the function is nearly flat: the gradient is small in norm but not zero, and the function is still descending, just very slowly. A **ridge** or valley floor is the anisotropic version — nearly flat along one direction while genuinely curved across it, so progress along the flat direction is slow while the perpendicular direction has already converged. The distinction matters because the two demand opposite responses. On a plateau there is still a direction that helps and patience or a longer step pays off. At a stationary point the first-order signal is exhausted and no amount of persistence in the same direction helps; if the point turns out to be a saddle you need curvature information to find the escape, and if it is a minimum you are done. ## Why the gradient norm alone is a weak test In finite precision you never see an exact zero, so "is the gradient zero" always becomes "is the gradient norm below some tolerance". That reduction loses two things. First, it loses **scale**. Replace `f` by `1000 * f`. Every gradient is multiplied by 1000 and every stationary point stays exactly where it was, because scaling by a positive constant cannot move an optimum. A fixed threshold like `1e-6` now fires in completely different places on the same problem. The same happens when you rescale the input variables. Any usable tolerance therefore has to be relative — compared against the magnitude of the objective, against the gradient norm at the start, or against the size of the variables — rather than an absolute number carried between problems. Second, it loses **direction**. A small gradient with consistent structure and a small gradient that is pure numerical noise carry completely different information, and a single norm collapses both to one number. ## Diagnostics that actually separate the two **Look at sign persistence.** Record the gradient over a window of evaluations. On a plateau the same components keep the same sign run after run: there is a genuine downhill direction, it is just shallow. Near a true stationary point the residual is dominated by numerical error, so component signs flip more or less arbitrarily and the average direction is close to nothing. **Take a long step and evaluate.** This is the most direct test available and it requires only function evaluations. Move a distance far larger than your usual step along the current negative-gradient direction and measure the objective. A meaningful drop means you were on a plateau and there was somewhere to go. No drop, in either direction along that line, is strong evidence you are at or near a stationary point. **Watch the objective, not only the gradient.** On a plateau the value declines slowly but persistently over a long window. At a stationary point it stops changing beyond the noise floor of the evaluation. **Probe curvature.** Sample the function along a few directions through the point and look at how it bends. A ridge shows near-zero curvature along one direction and clear curvature across it, which immediately explains why progress is slow without implying you have arrived. This also distinguishes a saddle, where some direction curves downward, from a minimum, where none does. ## Practical consequences Misreading a plateau as convergence is the expensive error: you stop early and report a point that was not optimal, and because the gradient really was small the stopping rule looked satisfied. Misreading a stationary point as a plateau is cheaper but wasteful: you keep spending evaluations on a direction that no longer exists. The safe stopping rule combines several signals rather than trusting one: a relative gradient tolerance, a relative change in the objective over a window, and a cap on total effort so a genuinely endless plateau terminates. State the tolerances relative to problem scale, log which criterion actually fired, and record enough of the run that the two cases can be told apart after the fact. Reporting only "converged" hides exactly the distinction this question is about.

  • Why is an absolute gradient-norm threshold a poor stopping criterion on its own?
    Because it is not scale-invariant. Multiplying the objective by 1000 multiplies every gradient by 1000 while leaving all optima exactly where they were, so a fixed threshold like `1e-6` fires somewhere completely different on the same problem. Rescaling the input variables has a similar effect. A usable criterion is relative to the objective magnitude, the initial gradient norm, or the variable scale.
  • What is the geometric difference between a plateau and a ridge?
    A plateau is nearly flat in every direction, so progress is slow whichever way you go. A ridge or narrow valley is flat only along one direction and clearly curved across it, so the curved directions converge quickly while the flat one drags. Probing curvature along a few directions separates them: a ridge shows a strong asymmetry, a plateau does not.
  • What is the cost of mistaking a plateau for convergence?
    You stop at a point that was still improving and report it as a solution, and the mistake is invisible because the stopping rule genuinely fired. The opposite error, treating a stationary point as a plateau, only wastes evaluations. That asymmetry argues for combining several stopping signals and logging which one triggered, rather than trusting a lone gradient threshold.

Walking a gently sloping salt flat feels the same underfoot as standing on a level floor. The way to tell is not to feel harder in one spot but to walk a long way and see whether your altitude changed.

saying these in an interview costs you the question

  • Treats any small gradient as a stationary point
  • Uses one absolute tolerance across differently scaled problems
  • Never looks at whether the objective is still decreasing
  • Ignores the direction of the residual gradient, only its norm
  • Assumes floating point will ever produce an exactly zero gradient

context