Why does the gradient of a differentiable function point in the direction of steepest ascent?
answer
- compare all directions at once
- rate in a direction is a fixed vector combined with u
- the angle is the only thing you control
- cos(theta) is capped at 1
- equality when the bearing aligns with the gradient
basics
~20 sThe rate of change in a unit direction equals the gradient's length times the cosine of the angle between them. Cosine peaks at 1 when the direction aligns with the gradient, so that bearing rises fastest.
solid answer
~50 sThe rate at which `f` changes as you step in a unit direction `u` — the directional derivative — is `grad f . u`, which equals `|grad f| * cos(theta)` where `theta` is the angle between `u` and the gradient. Since `cos(theta)` is at most 1 and hits 1 only when `theta = 0`, the rate is maximised by choosing `u` pointing the same way as `grad f`. So the gradient is the direction of steepest ascent, and the maximum rate itself is `|grad f|`. Turn 180 degrees and you get `-|grad f|`, the steepest descent. Directions at right angles to the gradient give zero first-order change. Two caveats worth saying out loud: this is a *local, first-order* claim — it names the best infinitesimal bearing, not a route to the maximum — and it needs `f` differentiable with a nonzero gradient at that point.
go deeper
Recall the two headline facts: the gradient points uphill fastest, and its negative points downhill fastest. Be able to say it in one sentence without hedging.
Derive it. Rate in a unit direction equals the gradient's length times cos(theta), cos(theta) tops out at 1, equality when the bearing aligns with the gradient, and the maximum rate is the gradient's own length.
Demonstrate the caveats you have been bitten by: the claim is local and first-order, it needs a nonzero gradient, and 'steepest' shifts the moment you rescale the input variables.
Own the framing that a steepest direction is only defined once you have chosen how to measure distance in input space. Choosing that scaling is a modelling decision with real consequences, not a preprocessing detail.
## The setup Stand at a point `a` on the surface described by a differentiable function `f`. You may step in any direction. Which bearing takes you uphill fastest? Answering that precisely is what makes the gradient more than a bookkeeping device for partial derivatives. ## Rate of change in a chosen direction Fix a direction as a vector `u` of length one. The **directional derivative** of `f` at `a` along `u` is `D_u f(a) = lim_{h -> 0} [ f(a + h*u) - f(a) ] / h` in words: the rate at which `f` changes per unit of distance travelled along that bearing. Insisting that `u` have length one is what makes it a rate *per unit distance* rather than per unit of whatever the arrow happened to measure. For a differentiable `f`, this limit has a closed form: it is the gradient combined componentwise with the direction, `D_u f(a) = grad f(a) . u = (df/dx1) * u1 + (df/dx2) * u2 + ...` That identity is the whole engine. Every direction's rate is read off one fixed vector — the gradient — rather than recomputed from scratch. ## The maximisation argument Write the combination geometrically. For any two vectors, `v . u = |v| * |u| * cos(theta)`, where `theta` is the angle between them. With `|u| = 1`, `D_u f(a) = |grad f(a)| * cos(theta)` Now the answer is immediate. `|grad f(a)|` is a fixed nonnegative number at this point — it does not depend on your choice of bearing. The only thing you control is `cos(theta)`, which ranges over `[-1, 1]`. Therefore: - **Maximum** rate `+|grad f(a)|`, attained at `theta = 0`: `u` points along the gradient. This is steepest ascent. - **Minimum** rate `-|grad f(a)|`, attained at `theta = 180` degrees: `u` points along the negative gradient. This is steepest descent. - **Zero** rate at `theta = 90` degrees: directions perpendicular to the gradient produce no first-order change. So two facts fall out of one line: the gradient's *direction* is the steepest-ascent bearing, and the gradient's *magnitude* is the actual rate you achieve by walking that way. ## The hiker picture Imagine a hiker on a hillside in thick fog, able to feel the ground underfoot but seeing nothing beyond a metre. Probing the slope in each compass bearing and keeping the steepest one is exactly evaluating `D_u f` over all `u` and taking the maximum. The gradient is the compass needle that answers this without probing: it already encodes the answer for every bearing at once. The negative gradient is the same needle reversed — the fastest way down. The fog is not decoration. It is the honest part of the picture: the hiker learns only about the ground within arm's reach. Following the steepest bearing does not mean walking toward the summit. On a long curving ridge the steepest local bearing may point across the ridge rather than along it, and the summit may lie in a direction that is barely uphill at all right now. ## Conditions and caveats **Differentiability.** The identity `D_u f = grad f . u` requires `f` to be differentiable at the point. Where it fails, individual partial derivatives may still exist while no single vector reproduces the rate in every direction, and the steepest-ascent story collapses. **A nonzero gradient.** If `grad f(a) = 0`, then every directional derivative is zero and no bearing is distinguished. There is no direction of steepest ascent at such a point — first-order information has run out. **Local and first-order.** The claim is about the instantaneous rate at one point. It says nothing about what happens a finite distance away, and following the steepest bearing repeatedly is not guaranteed to be an efficient route anywhere. **Not scale-invariant.** 'Steepest' silently depends on how you measured the inputs. Rescale one variable — hours to minutes, dollars to thousands of dollars — and the gradient's components change by different factors, so a different bearing becomes the steepest one. Steepest ascent is defined relative to the notion of distance you imposed when you chose units, which is why sensible input scaling is a real modelling decision rather than cosmetics. ## What interviewers listen for The weak answer asserts 'the gradient points uphill' as a memorised fact. The strong answer derives it: rate in a direction equals `|grad f| * cos(theta)`, maximised at angle zero, with the maximum rate being the gradient's own length. Adding the caveats — differentiable, nonzero, local, scale-dependent — is what separates a recital from understanding.
- What does the magnitude of the gradient tell you, as opposed to its direction?It is the actual rate of increase achieved along that best bearing: f rises by about |grad f| units per unit of distance moved. A long gradient means a steep spot, a short one a nearly flat spot. Direction answers 'which way', magnitude answers 'how fast'.
- Which direction gives the fastest decrease, and why is the answer so simple?The negative gradient. The rate in a unit direction is |grad f| * cos(theta), and cos(theta) bottoms out at -1 when the bearing is exactly opposite the gradient, giving -|grad f|. Steepest descent is the same computation with the inequality flipped, which is why it needs no separate derivation.
- Does the direction of steepest ascent depend on the units of the input variables?Yes. Rescaling an input divides or multiplies its partial derivative, so the gradient tilts and a different bearing becomes steepest. 'Steepest' is defined relative to the distance notion your units impose. That is why comparably scaled inputs matter before anyone reads meaning into which component is largest.
- What happens at a point where the gradient is the zero vector?Every directional derivative there is zero, so the surface is flat to first order and no bearing is uphill or downhill. There is simply no direction of steepest ascent. First-order information has been exhausted, and distinguishing a peak, a basin or a flat saddle requires looking beyond the gradient.
A hiker in thick fog can only feel the ground underfoot. The gradient is the compass needle that already knows which bearing rises fastest, without probing every direction by hand.
saying these in an interview costs you the question
- Says the gradient points toward the maximum of the function
- Treats the claim as global rather than local and first-order
- Describes the gradient as a scalar slope
- Forgets that the direction must have unit length for the comparison
- Claims steepest ascent is unaffected by rescaling the inputs
- Asserts a steepest direction exists even where the gradient vanishes