skip to content

Why does the L1 norm push coordinates to exactly zero while the L2 norm only shrinks them?

level: middleimportance: must knowfreq 68%

answer

  1. draw both unit balls in two dimensions
  2. one shape has corners, one does not
  3. where do the L1 corners sit?
  4. derivative of |w| versus derivative of w squared
  5. a constant pull versus a fading one

basics

~20 s

The L1 unit ball is a diamond with corners on the coordinate axes, so a constrained solution tends to land on a corner where some coordinates are exactly zero. The round L2 ball has no corners, so it only shrinks.

solid answer

~50 s

There are two ways to see it, and they agree. Geometrically, holding `||w||_1` below a budget confines `w` to a diamond (in higher dimensions a cross-polytope) whose vertices lie on the axes, and a smooth objective contour expanding outward is very likely to touch that shape first at a sharp corner, where every coordinate but one is exactly zero. Holding `||w||_2` below a budget confines `w` to a round ball with no corners at all; the contour touches it at a generic boundary point where all coordinates are small but nonzero. Analytically, the derivative of `|w|` is `+1` or `-1` right up to zero, so the pull toward zero never weakens and a coordinate can be driven all the way in and held there. The derivative of `w^2` is `2w`, which fades to nothing as `w` approaches zero, so shrinkage stalls before it arrives.

go deeper

for a junior

Recall the two shapes: L1 is a diamond with corners on the axes, L2 is a circle. Being able to sketch them and say which one has corners already answers most of what is asked at this level.

for a middle

Explain the tangency argument out loud and back it with the derivative story: constant pull from the absolute value, fading pull from the square. This is the tier where the mechanism, not the picture, is being tested.

for a senior

Show you know when each behaviour is desirable and when the corner geometry misleads, for example the instability of which corner wins when two coordinates carry nearly the same direction.

for a principal

Own the framing that exact zeros are a communication and maintenance property as much as a statistical one, and be ready to argue when a readable sparse result is worth more than a marginally better dense one.

## The two norms, side by side For a coefficient vector `w`, the L1 norm is `|w_1| + |w_2| + ... + |w_n|` and the L2 norm is `sqrt(w_1^2 + ... + w_n^2)`. Both are legitimate measures of how large `w` is, and both are minimised at `w = 0`. Yet if you hold an objective fixed and constrain the size of `w`, L1 produces solutions with exact zeros in them and L2 essentially never does. The reason is geometry, and it is worth being able to draw it. ## The unit balls The set of vectors with `||w||_1 <= 1` in two dimensions is the square rotated 45 degrees, with vertices at (1, 0), (0, 1), (-1, 0) and (0, -1). It is usually called the diamond. Every one of its four vertices lies **on a coordinate axis**, which means the other coordinate is exactly zero there. The set with `||w||_2 <= 1` is the ordinary circle: perfectly smooth, no vertices, no distinguished directions. The set with `||w||_inf <= 1` is the axis-aligned square, whose corners sit at (1, 1) and its reflections, so its corners are the points where *no* coordinate is zero. ## The tangency argument Picture the level sets of a smooth objective as nested contours in the plane, centred on the unconstrained optimum. Shrink the size budget until the constraint bites, and the solution is the first point where an expanding contour touches the ball. On a smooth round ball, tangency happens at a generic point on the boundary, essentially never exactly on an axis. Both coordinates come back toward zero, but neither arrives. That is what shrinkage means. On the diamond, the corners are the points of highest curvature, in fact points where the boundary is not differentiable at all. A whole range of contour orientations touches at the same corner, because a corner presents a fan of supporting directions rather than a single tangent line. That range is what makes exact zeros not merely possible but common: a corner is a target with positive width, whereas a specific point on a circle is a target of measure zero. In `n` dimensions the picture generalises. The L1 ball is a cross-polytope with `2n` vertices on the axes, plus lower-dimensional faces on which some subsets of coordinates are pinned at zero while others vary. So the solution does not have to collapse to a single nonzero coordinate; it can land on a face where, say, three of twenty coordinates are free and seventeen are exactly zero. ## The gradient argument The same conclusion falls out of calculus. Consider the pull that the size term exerts on one coordinate `w_j` as it approaches zero. - For the absolute value `|w_j|`, the derivative is `+1` for any positive value and `-1` for any negative value, however small. The push toward zero has constant strength all the way in. At zero itself the function has a kink and the derivative is undefined; the set of subgradients there is the whole interval from -1 to +1. If the objective's own pull on that coordinate is weaker than the size term's constant pull, zero is a genuine stationary point and the coordinate stays pinned there. - For the square `w_j^2`, the derivative is `2w_j`, which shrinks in proportion to the coordinate itself. As `w_j` gets small the restoring force gets small with it, so the two pulls balance at some small nonzero value. Zero is reached only in the limit of an infinite size penalty. The non-differentiable kink is the whole mechanism. Any smooth approximation to `|w|` restores a vanishing derivative near zero and destroys the sparsity, which is why smoothed absolute values are a poor substitute when exact zeros are the goal. ## Practical consequences and caveats Exact zeros mean a solution you can read: coordinates that are gone are gone, not merely small, so the result doubles as a selection. L2 shrinkage keeps every coordinate present, which is often the more faithful description when the coordinates genuinely share the signal. Two caveats belong with the geometry. First, when two coordinates point in nearly the same direction, the diamond offers several nearly equally good corners, so which one is selected can flip with a small change in the data. The L2 ball, having no corners, splits the weight between them stably instead. Second, exact zeros are a property of the constrained or penalised solution, not of the norm on its own: computing `||w||_1` for an existing vector creates no zeros at all. ## Checking your understanding If someone claims sparsity comes from L1 being smaller than L2, correct it. For the vector (3, 4) the L1 norm is 7 and the L2 norm is 5, so L1 is the *larger* of the two here. Magnitude is not the point. The corners are.

  • What does the L-infinity unit ball look like, and would constraining it produce zeros?
    It is the axis-aligned square, and its corners sit at points like (1, 1) where **no** coordinate is zero. Its corners instead encourage coordinates of equal magnitude, so an L-infinity budget tends to level coordinates toward a common absolute value rather than zeroing any of them.
  • Why does replacing the absolute value with a smooth approximation destroy the sparsity?
    Because the sparsity comes from the kink at zero. A smooth surrogate has a derivative that fades toward zero as the coordinate does, exactly like the square, so the restoring pull weakens and the coordinate settles at a small nonzero value. You get shrinkage back and lose the exact zeros.
  • Does computing the L1 norm of a vector make it sparse?
    No. A norm is only a measurement; it reports a number and changes nothing. Sparsity appears when the L1 value is *constrained* or added to an objective being minimised, so that the geometry of the L1 ball shapes where the solution can sit. The norm by itself never sets a coordinate to zero.

A marble rolling into a round bowl settles somewhere near the middle but off-centre; the same marble rolling into a diamond-shaped tray with sharp corners tends to wedge itself into a corner and stop dead.

saying these in an interview costs you the question

  • Says L1 gives zeros because it is smaller than L2
  • Cannot draw or describe the two unit balls
  • Claims L2 also produces exact zeros, just fewer
  • Thinks the L1 corners sit away from the axes
  • Believes computing a norm changes the vector

context