How does the constraint ||w|| <= t relate to minimizing loss plus lambda times ||w||?
answer
- the penalty weight is a multiplier
- one form's budget is the other's price
- the constant term drops from the argmin
- inverse, monotone, data-dependent mapping
- slack budget means zero penalty weight
basics
~20 sThe penalty weight lambda is the KKT multiplier of the constraint ||w|| <= t. For a convex loss the two forms trace the same family of solutions, with a larger budget t matching a smaller lambda.
solid answer
~50 sStart from the constrained form, `minimize loss(w) subject to ||w|| <= t`, and write its Lagrangian: `loss(w) + lambda*(||w|| - t)` with `lambda >= 0`. Since `-lambda*t` does not depend on `w`, minimizing that over `w` is the same as minimizing `loss(w) + lambda*||w||` - the penalized form, with the multiplier playing the role of the penalty weight. Complementary slackness supplies the boundary case: if the unconstrained minimizer already satisfies `||w|| <= t`, the constraint is slack, `lambda = 0`, and the penalty does nothing. Otherwise the constraint is active, the solution sits on the sphere `||w|| = t`, and `lambda > 0`. The correspondence between `t` and `lambda` is monotone in the sense that a larger budget `t` matches a smaller `lambda`, but it is data-dependent with no closed form, which is why you tune one by validation rather than converting from the other.
go deeper
Recall the headline: a hard size budget and a size penalty describe the same family of solutions, and turning one knob up corresponds to turning the other down.
Derive the link by writing the Lagrangian of the budgeted problem and showing the constant term drops out of the minimization over the parameters.
Explain the slack case where the penalty weight is zero, why the mapping between budget and weight has no closed form, and how you would sweep and select in practice.
Decide which parametrization the team exposes - an interpretable budget stakeholders can reason about, or a weight a solver handles directly - and defend the choice on tooling and communication grounds.
## Two formulations of the same idea There are two ways to say do not let the solution get too big. The **constrained** form imposes a hard budget: `minimize loss(w) subject to ||w|| <= t` The **penalized** form charges for size instead: `minimize loss(w) + lambda*||w||` Here `||w||` is any norm of the parameter vector, `t > 0` is a budget on its size, and `lambda >= 0` is a weight. The two are connected by exactly the machinery of constrained optimization: `lambda` is the multiplier on the budget constraint. ## Deriving the link Attach the constraint with a nonnegative multiplier: `L(w, lambda) = loss(w) + lambda*(||w|| - t)` For fixed `lambda`, the term `-lambda*t` is a constant with respect to `w`, so it cannot influence which `w` minimizes `L`. Dropping it leaves precisely `loss(w) + lambda*||w||`. So the penalized objective is the Lagrangian of the constrained problem up to an additive constant, and its weight is the constraint's multiplier - the shadow price of the size budget, in loss units per unit of norm. ## Which case you are in Complementary slackness, `lambda*(||w|| - t) = 0`, splits it cleanly. **Constraint slack.** If the unconstrained minimizer of the loss already has norm below `t`, it is feasible and therefore optimal for the constrained problem. The constraint does nothing, so `lambda = 0`, and the matching penalized problem has no penalty at all. This is the case people forget: a budget generous enough to contain the free solution has no effect whatsoever. **Constraint active.** If the unconstrained minimizer is too large, the constrained solution is pushed to the boundary `||w|| = t` and `lambda > 0`. There the two formulations return the same point, provided the pairing of `t` and `lambda` is the matching one. ## How tight the equivalence is For a convex loss and a convex norm, with a strictly feasible point available, the equivalence runs both ways: for each `t` where the constraint binds there is a `lambda >= 0` whose penalized solution coincides, and conversely, for each `lambda > 0` the penalized solution solves the constrained problem with `t` set to the norm of that solution. Tracing `lambda` from large to small traces `t` from small to large; the two parametrize the same one-dimensional path of solutions in opposite directions. The correspondence is **monotone but data-dependent**. There is no formula converting a chosen `t` into the equivalent `lambda` without solving the problem, because the mapping depends on the loss surface, the sample size and the scaling of the inputs. That is why in practice you pick one parametrization and sweep it against held-out performance rather than converting from a budget someone stated in the other units. Outside convexity the guarantee weakens. With a nonconvex loss, the penalized and constrained problems can have different solution sets - a duality gap - and a penalized solution need not solve any constrained version. The clean two-way correspondence is a convexity result, not a universal one. ## Practical consequences **Solvability.** The penalized form is unconstrained in `w`, so any general-purpose minimizer handles it, which is one reason it dominates in practice. The constrained form needs a solver that respects the feasible region, or a projection step back onto the ball. **Interpretability.** The constrained form has the more communicable parameter: `t` is a budget stated in the same units as the parameters, and it is easy to reason about. `lambda` has units of loss per unit of norm and no natural scale, so its value is meaningless without the accompanying problem. **Non-differentiability.** Norms are not differentiable at the origin, so stationarity is stated with subgradients rather than gradients when the solution can land exactly there. That is a technical refinement, not a change in the story: the multiplier reading survives. **Sanity checks.** If you sweep `lambda` and nothing changes over a whole range at the low end, you are in the slack regime - the budget is not binding and the penalty is inert. If the solution's norm sits exactly at your budget for every `t` you try, the constraint is always active and the loss alone would prefer something much larger. ## The takeaway A hard size budget and a size penalty are not competing ideas but the primal and Lagrangian views of one problem, joined by a single multiplier. Knowing that lets you move between whichever formulation is easier to solve and whichever is easier to explain, and it explains why the two knobs move in opposite directions.
- Does a larger budget t correspond to a larger or a smaller lambda?Smaller. A generous budget barely restricts the solution, so the constraint's shadow price is low; a tight budget forces the solution far from where the loss alone would put it, so the price of one more unit of norm is high. Push `t` beyond the norm of the unconstrained minimizer and `lambda` hits zero exactly.
- Why does the term -lambda*t disappear when you move to the penalized form?Because it contains no `w`. Minimizing over `w` is unaffected by adding or removing any constant, so `loss(w) + lambda*||w|| - lambda*t` and `loss(w) + lambda*||w||` have the same minimizer. The constant does still matter for the objective *value*, which is why it cannot be dropped when computing the dual function.
- Can you convert a stated budget t into the equivalent lambda in closed form?No. The mapping depends on the loss surface, the data and the feature scaling, so recovering the matching `lambda` generally requires solving the constrained problem and reading off the multiplier. In practice you sweep whichever parameter your solver takes and select by held-out performance instead.
saying these in an interview costs you the question
- Claims a fixed formula converts the budget t into lambda
- Says a larger budget corresponds to a larger penalty weight
- Assumes the penalty always moves the solution, even when the budget is slack
- Asserts the equivalence holds for arbitrary nonconvex objectives
- Treats the penalty weight as a unitless, transferable number