Under AdamW with a cosine-annealed learning rate, should the decay term shrink along with the rate?
answer
- Both terms carry the same step size
- Annealing the rate anneals the pull
- Late epochs lose their regularization
- Constant decay dominates a tiny gradient step
- Total shrinkage integrates the schedule
basics
~20 sIn the common formulation the decay step is multiplied by the current learning rate, so annealing the rate quietly anneals regularization too. Decide that deliberately: either hold the decay-to-step ratio fixed on purpose, or apply a rate-independent decay and accept the different endgame.
solid answer
~50 sThe usual written form is `w <- w - lr_t * (adaptive step) - lr_t * wd * w`, so both terms carry the current rate. Under a cosine schedule that means regularization pressure fades toward zero in the final epochs — the phase where the model is fitting the training set hardest. The alternative is a rate-independent decay, `w <- w - lr_t * (adaptive step) - wd * w`, which keeps shrinkage constant but then dominates late in training, when the gradient step has shrunk to almost nothing, pulling weights toward zero regardless of the loss. Neither is free. The judgment call is to pick one deliberately, document it, and keep the *ratio* of decay to gradient step in mind when transferring a recipe: two runs with the same coefficient and different schedules or epoch budgets do not apply the same total shrinkage.
go deeper
Know that the decay applied on a step is usually multiplied by the current learning rate, so a schedule that lowers the rate also lowers the shrinkage that step applies.
Be able to write both variants and say what each does at the end of an annealed schedule: one lets regularization fade with the rate, the other keeps pulling while the gradient step vanishes.
Show that you would check which variant is in play before comparing runs, and that you know shortening a run or reshaping a schedule changes the total shrinkage applied even at an unchanged coefficient.
Own the convention: pick one formulation, make it explicit across recipes, require decay to be revalidated whenever the schedule or epoch budget moves, and define the paired comparison that would settle the choice for a model family.
### The question behind the question Decoupling weight decay from the gradient removes one entanglement — the adaptive denominator — but it does not automatically remove the other one, the step size. In the formulation most people write down, the decay step is scaled by the same `lr_t` that scales the gradient step: `w <- w - lr_t * m/(sqrt(v) + eps) - lr_t * wd * w` So when a cosine schedule anneals `lr_t` from its peak toward zero, the per-step shrinkage anneals with it. The regularization strength is not a constant across the run; it is a curve shaped exactly like the learning-rate curve. ### Why that is arguably right Holding the two terms in fixed ratio means the balance between "follow the loss" and "pull toward zero" is the same at every point in training. The optimizer's direction of travel is scale-invariant in that sense: shrinking `lr_t` slows everything down uniformly rather than changing what the update rule is trying to do. It also means the final epochs, where the schedule is trying to let the model settle into a minimum, are not fighting a decay term that keeps yanking weights off it. It is also the property that makes the coefficient composable. If decay is tied to the rate, then halving the peak rate halves both terms and the character of the run is unchanged; if it is not, halving the rate changes the balance and the decay coefficient must be revisited. The stated motivation for decoupling in the first place is to make the two hyperparameters separable, and scaling both terms by the same schedule multiplier is what preserves that separability across schedules. ### Why it is arguably wrong The counter-argument is that regularization is supposed to be strongest when overfitting risk is highest, and under a long cosine tail the last stretch of training is exactly when the model has the most capacity left to memorize and the least shrinkage pushing back. Under rate-scaled decay, the final epochs are effectively unregularized. The rate-independent alternative, `w <- w - lr_t * (adaptive step) - wd * w`, keeps the pull constant. But it creates its own endgame problem: as `lr_t` approaches zero the gradient step shrinks toward nothing while the decay term does not, so the update becomes dominated by pure shrinkage. Weights drift toward zero for reasons that have nothing to do with the loss, and a long enough tail can undo the fit. It also re-couples the hyperparameters in the opposite direction: with a fixed decay, changing the peak rate changes the decay-to-step ratio at every point. ### The part people miss: total shrinkage is an integral Whichever form you pick, the total amount of shrinkage applied over a run is not the coefficient — it is the coefficient accumulated over every step, which under rate-scaled decay means it depends on the *integral of the schedule*. Two consequences follow, and both routinely surprise teams: - Cutting a run from ninety epochs to thirty at the same coefficient and the same schedule shape reduces total shrinkage roughly threefold. The shorter run is less regularized even though nothing labelled "regularization" was touched. - Changing the schedule shape — a longer warmup, a flatter plateau, a different peak — changes total shrinkage too. A recipe compared across schedule shapes at fixed coefficient is not a controlled comparison. This is why a lead should insist that decay be discussed alongside the schedule and the epoch budget, never as a standalone number in a config diff. ### What to actually own as a decision First, pick a convention and make it explicit across the team's recipes, because a silent difference here makes two runs incomparable in a way no metric names. Second, when moving a recipe to a new schedule or a new epoch budget, treat the decay coefficient as needing revalidation rather than assuming constancy — the accumulated shrinkage has changed even though the number has not. Third, define what evidence would settle it for a given model family: a short paired comparison at matched total shrinkage tells you more than an argument about which form is principled, and the answer can legitimately differ between a small model that needs the late-training pressure and a large one that needs to settle. A strong answer resists picking a universal winner. It names the mechanism — decay scaled by the current rate versus not — states what each buys and costs at the end of a schedule, and moves the discussion to how the choice gets made and recorded.
- If you shorten a run from ninety to thirty epochs at a fixed decay coefficient, what changes?The total shrinkage accumulated over the run falls by roughly the same factor, since the decay applied per step is summed over far fewer steps and, under rate-scaled decay, over a smaller schedule integral. The shorter run is effectively less regularized even though the coefficient is untouched, so it usually needs a larger one.
- What signal tells you the decay term has started dominating late in training?Compare the magnitude of the decay step against the adaptive step for a few representative parameters. If the decay contribution is the larger of the two while the training loss has stopped improving, the model is being pulled toward zero for reasons unrelated to the data, and parameter norms will be falling steadily.
- How would you settle this for a specific model family rather than in the abstract?Run a paired comparison at matched total shrinkage rather than matched coefficient, so the only difference is where in the schedule the pull is applied. Judge on the validation trajectory over the last stretch of training, which is precisely where the two formulations diverge.
saying these in an interview costs you the question
- Assumes decoupled means independent of the learning rate
- Compares two schedules at a fixed coefficient and calls it controlled
- Treats total shrinkage as the coefficient itself
- Declares one formulation universally correct
- Ignores that a constant decay can dominate a vanishing gradient step