Porting a recipe from Adam with L2 in the loss to AdamW, why re-tune the weight-decay value?
answer
- The number means different things
- Divided by the gradient RMS estimate
- Typical gradient scale is below one
- Amplification disappears when decoupled
- Re-tune rather than convert
basics
~20 sThe coefficient does not carry over. With the L2 term inside the gradient, the adaptive denominator rescales the penalty before it lands, so the same number produces very different shrinkage once decay is decoupled — in practice the decoupled value usually has to be much larger.
solid answer
~50 sIn the coupled setup the penalty is divided by the per-parameter adaptive denominator, so the shrinkage actually applied is roughly `wd * w / sqrt(v)` rather than `wd * w`. Typical gradient RMS values in a trained network are well below one, so that division tends to *amplify* the penalty — which is why a small coefficient sufficed. Decouple the decay and the amplification disappears, so the same number now barely regularizes and the tuned value commonly moves by one to two orders of magnitude. The mapping is not a constant either: the divisor differs per parameter and drifts over training, so there is no factor you can multiply the old value by. Re-tune it, and treat the fine-tuning run as under-regularized until you have. The payoff is that afterwards the decay coefficient and the learning rate move independently.
go deeper
Remember that a decay number copied from another training recipe is only meaningful alongside the optimizer that consumed it. Ask which form of decay the original used before reusing the value.
Be able to explain, with the update written out, why dividing the penalty by a per-parameter gradient statistic makes the coefficient non-transferable, and roughly which direction the tuned value moves.
Show the diagnosis, not just the theory: which signals you would watch on the first ported run, how you separate a too-small coefficient from a lost amplification, and how you avoid burning a long run to find out.
Take a position on how inherited recipes are validated at all before they become defaults, since a silently non-transferable hyperparameter costs a team far more than the single run that first exposed it.
### The failure this question is about A team lifts a fine-tuning recipe — architecture, schedule, batch size, learning rate, decay coefficient — from a setup where the L2 term was added to the gradient, and reuses it with decoupled decay. Nothing in the configuration file changed except which optimizer implements the decay, so everyone believes the run is comparable. It is not: the validation gap widens, or the model quietly overfits several epochs earlier, and the postmortem blames data or the schedule. ### Why the number does not transfer In the coupled form, the penalty term `wd * w` is summed into the gradient and then goes through the same per-parameter division as everything else. The shrinkage that actually reaches the weight per step is approximately `lr * wd * w / (sqrt(v) + eps)` where `sqrt(v)` is roughly the recent root-mean-square gradient magnitude for that parameter. In the decoupled form it is simply `lr * wd * w` So the two forms agree only when `sqrt(v)` happens to be near one. In practice it usually is not. Gradient magnitudes in a network past the first few epochs are typically well below one — think 1e-2 to 1e-4 — which means the division *multiplies* the penalty up by a factor of tens to thousands. That is why coupled recipes historically ran with small decay coefficients and decoupled recipes run with visibly larger ones: the same real shrinkage requires a bigger nominal number once the amplifier is removed. ### Why you cannot just apply a correction factor The tempting fix is to compute a single multiplier and scale the old value. It does not work, for two reasons. First, the divisor is per parameter. Two layers in the same network with different gradient scales sit at different points on that amplification curve, so no single scalar reproduces the old behaviour everywhere. The old run was, in effect, applying a different decay coefficient to every parameter — one you never chose and cannot recover with one number. Second, the divisor drifts. Gradient magnitudes shrink as training converges, so the effective coupled decay strengthens over the course of a run. The decoupled version is stationary in that sense. Even a perfectly matched value at epoch one diverges from the old behaviour by epoch thirty. The practical conclusion is that the decay coefficient becomes a hyperparameter to re-establish from scratch, not one to convert. The consolation is that it is now a cleaner one: because the decay no longer travels through the moment estimates, changing it does not perturb the gradient step, so it can be explored on a coarse logarithmic ladder without re-tuning the learning rate every time. ### How to confirm the diagnosis rather than guess The signal that separates "decay coefficient is too small" from "decay was silently amplified in the old setup" is the parameter-norm trajectory, tracked per block rather than for the whole model. In a coupled run the blocks with the smallest gradients shrink hardest and the busiest blocks barely shrink at all — an uneven pattern nobody configured. After decoupling, the same coefficient produces a uniform relative pull, and if the norms across every block now creep upward through training while the training loss keeps falling, that is the amplification you lost, not a data problem. A second check: compare the magnitude of the decay step against the magnitude of the adaptive step for a few representative parameters. If the decay contribution has dropped by orders of magnitude relative to the old run, you have measured the amplification directly rather than inferring it. ### What a strong answer sounds like Say that the coefficient means something different in the two forms because one of them is divided by an adaptive per-parameter quantity; that this normally means the decoupled value must be larger, though the exact factor is neither constant across parameters nor stable over training; that you re-tune rather than convert; and that you verify with per-block parameter norms rather than by staring at validation accuracy alone. Then note the upside: after the change, decay strength and learning rate stop interfering with each other, so the recipe is easier to move to the next model.
- What would you watch during the first ported run to confirm the coefficient is now doing what you think?Per-block parameter norms over training, not just validation accuracy. Under decoupled decay the relative pull should be uniform across blocks; if norms across every block creep upward while training loss keeps falling, the shrinkage is too weak. Comparing the magnitude of the decay step against the adaptive step for a few parameters measures it directly.
- Why can't you just multiply the old coefficient by a correction factor?Because the divisor is per parameter and it moves. Layers with different gradient scales sat at different amplification factors, so no scalar reproduces the old per-parameter pattern, and gradient magnitudes shrink as training converges, so even a matched value at the start drifts apart by the end.
- Is there any upside to the re-tuning cost?Yes. Once the decay bypasses the moment estimates, changing it no longer perturbs the direction or size of the gradient step, so decay strength and learning rate can be explored more independently. That is the practical argument for the decoupled form beyond the uneven-shrinkage fix.
saying these in an interview costs you the question
- Assumes the coefficient transfers unchanged between the two forms
- Proposes one constant factor to convert the old value
- Blames data or schedule before checking the decay path
- Judges regularization strength from validation accuracy alone
- Claims gradient magnitudes are near one so nothing changes