skip to content

Sweeping weight decay across four decades on a wearable-accelerometer activity classifier, how do you pick the value?

level: seniorimportance: should knowfreq 55%

answer

  1. one curve alone tells you nothing
  2. select on the held-out side
  3. search multiplicatively, not additively
  4. U-shape with a broad flat top
  5. both curves falling means underfitting

basics

~20 s

Read the held-out curve, not the training curve. Across a logarithmic sweep the held-out metric is U-shaped: too little decay changes nothing, too much underfits. Pick from the flat top of that curve, then re-sweep when data volume or schedule changes.

solid answer

~50 s

Sweep multiplicatively — powers of ten first, then refine by factors of about three around the best region — because the coefficient acts on a log scale and a linear grid wastes almost every run. Plot both curves. The training metric on the six-channel window data degrades monotonically as decay rises; that curve tells you the knob is connected, nothing more. The held-out metric is the one that decides: at the small end it is flat, because the penalty is too weak to move the equilibrium weight norm at all; through the middle it improves as unsupported weights are suppressed; past the peak both curves fall together, which is the signature of underfitting rather than regularising. Pick from the flat region near the peak rather than the exact argmax, since run-to-run noise on 20,000 windows easily exceeds the difference between neighbouring grid points. Then hold the rest of the recipe fixed while you sweep, and re-tune whenever dataset size, run length or the learning-rate schedule changes.

go deeper

for a junior

Know that the decay coefficient is chosen by trying several values on a logarithmic grid and comparing held-out performance, never by looking at how well the model fits the training set.

for a middle

Describe the three regions of the held-out curve — flat, rising, falling — and be able to say that both curves declining together is the signature of underfitting rather than overfitting.

for a senior

Show operational judgment: pick from a plateau rather than an argmax, check the training curve responds at all as a wiring sanity test, and re-sweep after any change to data volume, run length or schedule.

for a principal

Own the cost argument for joint versus sequential search over learning rate and decay, and set the team convention for when a tuning result is declared stale and must be re-run.

## Why a four-decade logarithmic sweep The decay coefficient enters the update multiplicatively and useful values on real problems span several orders of magnitude, so the only sensible search is a logarithmic one. Start with one run per decade across four decades, look at the shape, then refine with a factor-of-three grid in the one or two decades that matter. A linear grid over the same range spends almost every run in a regime where decay does nothing at all, and will report — correctly and uselessly — that the knob has no effect. ## The two curves and what each one is for Run the sweep on a fixed recipe: same architecture over the six sensor channels, same window length, same learning-rate schedule, same number of epochs, same seed policy. Record both a training metric and a held-out metric. The **training metric** is a sanity instrument, not a selection instrument. It should fall monotonically as decay rises. If it does not move at all across the whole sweep, decay is not reaching the parameters you think it is — a grouping bug, or values so small they are swamped by the gradient. If it collapses immediately at the smallest nonzero value, your grid starts too high. The **held-out metric** is what you select on, and its shape has three regions: - **Flat left region.** The penalty is too weak to move the equilibrium weight norm perceptibly, so the run is indistinguishable from no decay. Every point here is the same model. - **Rising middle region.** Decay is now suppressing weights the 20,000 windows do not support, and held-out performance improves while training performance slips. This is the regime you want. - **Falling right region.** Both curves now decline together. That joint decline is the diagnostic signature of underfitting: the penalty is suppressing structure the data genuinely supports. There is no ambiguity here — regularising too hard and overfitting look nothing alike on a two-curve plot. ## Choosing a point Take the flat top, not the argmax. On a dataset of this size, run-to-run variation from initialisation and data ordering routinely exceeds the gap between adjacent grid points, so the single best number in a sweep is partly luck. A value sitting in a broad plateau will reproduce; a spike that beats its neighbours by a hair usually will not. If you can afford it, run two or three seeds at the two or three candidate values and choose on the mean, preferring the slightly stronger decay when the means tie, since that is the side that degrades gracefully. ## What invalidates the result The decay coefficient is not a property of the architecture. It is calibrated against a particular amount of data, a particular number of update steps and a particular learning-rate schedule, and changing any of these moves the right value. - **Data volume.** More data does more of the regularising by itself, so the best explicit penalty falls as the dataset grows. - **Run length.** Decay acts once per update, so total shrinkage depends on the number of steps. Halving the epochs weakens the effective regularisation at a fixed coefficient. - **Learning-rate schedule.** In the usual update the shrinkage applied per step is scaled by the learning rate, so a change to the schedule changes how much decay is actually applied. On layers followed by normalization the coupling is stronger still, since the two together determine the effective step size. - **Other regularisers.** Adding augmentation or dropout after tuning decay means part of the job is now being done twice. Any of these means the sweep is stale and needs re-running, at least coarsely. ## Sweeping decay together with other knobs A joint search over learning rate and decay is more honest than a sequential one, because the two interact, but it costs a product of runs rather than a sum. A workable compromise is a coarse joint grid — three learning rates by four decay values — to locate the region, then a one-dimensional refinement of decay at the winning learning rate. State the compromise explicitly rather than pretending the sequential search was principled; interviewers are listening for whether you know the knobs interact. ## What a weak answer looks like Selecting on the training metric, sweeping linearly, tuning once and never revisiting, or treating more decay as inherently safer. The last is the most common: candidates internalise that regularisation prevents overfitting and conclude that erring high is conservative. On the right-hand side of the curve it is simply a worse model, and unlike overfitting it will not announce itself as a train-versus-held-out gap.

  • Your training set grows tenfold and the old decay coefficient now underfits — why?
    More data does more of the regularising on its own, so the optimal explicit penalty falls. The coefficient calibrated on the small set now suppresses weights that the larger, richer set genuinely supports, and both training and held-out metrics sit below what a weaker penalty reaches. Re-sweep downward whenever data volume changes materially; the old value is stale, not wrong-in-principle.
  • On an ad click-through ranking model, raising decay lifts held-out ranking quality while training quality falls. When do you stop raising it?
    At the point where held-out quality stops improving. Keep raising while only the training metric declines; stop the moment the held-out curve flattens, and back off as soon as it turns down, since both falling together means you have crossed into underfitting. Choose from the flat region rather than the exact peak so the choice survives another seed.
  • How would you tell a decay sweep apart from a broken parameter grouping?
    Look at whether the training metric responds at all. A correctly wired sweep shows training performance degrading monotonically as the coefficient rises across four decades. A completely flat training curve means the shrinkage is not reaching the weights — a grouping mistake or a grid centred far too low — and no held-out conclusion from that sweep is trustworthy.

saying these in an interview costs you the question

  • Picks the coefficient by best training metric
  • Sweeps on a linear grid instead of a logarithmic one
  • Reads both curves falling as continued overfitting
  • Tunes decay once and never revisits after data or schedule changes
  • Assumes stronger decay is always the safer error

context