skip to content

Learning-Rate Range Test

Sweeping the rate upward over a few hundred steps and plotting loss against it exposes the band where loss falls fastest and the point where it blows up. It answers 'what rate would you start at?'.

on this pageshow

questions

3

How do you run a learning-rate range test to pick a first rate for a new model?

level: middleimportance: must knowfreq 52%

answer

  1. one short throwaway run
  2. multiply the rate, never add
  3. log x-axis, smoothed loss
  4. the knee, not the bottom
  5. about ten times below the minimum

basics

~20 s

Run one short pass, multiplying the learning rate each step from tiny to divergent, and plot smoothed loss against rate on a log axis. Start the real run about a factor of ten below where the loss bottoms out.

solid answer

~50 s

Run one short throwaway pass from the model's initial weights, multiplying the learning rate by a constant factor after every step so it sweeps geometrically from far too small to certainly divergent — on a keyword-spotting model over 16 kHz voice-command clips, `1e-7` to `10` across about 200 steps. Record each step's mini-batch loss, smooth it with an exponential moving average, and plot it against the rate on a log x-axis. You get a flat region, a descent, a minimum, then a sharp rise. Do not take the minimum: it sits at the edge of instability, and the curve is a trailing measurement that credits the smaller rates that came before it. Take the point of steepest descent, or roughly an order of magnitude below the minimum. Then discard the swept weights, reload the initial state, and start the real run at that rate.

go deeper

for a junior

Be ready to say what the plot's axes are: loss against learning rate on a log scale, produced by one short run. Knowing that the rate is swept upward step by step, rather than tuned by many separate runs, is the recall you need.

for a middle

Explain the mechanics: a constant multiplier per step, an exponential moving average over per-step losses, a stop rule once the loss blows past its best, and why the rate you choose sits below the curve's minimum rather than at it.

for a senior

Show that you use it as a cheap guard on real work: enough steps per decade for the curve to be readable, weights restored from the initial checkpoint afterwards, and the number treated as an upper bound you confirm on the first hundred steps of the actual run.

for a principal

Own the cost argument. A two-hundred-step sweep prevents a class of silent multi-day failures, so make it the convention for any new configuration, and make the curve and the chosen rate part of what every run records.

## Why the test exists The learning rate is the single hyperparameter that decides whether a run learns at all. Too small and the loss barely moves for hours; too large and it climbs, oscillates or blows up. The usable band spans several orders of magnitude, and it moves with almost every change to the recipe, so guessing costs whole runs. A range test buys a measurement of that band for the price of a few hundred steps. ## The procedure Start from the initial weights of the model you actually intend to train, and save them first. Set the rate to something certainly too small — `1e-7` is a common floor — and pick a ceiling certainly too large, say `10`. Choose a step count N, then after every step multiply the rate by the constant factor `(max/min)^(1/N)`, so the rate traverses the range geometrically rather than in equal increments. Take one mini-batch per step, perform one update, record that batch's loss, and continue until you reach the ceiling or the smoothed loss exceeds its best value by a large factor — a stop rule of roughly four times the best loss is common, because a diverged model produces nothing readable afterwards. On the keyword-spotting model over 16 kHz voice-command clips, a sweep from `1e-7` to `10` across 200 steps is a typical shape: 200 steps across eight decades of rate is about 25 measurements per decade, enough to see structure. When the sweep ends, throw the weights away and reload the saved initial state. The sweep has taken a couple of hundred updates at rates ranging from useless to explosive, several of them past the point of blow-up; the model it leaves behind is damaged, not warmed up. The only artifact you keep is the number. ## Reading the plot Plot loss on the y-axis against rate on a logarithmic x-axis, and smooth the loss first — an exponential moving average over the per-step values — because a single mini-batch's loss is noisy enough to invent dips that are not there. Four regions appear: 1. **A flat left region.** The rate is so small that nothing measurable changes. 2. **A descent.** The loss falls, and somewhere in here it falls fastest. 3. **A minimum.** The smoothed loss stops improving. 4. **A sharp rise.** Updates overshoot and the run diverges. The trap is the minimum. It is the most visible feature, so it is the one people point at, and it is already too hot. Two reasons. First, the curve is a *trailing* measurement: the loss recorded at rate r reflects every earlier, smaller update too, so part of the improvement visible there was earned at lower rates. It is not the loss you would obtain by training at r. Second, by the time improvement stops, the rate is essentially at the edge of the region where updates still help; a full run held there has thousands more steps in which to meet a batch that overshoots. The two standard readings are the point of steepest descent — where the smoothed curve falls fastest — and roughly one order of magnitude below the minimum. On a curve that bottoms near `3e-2`, both rules land somewhere around `1e-3` to `3e-3`. ## How long the sweep has to be Count steps per decade, not steps. A 40-step sweep across eight decades is five points per decade, each point a single noisy mini-batch drawn from a million-example corpus: the result is jagged noise in which any dip can be misread as a knee. A few hundred steps over the full range is the usual working point. It is also normal to run a coarse sweep, see roughly where the interesting region is, then re-run over three decades with the same number of steps for a much cleaner curve. ## What the number is and is not Treat the output as an upper bound plus a sensible starting point, not an oracle. It is a short-horizon statement about training loss on one exact configuration. Confirm it on the first hundred steps of the real run: the loss should fall steadily rather than flatline or spike. And remember that the curve belongs to the whole recipe — the same architecture with a different batch size or a pretrained encoder produces a different curve and a different answer.

  • Why sweep the rate multiplicatively rather than in equal increments?
    Learning-rate effects are multiplicative: `1e-4` to `2e-4` is a real change, `0.5` to `0.5001` is not. The usable band can sit anywhere across eight orders of magnitude, so equal increments spend nearly every sample inside the diverged region and none in the decades where the answer lives. A constant multiplier gives equal resolution per decade, which is exactly how the log-axis plot is read.
  • A 40-step sweep over a million-example corpus produced an unreadable, jagged curve. What went wrong?
    Forty steps across eight decades is five samples per decade, and each sample is one noisy mini-batch, so batch-to-batch variation swamps any trend. Fix it by counting steps per decade instead of total steps: a few hundred steps over the same range, or the same 40 steps over a narrowed three-decade range. Smooth the loss before plotting rather than reading raw per-step values.
  • Do you keep training from the weights the sweep produced?
    No. The sweep took a couple of hundred updates at rates spanning useless to divergent, including several past the blow-up point, so the resulting weights are damaged rather than warmed up. Save the initial state before the sweep and reload it afterwards. The number is the only thing the test produces that you keep.

Like turning a volume knob up steadily until the speaker distorts. You learn where the limit is, then set the dial comfortably below it — not at the loudest point that still sounded acceptable.

saying these in an interview costs you the question

  • Picks the rate at the very bottom of the loss curve
  • Sweeps learning rates in equal linear increments
  • Reads the raw unsmoothed loss and picks a noisy dip
  • Runs the sweep for a full epoch before plotting anything
  • Continues training from the weights the sweep left behind
  • Treats the curve as the loss training at that rate would give

context

open as a page

Why must you re-run a learning-rate range test after the batch size or initialization changes?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The curve describes a configuration, not an architecture. A larger batch averages away gradient noise, so the divergence knee moves up; a pretrained encoder sits near a good solution, so rates that were fine from scratch wreck its features.

open as a page

When is a learning-rate range test the wrong tool for choosing the rate you ship?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

A range test measures one thing: the largest rate that keeps training loss falling over a couple of hundred steps. It is silent on generalization, on stability thousands of steps later, and on recipes whose good rate is already known.

open as a page