skip to content

How do you set warmup length - a fixed step count or a fraction of the run - when porting a recipe to a far larger dataset?

level: principalimportance: should knowfreq 38%

answer

  1. what is the ramp buying, in what currency
  2. steps, not epochs, not percent
  3. epoch-based ramps move with batch size too
  4. err long, the cost is bounded
  5. decay stretches with budget, the ramp does not

basics

~20 s

Warmup's job is finished after a certain number of optimizer steps, not after a share of the data, so a fixed step count is the defensible unit. A percentage silently stretches tenfold on a tenfold dataset.

solid answer

~50 s

Ask what the ramp is buying. It buys time for the weights to leave the sharp initial region and for the optimizer's running statistics to accumulate - both counted in optimizer steps, neither caring how much data remains. So a fixed step count, sized to the optimizer's averaging window and verified on the opening steps, ports correctly. "Two percent of the run" or "one epoch" does not: on a dataset ten times larger an epoch-based ramp runs ten times longer for no gain, and on a short fine-tune it shrinks until it protects nothing. The unit also breaks quietly when batch size changes, since steps per epoch move with it. For a shared recipe: express the ramp in steps, err long because the cost is a fraction of a percent, and let the decay - not the ramp - stretch with the budget.

go deeper

for a junior

Know that warmup length is configured in some unit and that the units are not interchangeable - steps, epochs, and a percentage of the run mean different things when the dataset changes.

for a middle

Explain why the ramp's job is counted in optimizer steps: leaving the sharp initial region and accumulating the optimizer's running statistics both depend on steps taken, not on data remaining.

for a senior

Show the verification loop - log unsmoothed loss and gradient norm over the opening steps, and read a break that lands exactly where the ramp ends as evidence to lengthen the ramp rather than lower the target rate.

for a principal

Own the policy: one unit across every recipe, sized with margin because the downside is asymmetric, the derived percentage logged for comparability, and a clear separation between the ramp and the decay that stretches with the budget.

## The question behind the question Every training recipe expresses warmup in some unit: a step count, a percentage of total steps, or a number of epochs. Nobody thinks about the unit until the recipe moves - to a bigger corpus, a shorter fine-tune, a different batch size - and then the choice of unit silently changes the actual ramp. The interviewer wants to see whether you reason from what warmup is for, or from what the config file happened to say. ## What the ramp is buying, in what currency Warmup exists to carry the run through a transient regime with two components. First, the iterate has to leave the neighbourhood of a random initialization, where the curvature can be sharper than anywhere the target rate was tuned for. Second, the optimizer's running averages - momentum, and for adaptive methods a per-coordinate second moment - have to accumulate enough samples to be stable. Both are measured in **optimizer steps**. Neither has any dependence on how many more batches exist afterward. That single observation settles the unit argument. Steps are the currency in which the ramp's job is denominated, so steps are the unit in which it should be written down. ## What goes wrong with the alternatives **A fraction of total steps.** Convenient, and roughly harmless within one order of magnitude of budget. Take the recipe to a dataset ten times larger, keep "two percent of the run", and the ramp goes from 500 steps to 5,000. That is not a disaster - the extra steps are still training steps, just at a lower rate - but it is 4,500 steps of deliberately reduced progress bought for nothing. Take the same recipe the other way, to a 2,000-step fine-tune, and two percent is 40 steps, which is inside the window where the second-moment estimate is still nonsense. The failure at the short end is real, and it is the one that bites, because short runs are exactly where people reuse a recipe unthinkingly. **A number of epochs.** Worse, because it couples the ramp to two independent quantities. Steps per epoch equal dataset size divided by batch size, so an epoch-based ramp changes when the corpus grows *and* when someone changes the batch size for throughput reasons. A team member doubling the batch to fill a machine has just halved the warmup without touching the warmup config, and the review diff shows nothing suspicious. **A fixed step count.** Ports on the axis that matters. It has one honest weakness: on a very short run, a fixed ramp can become a large share of the budget, so it needs a cap - something like "500 steps, or ten percent of the run, whichever is smaller" is a reasonable compromise that keeps the step-count intent while preventing an absurd result on a tiny fine-tune. ## Sizing the number, not just the unit A defensible starting point is the optimizer's averaging window. With a second-moment decay of 0.999 the effective window is on the order of a thousand steps, and ramps in the hundreds to low thousands are what practitioners converge on for that setting. Then verify on the run itself: log unsmoothed per-step loss and gradient norm over the first couple of thousand steps. If the trace breaks - spikes, plateaus, or diverges - at or just after the step where the ramp ends, the ramp was too short; that coincidence in timing is the signature, and the fix is a longer ramp rather than a lower target rate. If nothing interesting happens anywhere in the opening steps, you may be able to shorten it, though the saving is rarely worth the risk. The asymmetry drives the policy. An over-long ramp costs a small, bounded amount of progress. An under-long ramp costs a dead run, sometimes only on some seeds, discovered hours in. When the downside is that lopsided, err long. ## The organizational call This is where the principal-level answer lives. Pick one unit and make every recipe in the organization use it, because the cost of mixed units is not a bad ramp on one run - it is that nobody can tell by reading two configs whether they are training the same way. Write the ramp in steps, record the derived percentage in the run's logged metadata so comparisons across budgets are readable, and make the opening-steps trace a standard artifact of every launch so an under-warmed run is caught in the first two minutes rather than at the first evaluation. And keep the roles separate: the ramp is sized by the optimizer and the architecture, while the part of the schedule that stretches to fill a bigger budget is the decay that follows it. ## Also on the checklist The ramp shape is a minor knob compared with its length - a linear ramp and a gradual exponential ramp behave similarly, and a short constant-low-rate phase before the ramp is a variation, not a different technique. Spend the argument on length and unit; nobody has ever lost a run to the shape.

  • Does the ramp shape matter - linear versus a gradual exponential rise?
    Much less than the length. Both spend the opening steps at a reduced rate and reach the target at the same point, and runs are not typically won or lost on the curvature of the ramp. A short constant-low-rate phase before the ramp is a similar variation. Standardize on linear because it is the easiest to read and reason about, and spend your attention on how many steps the ramp runs for.
  • Someone proposes shortening the ramp to save budget. How do you evaluate it?
    Price both sides. The saving is the difference in steps times the fraction of a step's value lost at a reduced rate - a fraction of a percent on any long run. The risk is a dead run, discovered hours in, possibly only on some seeds. Ask for evidence across several seeds, not one green launch, and require the opening-steps trace to be clean at the shorter setting before adopting it.
  • Would you ever set warmup as a percentage anyway?
    Yes, as a cap rather than the primary rule: a fixed step count, bounded above by a share of the run, keeps the step-count intent while preventing a fixed ramp from swallowing a very short fine-tune. State it in that order in the config so the intent is legible, and log the resulting absolute step count with the run.

saying these in an interview costs you the question

  • Keeps a percentage-of-run ramp without asking what it becomes at the new scale
  • Expresses warmup in epochs, unaware batch size changes steps per epoch
  • Treats a shorter ramp as a meaningful budget saving
  • Lowers the target rate instead of lengthening a ramp that ends where the run breaks
  • Assumes the ramp should stretch with the budget the way decay does

context