skip to content

Should a weekly hyperparameter search be warm-started from last week's best configuration?

level: seniorimportance: nice to knowfreq 26%

answer

  1. two different things share this name
  2. prior evidence, not a prior conclusion
  3. a score belongs to the data it was measured on
  4. the search can stop looking anywhere new
  5. cold-start occasionally to check

basics

~20 s

Warm-start by seeding the surrogate with past trials, not by trusting last week's winner. Old scores do not transfer to new data, so re-evaluate the incumbent, keep some random exploration, and cold-start periodically to catch a moved optimum.

solid answer

~50 s

Take the delivery-ETA model retrained every week. Warm-starting the search means seeding the surrogate with the (configuration, score) pairs from earlier weeks so the new search starts informed rather than blind — that is different from warm-starting model weights, which is about initialising training. It is usually worth doing, with three guards. First, always re-evaluate the incumbent configuration on this week's data; an old score was measured on old data and does not carry over. Second, down-weight or expire old observations, because feature changes, new cities or a shifted delivery mix move the objective surface and stale points will pin the surrogate to a basin that no longer wins. Third, keep a slice of the budget on fresh independent draws so the search cannot collapse onto the incumbent's neighbourhood. On a slower cadence — say monthly — run a full cold search to check the optimum has not migrated somewhere your warm searches would never look.

go deeper

for a junior

Know that a hyperparameter search can start from what earlier runs learned instead of from nothing, and that a score measured on old data has to be re-measured before you trust it.

for a middle

Distinguish warm-starting the search history from warm-starting model parameters, and explain how seeded observations change what the acquisition rule proposes first.

for a senior

Give a concrete recurring-retrain policy: age-decay or expiry of old trials, a re-scored incumbent, a reserved share of fresh draws, and history invalidation when the feature set changes.

for a principal

Own the cadence tradeoff for the organisation: how often a cold search is worth its compute, what margin justifies changing a deployed configuration, and who notices when tuning has quietly frozen.

## Two things called warm starting Be explicit about which one is being discussed, because interviewers use the phrase loosely. - **Warm-starting training**: initialising a new model's parameters from a previously fitted model so training converges faster. That is about the model. - **Warm-starting the search**: initialising the hyperparameter search with knowledge from earlier searches — a history of (configuration, validation score) pairs, or at minimum the previous winner evaluated as an early trial. That is about the search, and it is what this question is about. ## Why it helps A surrogate-based search spends its first several trials just building enough of a picture to be useful. On a recurring retrain, that picture already exists. Seeding the surrogate with prior observations skips the uninformed phase, so a 30-trial weekly budget behaves more like the tail of a 200-trial search. On a bandit-style budget schedule the equivalent move is to inject prior good configurations into the first rung alongside fresh draws, so at least some strong candidates are in the race from the start. There is a second, cheaper benefit: continuity. If the tuner reproposes roughly the same configuration each week, the model's behaviour is stable and any change in metrics is more attributable to the data than to a tuning coin-flip. ## Why it is dangerous **Old scores are measurements, not properties.** Last week's configuration scored what it scored *on last week's data*. New rows, a re-derived feature, a fixed label bug, a seasonal shift in delivery patterns — any of these change the objective. Seeding the surrogate with stale `y` values makes it confident about a surface that has moved. The minimum discipline is: re-score the incumbent this week, and treat only that fresh score as authoritative. **Exploration collapses.** An acquisition rule fed a history that says one region is excellent will keep proposing near that region. Each week the history grows more lopsided, uncertainty everywhere else stays untouched by fresh evidence, and the search quietly becomes a local refinement around a configuration nobody has re-justified in months. The optimum can drift out from under you and no trial would ever notice. **Hidden coupling to a stale space.** If the feature set changed, some knobs mean something different than they did — the right regularisation strength for 40 features is not the right one for 120. Prior observations from before that change are not merely old, they describe a different function. ## A workable policy 1. **Seed, do not conclude.** Load prior trials into the surrogate as observations, not as a fixed answer. 2. **Decay by age.** Weight observations down as they get older, or expire them past a horizon, so the surrogate's confidence reflects how much of its evidence is still current. 3. **Always re-run the incumbent.** It costs one trial and it gives this week's baseline a fresh, comparable number. 4. **Reserve exploration.** Force a fixed share of trials — a quarter is a reasonable default — to be drawn independently of the surrogate. 5. **Invalidate on schema change.** When features, target definition or data pipeline change materially, drop the history rather than decaying it. 6. **Cold-start on a cadence.** Periodically run a full search from scratch and compare its winner to the warm-started one. If the cold search keeps finding the same region, you have earned the right to warm-start more aggressively; if it keeps finding somewhere else, your warm searches are stuck. 7. **Adopt only on a real margin.** Switching configuration every week because a new candidate beat the incumbent by less than the fold-to-fold noise is churn, not improvement. Require the gap to exceed the measurement spread before changing what is deployed. ## What interviewers listen for The distinction between the two kinds of warm start, awareness that a stale score is not a valid observation, and a concrete guard against exploration collapse. Candidates who answer only "yes, it saves time" have not thought about drift; candidates who answer only "no, always search fresh" are burning budget rediscovering the same region every week.

  • How would you tell that a warm-started search has stopped exploring?
    Track the spread of proposed configurations over weeks. If every proposal lands in a shrinking neighbourhood of the incumbent while the fresh random draws you reserved keep scoring close behind, the surrogate is refining rather than searching. A periodic cold search that lands somewhere else is the clearest confirmation.
  • When should the search history be discarded rather than down-weighted?
    When the objective is no longer the same function: a changed feature set, a redefined target, a fixed data bug, or a new evaluation split policy. Age-decay assumes the surface is drifting slowly; a schema change means old observations describe a different problem and will actively mislead the surrogate.

It is like reusing last season's scouting notes. Useful for deciding where to look first, worthless as a verdict on a player you have not watched this season.

saying these in an interview costs you the question

  • Reuses last week's validation score as if still valid
  • Confuses warm-starting the search with warm-starting model weights
  • Fixes the incumbent as the answer and only searches nearby
  • Never cold-starts, so a moved optimum goes unnoticed
  • Swaps the deployed configuration on a difference within noise

context