skip to content

With a fixed 60-trial budget and six hyperparameters, how do you decide what to search?

level: principalimportance: should knowfreq 38%

answer

  1. start from wall-clock compute
  2. trials times folds times fit time
  3. fix the inert axes deliberately
  4. coarse round, inspect, then refine
  5. stop when gains sink below fold noise

basics

~20 s

Start from the compute you have, not the size of the space. Fix the hyperparameters that rarely matter at defaults and give the whole budget to the two or three that plausibly move the score.

solid answer

~50 s

The budget is a compute decision: 60 trials times the fold count times the fit time is a wall-clock number I want stated before starting. Within it I fix the axes that are conventionally inert for the model family and give the whole 60 trials to the two or three that matter — six axes at 60 trials spreads coverage thinly, and four of them will turn out flat. Random draws suit a fixed budget precisely because the trial count is an input rather than a consequence of the space. I search wide coarse ranges first, inspect where the good trials cluster, then spend a second round narrowed there. Cheaper scoring buys trials: fewer folds during the search, with the shortlist re-scored under the full protocol. I stop when best-so-far gains fall below the fold-to-fold variation of the score.

go deeper

for a junior

Know that every trial costs one fit per fold, so the budget is trials times folds times fit time, and that you should work that number out before launching a search.

for a middle

Be able to explain why a fixed budget suits random draws better than a grid, and why searching two axes wide beats searching six axes thinly at the same trial count.

for a senior

Show the operating plan: fix inert axes deliberately, score cheaply during the search and re-score finalists fully, run coarse then fine with an inspection between, and stop on a noise-based rule rather than on exhaustion.

for a principal

Own the allocation question. Argue explicitly whether the next block of compute belongs in tuning at all versus in features, labels or data quality, and set the team's convention for logging ranges, seeds and spend so searches compound instead of restarting.

## The budget is set by compute, not by the space The first thing to make explicit is the multiplication: trials x folds x fit time. Sixty trials, 5-fold cross-validation and a 30-minute fit is 300 fits and roughly 150 hours of serial compute — about six hours on 25 parallel workers. If nobody has done that arithmetic, the search plan is not a plan. This is also the strongest structural argument for random draws: a grid's cost is dictated by the axis sizes, so fitting it to a budget means uniformly coarsening every axis, whereas a random search takes the trial count as an input. ## Choose the axes before choosing the values With six hyperparameters and 60 trials, the temptation is to give each axis a range and let the sampler sort it out. It is usually better to cut the space first: - **Fix what is conventionally inert for the model family** at a documented default, and say in the write-up that it was fixed. Deciding not to search something is a decision you should be able to defend, not an omission. - **Search wide on what plausibly matters.** For most model families this is two or three axes: a capacity control, a regularisation strength, and a learning-rate-like quantity for iterative learners. - **Prefer a wide, coarse range to a narrow, precise one.** A narrow range that excludes the optimum cannot be rescued by more trials; a wide range that includes it can be refined in a second round. The payoff is directly countable. Sixty trials over two live axes give you 60 distinct values of each and a readable two-dimensional scatter of score against configuration. Sixty trials over six axes give the same 60 values per axis but a far sparser joint picture, and four of the six panels will be flat clouds you learn nothing from. ## Buy trials with cheaper scoring The fold count is a lever, not a constant. Scoring each trial on a single held-out split, or on 3 folds instead of 10, multiplies the number of configurations you can try for the same compute. The cost is a noisier score per trial, which makes the ranking of near-ties unreliable — so the standard pattern is a cheap score during the search to find the good region, then a full cross-validation on a shortlist of the top handful of configurations before committing. Similarly, searching on a stratified subsample of the training rows is often enough to rank configurations, with the finalists refit on everything. ## Two rounds beat one A coarse-then-fine split of the budget — say 40 trials over wide ranges, then 20 concentrated where the good trials clustered — is usually stronger than 60 trials at one resolution, and it gives you an inspection point in the middle. That inspection is where you catch the two failure signals: best values sitting on a range boundary (the range was wrong) and a score that is flat across an entire axis (that axis is not live and its budget should move elsewhere). ## Knowing when to stop Stop when the improvement in the best-so-far score is smaller than the fold-to-fold variation of the score itself. If the standard deviation across folds is 0.01 AUC and your last 20 trials moved the best score by 0.002, further trials are selecting on noise rather than finding a better model. The same yardstick answers the more important question: whether tuning is where the remaining value is at all. A search that has moved the metric by half a point while a data-quality problem or a missing feature is worth three points is a badly allocated budget, however well executed. ## What to write down Record every trial's full configuration and score, the ranges and scales searched, what was fixed and why, the seed, and the wall-clock spend. That record is what lets the next person start from your ranges instead of re-deriving them, and it is what makes a second round cheap. ## What an interviewer is listening for A number for the compute, an explicit decision about which axes are searched versus fixed, the fold count treated as a lever, a two-round plan with an inspection between rounds, and a stopping rule tied to the noise in the score rather than to the budget running out.

  • Would you cut cross-validation folds from 10 to 5 to double the number of trials?
    Usually yes during the search, then no at the end. Five folds doubles the configurations you can try at the cost of a noisier score per trial, which only matters for separating near-ties. So search cheaply, then re-score the top handful under the full protocol before choosing.
  • How do you know when to stop adding trials?
    Compare the improvement in the best-so-far score to the fold-to-fold variation of the score. Once gains are inside that noise band, extra trials are selecting on randomness. That is also the moment to ask whether more tuning beats spending the same compute on features or data quality.
  • How would you spend the budget differently if a single fit took 20 seconds instead of 30 minutes?
    The compute constraint largely disappears, so the binding limit becomes how reliably you can tell configurations apart. I would use more folds or repeated cross-validation per trial rather than more trials, since with cheap fits the noise in the score, not the number of configurations, is what limits the choice.

saying these in an interview costs you the question

  • Sizes the search from the space instead of the compute
  • Searches all six hyperparameters because they exist
  • Never states the wall-clock cost of the plan
  • Treats the fold count as fixed rather than a lever
  • Keeps searching for gains smaller than the score's noise
  • Chases a tiny metric gain while a data problem is unaddressed

context