skip to content

Your validation curve is still improving at the largest hyperparameter value you swept — what do you do?

level: seniorimportance: nice to knowfreq 32%

answer

  1. the edge is your grid, not the optimum
  2. extend by doubling, not by nudging
  3. compare the gain to fold-to-fold jitter
  4. a flat tail can mean saturation

basics

~10 s

An edge value is a property of your grid, not of the model: the sweep ended before the optimum. Extend the range and refit until validation clearly flattens or turns over, then choose.

solid answer

~50 s

A curve still rising at the edge of the grid has not shown you an optimum; it has shown you where you stopped looking. The response is to extend the sweep by doubling rather than nudging, until validation visibly flattens or turns down. Two checks come with that. First, confirm the last gains are real by looking at how much the score jitters across folds and seeds — a rise smaller than that jitter is not a rise. Second, watch for dials that saturate instead of turning over: past the depth at which every leaf is already pure or blocked by a minimum-size constraint, deeper trees are literally the same model, so the tail flattens by construction. If the curve keeps improving past what you can afford to serve, report that as a resource-bounded choice.

go deeper

for a junior

Know that if the best score sits at the end of the range you tried, you have not found the best setting — you have found the edge of your search. Say you would try larger values.

for a middle

Explain how you would extend the sweep and what the three possible shapes mean: a turn-over, a flat tail, or continued improvement. Be ready to say why doubling beats stepping.

for a senior

Show that you check the last gains against fold-to-fold noise before acting, recognise a saturating dial from a genuinely tuned one, and weigh training and serving cost when the curve is flat.

for a principal

Own the review norm that every reported hyperparameter is quoted with the range it was searched over, so boundary artefacts are visible to reviewers instead of surviving into production model cards.

## What an edge maximum actually means When the best validation score on a sweep sits at the first or last value you tried, the plot contains no evidence about what lies beyond it. The apparent optimum is an artefact of the grid boundary. Reporting "depth 8 is best" when 8 was the largest depth you swept is a statement about your search range, not about the model. The correction is mechanical: extend the range in the direction the curve is still improving and refit. Extend multiplicatively — 8, then 16, then 32 — rather than nudging one step at a time, because you are looking for where the curve turns, and stepping cautiously turns one diagnosis into ten fits. ## Three things the extended curve can show you **It turns over.** The best case: validation rises, peaks and declines. You now have a genuine interior optimum and can pick it. **It flattens.** Validation stops improving and stays flat. This is common and usually means the dial has stopped biting rather than that you are perfectly tuned. The clearest example is tree depth: once every leaf is already pure, or every leaf has hit the minimum-samples constraint, the tree stops growing on its own, and raising the maximum depth from 12 to 30 produces the *identical* model. A dead-flat right tail on such a dial is a construction artefact. When that is what you see, take the smallest value on the flat region — larger settings buy nothing and cost training time, memory and inference latency. **It keeps improving past what you can afford.** Sometimes capacity really does keep paying, but each step doubles training cost or pushes the model past a latency budget. That is a legitimate stopping point, but it should be reported honestly as a resource-bounded choice: "validation was still improving at the largest setting we could serve," not "we tuned this dial." ## Distinguishing a real gain from noise at the edge The scores at the far end of a sweep are usually the noisiest, because high-capacity models vary more from split to split. Before chasing an apparent rise, look at the spread of the per-fold scores at those points and repeat the sweep under a different random seed. If the improvement from one value to the next is smaller than the movement you see across folds and seeds, the curve is flat there and you are extending the grid to chase noise. This is how a well-intentioned "just try a bit further" turns into a model that is tuned to the validation split. The symmetric mistake exists at the other edge. If the best value is the *smallest* one you swept — the simplest model on the grid — the sweep is equally uninformative, and the model may want to be simpler still. On a dial whose capacity runs backwards, such as the number of neighbours, an edge maximum at the largest value means the model wants *more* smoothing, not more flexibility. ## The honest reporting habit Always state the swept range alongside the chosen value: "depth 5, swept over 1 to 30" is a reviewable claim; "depth 5" is not. A reviewer who can see the range immediately knows whether the optimum was interior. Making the range explicit in a model card or an experiment log is what lets someone else catch a boundary artefact months later. ## Where the boundary judgment lands Extending a grid is cheap on a small tabular problem and expensive on a large one, so the decision is a cost call as much as a statistical one. A reasonable default: extend once, aggressively, in the direction of improvement. If the curve turns over, you are done. If it flattens, take the cheap end of the flat region. If it is still climbing after one aggressive extension, stop turning this dial and ask whether capacity is really the binding constraint — a model that wants unbounded capacity is often signalling something about the features or the target rather than about the hyperparameter.

  • How do you tell a real improvement at the edge of the sweep from noise?
    Look at how much the score moves across folds and across random seeds at those settings, then compare the step you are excited about to that movement. High-capacity settings are the noisiest points on the curve, so a gain smaller than the fold-to-fold spread is not a gain. Repeating the sweep under a different seed and seeing the same shape is the cheap confirmation.
  • A tree-depth curve is completely flat from depth 12 to depth 30. What is the likely explanation?
    The trees are not actually reaching those depths. Once every leaf is pure or has hit the minimum-samples constraint, growth stops on its own, so raising the cap produces the identical model and the identical score. Take depth 12, or better the smallest value on the flat region — the larger settings buy nothing and only cost training time and memory.
  • You extend the range and validation still improves at the largest setting you can afford to train. What do you report?
    That the choice is resource-bounded, not tuned: state the range swept, the setting chosen, and that validation was still rising at the boundary. That framing keeps the door open for someone with a bigger budget and stops a reader inferring an optimum that the experiment never demonstrated.

saying these in an interview costs you the question

  • Ships the largest value swept because it scored best
  • Treats the edge of the grid as the model's optimum
  • Nudges the range outward one step at a time indefinitely
  • Reads a flat tail as proof the dial is well tuned
  • Reports a chosen value without saying what range was swept

context