skip to content

Which tuned hyperparameters stop being correct once you refit on the full training set?

level: seniorimportance: should knowfreq 42%

answer

  1. tuned at 80% of the rows
  2. some optima move with n
  3. counts do not rescale, fractions do
  4. more data supports more capacity
  5. boosting rounds times k over k minus one

basics

~20 s

The ones whose best value depends on how many rows you trained on: the number of boosting rounds, the neighbour count k in k-nearest-neighbours, penalty strength, and any leaf-size or split threshold written as an absolute row count.

solid answer

~50 s

Cross-validation tunes on `(k-1)/k` of the data, so any setting whose optimum moves with sample size arrives at the refit with a stale value. The classic case is the number of boosting rounds: chosen on 80% of the rows, it is typically a little too small once the model trains on 100%, because more data supports more rounds before overfitting. `k` in a k-nearest-neighbour model tuned at n=8,000 and shipped at n=10,000 is similarly slightly too small — the optimal neighbourhood size grows with n. Regularization strength usually wants to be a touch weaker with more rows. And any constraint written as an absolute count, such as a minimum of 20 samples per leaf, silently becomes a looser brake at larger n. Two habits fix this: express size-sensitive settings as fractions of n, and rescale or re-derive the rest rather than copying them.

go deeper

for a junior

Know that hyperparameters are tuned on only part of the data, and that a few of them — such as the number of boosting rounds — depend on how many rows were used for training.

for a middle

Explain which settings move with sample size and in which direction: more rows generally support more boosting rounds, a larger neighbour count, and a weaker penalty.

for a senior

Demonstrate the habit of auditing the setting list before a refit, separating size-sensitive values from size-neutral ones, and rescaling or re-deriving the sensitive ones rather than copying them.

for a principal

Own the convention: search spaces written in proportional terms wherever the learner allows it, and a documented correction rule for the settings that cannot be, so refits are reproducible across a team instead of folklore.

## The gap nobody looks at Selection happens at one training-set size and deployment happens at another. With 5-fold cross-validation every candidate was judged by models trained on 80% of the rows; the artifact you ship trains on 100%. That is a 25% increase in training data between the measurement and the thing measured. Most hyperparameters do not care. A few care a lot, and copying those across is one of the quieter ways a refit ends up worse than the fold models it replaced. ## Which settings move with n, and in which direction **Number of boosting rounds (gradient-boosted trees).** Each round adds capacity. With more training data, the point at which additional rounds start fitting noise arrives later, so the optimal count generally **rises** with n. A count tuned on 80% of the rows therefore tends to underfit the full refit. A widely used heuristic is to scale the tuned count by `k/(k-1)` — a factor of 1.25 at k=5, 1.11 at k=10 — on the argument that the refit sees that much more data. It is a rough correction, not a law: with a small learning rate the loss curve is flat near its minimum and the exact count barely matters, while with an aggressive learning rate it matters a great deal. **Neighbour count k in a nearest-neighbour model.** For consistency the neighbourhood must grow in absolute size while shrinking as a share of the sample — `k` increases with n while `k/n` goes to zero. Practically, a k tuned at n=8,000 is modestly too small at n=10,000: the same k now spans a physically smaller region, giving a slightly noisier, lower-bias estimate than the tuning suggested. Scaling with the square root of n is the usual rule of thumb, which would nudge k up by roughly 12% for that 25% increase in rows. **Regularization strength.** A penalty exists to control variance. Variance falls as n grows, so the optimal penalty typically **weakens** with more data — the fold-tuned value is usually a little too strong for the refit. The effect is mild for a 25% jump and material when the jump is large. **Anything written as an absolute count.** A minimum of 20 samples per leaf is a constraint of 20/6,400 of the data during a fold and 20/8,000 after the refit: the same number is a *weaker* brake at larger n, so the refit tree grows deeper than the tuning implied. Early-stopping windows measured in rounds, minimum samples per split, and minimum counts for a category to survive encoding all share this property. ## Which settings travel unchanged Settings that describe the *shape* of the hypothesis rather than how much evidence supports it: - the learning rate and the loss function, - the distance metric or kernel choice, - the encoding strategy for categorical features, - the number of trees in a random forest — more trees reduce prediction variance and do not overfit, they only cost time, - anything already expressed as a fraction: a minimum leaf as a share of n, a row-subsample or column-subsample rate, a validation share. That last bullet is the design lesson. **A setting expressed as a proportion rescales itself; a setting expressed as a count does not.** Where a learner offers both forms, prefer the proportional one and the whole problem disappears. ## Three ways to handle the size-sensitive ones 1. **Rescale by the size ratio.** Multiply round counts by `k/(k-1)`, nudge k for neighbours by the square-root rule, ease the penalty slightly. Cheap, approximate, and much better than doing nothing. 2. **Re-derive on the full set.** Carve a small internal validation slice out of the full training data and read the size-sensitive value off it directly, keeping every other setting at its cross-validated value. Costs one extra fit and gives the value at the correct sample size. 3. **Tune only proportional forms.** Restate the search space so nothing is size-dependent in the first place. This is the version that survives contact with a team, because it needs no one to remember the correction. ## The tell in an interview A candidate who says "the hyperparameters transfer, that is the whole point of tuning" has never watched a refit get worse. A candidate who says "I audit the setting list before the refit and split it into size-sensitive and size-neutral" has. The direction matters too: more data supports **more** capacity and **less** penalty. Getting that backwards — claiming a bigger training set needs stronger regularization — is the misconception this question is built to catch. ## Keeping it in proportion For a 25% increase in rows the corrections are usually small, and on a flat loss surface they may be invisible. The point is not that the refit will fail without them; it is that you know which numbers are assumptions about sample size and can say so. On a bigger jump — say tuning on a 10% subsample for speed and refitting on everything — the same corrections stop being cosmetic and start deciding whether the model works.

  • What is a practical rule for scaling the boosting round count from k-fold tuning to the full refit?
    Multiply the tuned count by k/(k-1) — 1.25 for 5-fold — on the reasoning that the refit sees that factor more data. It is a heuristic, not a law: with a small learning rate the loss curve is flat near its optimum and the exact count hardly matters, while with a large one it does. If a refit is cheap, re-derive the count on the full training set instead.
  • Which hyperparameters are safe to copy across unchanged?
    Those describing the shape of the model rather than how much data supports it: the learning rate, the loss, the distance metric, categorical encoding choices, and the number of trees in a random forest, where more trees only cost time. Anything already written as a fraction of n — a subsample rate, a minimum leaf as a share of rows — also travels, because it rescales itself.
  • You tuned on a 10% subsample for speed and now refit on everything. Is this still a minor correction?
    No. A tenfold jump in training data moves the size-sensitive optima far enough that copying is unsafe: the round count is badly under-set, the penalty far too strong, and count-based constraints effectively toothless. At that gap, re-derive the size-sensitive settings on the full data rather than rescaling, and treat the subsample tuning as a coarse search that narrowed the range.

saying these in an interview costs you the question

  • Assumes every hyperparameter transfers unchanged to more data
  • Copies the fold-tuned boosting round count blindly
  • Claims more training data calls for stronger regularization
  • Writes leaf-size limits as raw counts and never revisits them
  • Retunes everything on the full set, discarding the selection

context