skip to content

What determines how large the gap between training error and true error will be?

level: middleimportance: should knowfreq 55%

answer

  1. three levers, not one
  2. flexibility, noise, sample size
  3. count freedom, not columns
  4. roughly one leaf per row
  5. gap scales with df over n

basics

~20 s

Three things set the gap: the fitting procedure's effective degrees of freedom, the noise variance of the target, and the number of rows. The expected gap scales as degrees of freedom times noise, over rows.

solid answer

~50 s

The optimism is not a fixed penalty, it scales. For squared-error loss the expected gap is about `2 * df * sigma^2 / n`, where `df` is the effective degrees of freedom of the fitting procedure, `sigma^2` is the noise variance of the target and `n` is the number of training rows. So the gap grows with flexibility, grows with noise, and shrinks as rows are added — but only if the model does not grow along with the data. Effective degrees of freedom is the honest measure of flexibility: for least squares it equals the number of fitted coefficients; for a tree grown until every leaf is pure it is roughly one per leaf, which is one per row. That tree shows zero training error and the largest gap available, while a two-parameter model on a million rows has an optimism you can ignore.

go deeper

for a junior

Be able to say the gap grows when the model is more flexible or the data is noisier, and shrinks when there are more rows. The formula can wait.

for a middle

An interviewer expects the mechanics: the expected gap scales as degrees of freedom times noise variance over the number of rows, and effective degrees of freedom is what a flexible method really spends.

for a senior

Demonstrate that you estimate the gap before you measure it — from the flexibility you allowed and the noise in the target — and that you count freedom spent on thresholds and hand-picked features, not just fitted coefficients.

for a principal

Own the tradeoff when a team asks for a bigger model on a noisy target: extra capacity is paid for in optimism per row, so the decision is whether the data budget can fund it, not whether the model is fashionable.

## The gap has a size, and the size is predictable Saying "training error is optimistic" is only half the story; the interview question one layer down is *how* optimistic. For squared-error loss there is a clean expression. Write the expected optimism as ``` E[optimism] = (2 / n) * sum_i Cov(yhat_i, y_i) ``` and define the **effective degrees of freedom** of the fitting procedure as ``` df = (1 / sigma^2) * sum_i Cov(yhat_i, y_i) ``` where `sigma^2` is the variance of the noise in the target. Substituting gives the headline result: ``` E[optimism] = 2 * df * sigma^2 / n ``` Three levers, and only three: 1. **`df` — how much freedom the fit has.** Not the number of columns you happen to have, but how much each fitted value is allowed to follow its own observation. 2. **`sigma^2` — how noisy the target is.** In a low-noise problem there is little noise to overfit, so even a flexible model shows a small gap. In a very noisy one — churn, click-through, next-day returns — the same model overfits badly. 3. **`n` — how many rows.** The gap is per-row work spread over the sample. Doubling `n` with the model held fixed roughly halves the gap. ## What effective degrees of freedom actually counts The intuition behind `df` is: *if I nudge row i's observed target upward, how much does the model's own prediction for row i follow it?* Sum that responsiveness over all rows. - **Least squares with `d` fitted coefficients:** `df = d` exactly. This is why the classical "degrees of freedom" and the effective version agree for ordinary regression. - **A nearest-neighbour rule that predicts each row by its single closest neighbour:** each training row is its own nearest neighbour, so the prediction follows the observation one-for-one and `df = n`. - **A decision tree grown until every leaf is pure:** each leaf stores one fitted value, so `df` is roughly the number of leaves. Grow it all the way and that is about one leaf per row, giving `df` near `n` again. Plug it in: `E[optimism] ≈ 2 * sigma^2`, while the training error is exactly zero. The measured number and the truth have nothing to do with each other. The key move is that `df` counts *what the procedure actually spent*, not what a naive parameter count would suggest. A method can burn degrees of freedom through split searching, feature selection, or thresholds chosen by looking at the data, and every one of those is paid for in optimism. ## The noise-column demonstration The cleanest way to feel this is to add predictors that cannot possibly help. Take a 500-row regression and append 50 columns of pure random noise, uncorrelated with the target by construction. Refit and watch the training R-squared climb. Nothing was learned — the columns are noise — but each new coefficient is free to grab whatever slice of the residual happens to line up with it in these particular 500 rows. Under the null of no real signal, the expected training R-squared from `p` useless predictors on `n` rows is about `p / n`, so 50 noise columns on 500 rows buys roughly 0.10 of R-squared for nothing. The out-of-sample R-squared meanwhile drifts towards zero or below. The same arithmetic runs through the optimism formula: raising `df` by 50 on 500 rows raises the expected gap by `2 * 50 * sigma^2 / 500 = 0.2 * sigma^2`. Fifty free parameters cost you a fifth of the noise variance, every time, whether or not the columns mean anything. ## Practical consequences - **A perfect training score is a red flag proportional to `df`.** Zero training error from a fully grown tree or a memorising neighbour rule tells you `df ≈ n`, and therefore that the gap is as large as it can be. - **Noisy targets deserve simpler models.** Two teams with the same model and the same sample size can get very different gaps purely because one predicts a near-deterministic physical quantity and the other predicts human behaviour. - **Adding data helps only if the model stays put.** If you double the rows and simultaneously let the tree grow twice as deep or add twice as many features, `df / n` is unchanged and so is the optimism. - **Count the degrees of freedom you spent outside the fit too.** Thresholds you tuned by eye, features you kept because they looked promising in a plot, and variants you tried and discarded are all freedom spent on the same rows. ## Interview register Give the three levers, name effective degrees of freedom as the honest measure of flexibility rather than raw parameter count, and offer one concrete case — a fully grown tree with `df` near `n`, or the noise-column experiment — to prove you have actually seen it rather than only read about it.

  • Why does appending 50 pure-noise columns to a 500-row regression raise the training R-squared at all?
    Each extra coefficient is free to pick up whatever part of the residual happens to correlate with that column in these particular rows. Under no real signal, `p` useless predictors on `n` rows lift the expected training R-squared by roughly `p / n` — here about 0.10. Nothing was learned; the fit simply had 50 more chances to be lucky, and none of that luck repeats on new rows.
  • How would you measure effective degrees of freedom for a method that is not a linear regression?
    Use the definition rather than a parameter count: `df` is the sum over rows of how much the fitted value for a row moves when that row's own observed target moves, scaled by the noise variance. For least squares it collapses to the number of coefficients; for a one-nearest-neighbour rule the fitted value tracks the observation exactly, so `df` equals the number of rows.
  • Does collecting more training rows always shrink the gap?
    Only if the model's flexibility is held fixed, since the gap scales with degrees of freedom divided by rows. In practice more data usually tempts you into a bigger model — deeper trees, more features, more interactions — and if `df` grows in step with `n` the ratio is unchanged and the optimism does not improve at all.

saying these in an interview costs you the question

  • Equates degrees of freedom with the number of raw columns
  • Thinks useless predictors get exactly zero training benefit
  • Calls a zero-training-error tree a low-complexity model
  • Ignores target noise when judging overfitting risk
  • Says more data always fixes the gap

context