Why is a model's error on its own training rows an optimistically biased estimate of future error?
answer
- the exam and the answer key
- the noise got fitted too
- one-sided bias, never pessimistic
- the gap has a name
- zero when parameters were fixed beforehand
basics
~20 sTraining error is optimistic because the fitting procedure tuned the model to those exact rows, absorbing their random noise as if it were signal. On fresh rows the noise does not repeat, so error rises.
solid answer
~50 sTraining error is measured on the same rows that were used to pick the parameters, so the number is contaminated by the fitting step. Any flexible fit will chase the particular noise in those rows — a quirk in one patient, a mislabelled record, a lucky coincidence between a feature and the target — and that part of the fit does not transport. Formally, for a training set drawn from some distribution, the expected error on the training rows is below the expected error on fresh rows drawn from the same distribution; the difference is the `optimism`, or the generalization gap. The bias is one-sided: training error is too good on average, never too bad, so a great training score is uninformative while a bad one is real news. The only honest number comes from rows that played no part in choosing the model.
go deeper
Be ready to say in one sentence why you never report the score measured on the rows you fitted on, and to name the gap between that score and future performance.
An interviewer expects the mechanism: the optimiser matched the noise in those specific rows, the bias is one-sided, and the size of the gap depends on how much freedom the fit had.
Show you use the asymmetry in practice — a bad training score is a real diagnostic signal you act on, a good one is never evidence you present to anyone.
Own the norm: define what counts as a reportable performance number in your organisation, so that no dashboard, model card, or executive slide ever shows a score computed on the fitting data.
## Two different quantities When you say "the model's error", you can mean one of two things, and interviews on this topic exist because people conflate them. - **Training error** (`err`): the average loss of the fitted model over the very rows used to fit it. If you fit on 500 rows and score on those same 500 rows, this is what you get. - **True (out-of-sample) risk** (`Err`): the expected loss of that same fitted model on a new row drawn from the same distribution the training rows came from. This is the number you actually care about, because it predicts what happens after deployment. The **optimism** is defined as the difference: ``` optimism = Err - err ``` and the empirical fact this leaf is about is that, in expectation, `optimism >= 0`. Training error is a downward-biased estimate of risk. ## Where the bias comes from Fitting is an *optimisation over the training rows*. Whatever knob the procedure has — a regression coefficient, a split threshold, a cluster centre, the choice of which features to keep — it is turned in whichever direction lowers the loss on those specific rows. Each row's target value has two parts: the systematic part that a future row with the same features would also show, and a random part specific to that row (measurement error, an unlucky day, an idiosyncratic customer). The optimiser cannot tell them apart. It lowers the loss by matching **both**. The systematic part transports to fresh rows; the random part does not, and on fresh data it turns from a discount into a penalty. A useful mental model: the training rows are both the exam and the answer key. Any procedure that studied the answer key will score higher on that exam than on a new one, and the more of the key it was able to memorise, the wider the difference. A formal version makes the mechanism explicit. For squared-error loss the expected optimism can be written as ``` E[optimism] = (2 / n) * sum_i Cov(yhat_i, y_i) ``` where `yhat_i` is the model's fitted value for row `i` and `y_i` is that row's observed target. The covariance term asks: *how much does the prediction for a row move when that row's own observed value moves?* A procedure that ignores the training labels has zero covariance and therefore no optimism. A procedure that lets each row pull its own fitted value towards itself — which is what fitting does — has positive covariance, and the optimism is exactly the price of that pull. ## The bias is one-sided, which is useful Because the optimism is non-negative in expectation, training error behaves as an optimistic *floor*: - A **low** training error tells you almost nothing. It is consistent with a model that generalizes beautifully and with one that memorised the data. - A **high** training error is informative. If the model cannot even fit the rows it was allowed to look at, it will not do better on fresh ones. This is how you diagnose underfitting, a broken feature, or a bug — the one legitimate everyday use of a training score. ## The extreme case: more columns than rows Consider a gene-expression classifier: 200 patients, 3,000 gene columns, and it reports **100% training accuracy**. That number carries no information at all. With far more columns than rows, a linear rule can separate almost any assignment of labels to those 200 patients, including labels shuffled at random. Perfect separation is a property of the model's capacity relative to the sample size, not evidence about biology. The same is true of a nearest-neighbour rule that predicts each training row by looking it up: its training error is exactly zero on every dataset ever collected, informative or not. ## When the optimism is small The gap is not always large. It shrinks when the fitting procedure has little freedom relative to the amount of data — a two-parameter model on a million rows has an optimism close to nothing — and it vanishes entirely when the model was fixed *before* the rows were seen, since then no parameter was tuned to them. It grows with the model's flexibility, with the noise in the target, and as the sample shrinks. ## What to say in an interview Name the phenomenon (optimism / generalization gap), give the mechanism in one line (the parameters were chosen to fit those rows' noise, which does not repeat), state that the bias is one-sided, and finish with the consequence: the only trustworthy performance number is measured on data that took no part in fitting or choosing the model.
- Is there any model for which the training error is an unbiased estimate of future error?Yes — one whose parameters were fixed before those rows were seen. If no knob was turned in response to the training labels, the fitted values carry no information about those rows' noise, the covariance between prediction and observed target is zero, and the training score is just an ordinary sample average. The moment you tune anything on the data, the optimism reappears.
- A model reaches only 62% training accuracy on a balanced two-class problem. Is that number as uninformative as a perfect one?No — it is genuinely bad news you can trust. Because the bias runs one way, training error sits at or below true error in expectation, so a model that cannot fit the data it was given will not do better on fresh data. Treat it as underfitting, a missing signal, a broken feature, or a bug, and fix that before worrying about generalization.
- Does shuffling the rows or refitting on a different random seed remove the optimism?No. The bias comes from scoring on the same rows that drove the fit, not from their order or from any randomness in the optimiser. Shuffling, reseeding, or refitting several times and averaging all still score the model on data it was fitted to, so every one of those numbers is optimistic by roughly the same amount.
saying these in an interview costs you the question
- Treats 100% training accuracy as proof the model works
- Thinks shuffling rows or changing the seed removes the bias
- Calls training error a safe or conservative estimate
- Believes the gap is a coding bug rather than a statistical fact
- Assumes more rows always fixes it even as the model grows