skip to content

Expected squared error splits into bias squared, variance and noise - what is the expectation over?

level: middleimportance: must knowfreq 62%

answer

  1. not the test set
  2. the training data is what is random
  3. one input, many parallel refits
  4. the procedure is fixed, the fitted model is not

basics

~20 s

The expectation runs over two random things: the training set the model was fitted on, drawn afresh from the same population, and the noise in the new label being predicted. The input point and the fitting procedure are held fixed.

solid answer

~50 s

It is an average over the training sets the model could have been fitted on, plus the noise in the label you are trying to predict. What is held fixed is the input at which error is measured and the *procedure* - the model family, the hyperparameters, the training-set size - not the fitted model, which is itself a random object because the data it saw was random. So the variance term is the spread of the prediction at one input across refits on independently drawn training samples, and the noise term is the variance of the observed label around the true value there. Total test MSE is the average of this pointwise decomposition over the distribution of inputs. In practice you approximate the expectation: refit on repeated resamples and read the spread of the predictions at a chosen input.

code

python · 17 lines
python
from statistics import mean

truth = 1.62  # true output power in MW at 8 m/s
# prediction at 8 m/s from each of ten refits on a resampled month
preds = [1.51, 1.74, 1.38, 1.66, 1.81, 1.44, 1.59, 1.70, 1.35, 1.62]

avg = mean(preds)                              # gbar at this input
bias_sq = (avg - truth) ** 2                   # systematic offset, squared
variance = mean((p - avg) ** 2 for p in preds) # spread across refits

print("mean prediction:", round(avg, 3))
print("squared bias:   ", round(bias_sq, 5))
print("variance:       ", round(variance, 5))
# mean prediction: 1.58
# squared bias:    0.0016
# variance:        0.02184
# the average refit is nearly right; the instability is what costs.

go deeper

for a junior

Know at least that the averaging is over different possible training samples, not over the rows you score on. Being able to say that much already puts you ahead of most screening answers.

for a middle

This is your tier. Write the three terms out, name what is fixed and what is random, and explain why exactly three terms survive when the square is expanded.

for a senior

Show you can operationalise it: describe how repeated resampling approximates the expectation, and be candid about why the estimate is optimistic and why bias is the hard term to measure.

for a principal

Be ready to say where this framing pays for itself on a real programme and where it becomes theatre - measuring parallel refits costs compute, and the decision it informs is often obvious from cheaper evidence.

## Two sources of randomness, one fixed point The decomposition is a statement about an average, and everything turns on what is being averaged. Take a wind-turbine power curve: a model that predicts output power from wind speed, and ask about its error at exactly 8 m/s. The expectation runs over: 1. **The training set `D`.** You have ten months of turbine telemetry; imagine drawing ten *other* months from the same regime. Each draw produces a different fitted curve and therefore a different predicted power at 8 m/s. The fitted model is a random object because its input data was random. 2. **The noise in the new label.** The power actually recorded on a future 8 m/s reading is not a fixed number; turbulence, air density and sensor error move it around some underlying true value. What is **fixed**: - **The input.** The decomposition is stated pointwise, at 8 m/s. Fix a different wind speed and you get different bias and variance numbers. - **The procedure.** The model family, the hyperparameters, and the number of training rows are all held constant. Change any of them and you are describing a different estimator with a different decomposition. Writing `g_D(x)` for the prediction of the model fitted on `D`, `gbar(x)` for its average over draws of `D`, and `f(x)` for the true value: ``` E_{D, noise} [ (y - g_D(x))^2 ] = (f(x) - gbar(x))^2 + E_D[(g_D(x) - gbar(x))^2] + noise variance ^ squared bias ^ variance ^ irreducible ``` ## What the cross terms do The reason exactly three terms come out is algebraic. Split the error `y - g_D(x)` into `(y - f(x)) + (f(x) - gbar(x)) + (gbar(x) - g_D(x))`, square it, and take the expectation. The three squares give the three terms. Every cross term dies: the label noise has mean zero and is independent of the training set, and `g_D(x) - gbar(x)` has mean zero over `D` by the definition of `gbar`. Nothing else is assumed - the noise does not need to be normal, and the model does not need to be linear or fitted under squared error. ## Common confusions this exposes - **"Variance is how much the predictions vary across the test set."** No. That is how much the target varies from case to case, which would be large even for a perfect model. Variance in this decomposition is measured at a *single* input, across training sets. - **"Bias is the average error on the training data."** No. Bias compares truth to the mean prediction over refits, a quantity that cannot be read off one fit. - **"It only applies at one point, so it is useless."** It aggregates: average the pointwise decomposition over the distribution of inputs, and total expected squared error is average squared bias plus average variance plus average noise. Each term keeps its meaning. ## Estimating it when you only have one dataset You never observe the expectation, but you can approximate it. Draw many resamples of your data - bootstrap resamples, or repeated random splits of the same size - refit the same procedure on each, and collect the prediction at a chosen input. The mean of those predictions estimates `gbar(x)`; their spread estimates the variance term; the gap between their mean and a known or well-estimated truth estimates the bias. Two honest caveats: resamples of one dataset overlap heavily, so the spread underestimates the true variance across genuinely independent samples, and you rarely know `f(x)`, so the bias term is usually the hardest one to pin down outside a simulation where you generated the truth yourself. This is why the decomposition is most often used as a *conceptual* instrument - a way to reason about which of the two failure modes a change to the procedure will attack - rather than as three numbers reported on a dashboard. ## Why 'the procedure, not the fit' matters in interviews The crispest way to show you understand this is to say what changes when you tune. Halve the training-set size and the variance term grows even though the model family is unchanged, because each draw of `D` now pins the fit down less. Tighten a penalty on the coefficients and the variance term shrinks while the average prediction moves away from the truth, raising bias. Both statements are about the procedure averaged over data, not about the single model sitting in your notebook - and that is the distinction the question is really probing.

  • How would you actually estimate the variance term when you only have one dataset?
    Refit the same procedure many times on resamples of the data - bootstrap resamples or repeated splits of equal size - and read the spread of the predictions at a chosen input. Two caveats: the resamples overlap heavily, so the spread understates the variance across genuinely independent samples, and without knowing the true value you cannot separate bias from it, which is why full decompositions are usually demonstrated on simulated data.
  • Does the decomposition hold for overall test MSE, or only at a single input?
    Both. It is derived pointwise, but averaging the pointwise identity over the distribution of inputs gives total expected squared error as average squared bias plus average variance plus average noise. The aggregate can hide structure, though: a model may be badly biased in a sparse region and merely unstable elsewhere, and only the pointwise view shows that.
  • If you halve the training-set size, which terms move and in which direction?
    Variance rises, because each draw of the training data pins the fit down less and the predictions at a fixed input scatter more. The noise term is untouched - it belongs to the label, not the model. Bias is roughly unchanged for a fixed model family, since the average prediction is still constrained by what that family can express, though heavily penalised procedures can shift it too.

saying these in an interview costs you the question

  • Says the expectation is over rows of the test set
  • Treats the fitted model as fixed and the input as random
  • Assumes the noise must be normally distributed for the split to hold
  • Thinks the decomposition only works for models fitted under squared error
  • Cannot say what changes when the training-set size changes

context