In the bias-variance decomposition of prediction error, what do bias and variance mean?
answer
- two different ways of being wrong
- retrain on a fresh sample, look again
- one is a systematic offset from truth
- the other is spread across refits, one input
basics
~20 sBias is the gap between the true value and the average prediction a model makes when it is retrained on different training samples. Variance is how much that prediction moves from one training sample to the next.
solid answer
~50 sBoth are defined at a single fixed input, and both are averages over the training sets the model could have been fitted on. **Bias** is the distance between the true value at that input and the average prediction across those refits - a systematic offset that would survive even if you averaged over infinitely many training sets of the same size. **Variance** is how far one individual refit's prediction sits from that average - pure instability, the part of the error caused by which rows you happened to train on. Expected squared error at the input is `bias^2 + variance + noise`. High bias means the model is wrong in the same direction every time; high variance means it is wrong in a different direction each time, and the two are diagnosed and fixed differently.
go deeper
Be ready to state both definitions in one breath and say what each looks like in practice: same wrong direction every refit versus a different wrong direction every refit. Screeners stop there.
Expect to be pushed on what the averaging runs over, and to write the decomposition down with the squared bias, the variance and the noise term named separately.
Show that you diagnose before you fix: say which observable evidence points to a systematic offset and which points to instability, and note that the diagnosis, not the label, drives the next move.
Own the framing itself. Be able to say when the bias-variance vocabulary earns its keep on a project and when it is being used as a slogan to justify a change nobody measured.
## The setup Fix one input - say the electricity demand of a city at 32 C, on a Tuesday, at 18:00. Two things are random in this picture, and keeping them straight is the whole point: 1. **The training set.** You fitted your model on the rows you happened to collect. A parallel universe collected a different sample from the same population and got a slightly different fitted model, and therefore a slightly different prediction at 32 C / Tuesday / 18:00. 2. **The label itself.** The demand actually observed on such an evening is not a fixed number either; it wobbles around some underlying true value for reasons no feature in your table records. Write `f(x)` for the underlying true value at that input, `g_D(x)` for the prediction made by the model fitted on training set `D`, and `gbar(x)` for the average of `g_D(x)` over all the training sets `D` you could have drawn. Then: ``` bias(x) = f(x) - gbar(x) variance(x) = average over D of (g_D(x) - gbar(x))^2 ``` and the expected squared error at that input is `bias(x)^2 + variance(x) + noise`. ## Bias: a systematic offset Bias is what is left over after you average away all the luck. If every refit of your demand model predicts roughly 15% below the truth at 32 C, the average prediction is 15% low, and no amount of re-drawing training samples of the same size will fix it. Bias is a property of the *procedure* - the family of shapes the model is allowed to take, the features it is given, the penalty it is fitted under - not of the particular dataset you hold. Two things bias is **not**: - It is not the average residual on the training data. A flexible model can drive its training residuals to nearly zero and still be badly biased at inputs it saw few examples of. - It is not sampling bias or fairness bias. The word is overloaded; here it means only the offset between truth and mean prediction. Bias can be positive or negative at a given input, and the decomposition uses its square, so offsets in either direction cost the same. ## Variance: instability across refits Variance is the spread of `g_D(x)` around its own average, at the *same* input. Concretely: refit the model ten times, each time on a different sample of the same size, and read off the ten predictions for 32 C / Tuesday / 18:00. If they run from 610 MW to 890 MW, that spread is variance - even if their average happens to land exactly on the truth. Variance measures how much the model is chasing the particular rows rather than the underlying relationship. The most common mistake here is to compute the spread of predictions *across the test set*. That is just how much demand varies from evening to evening; it says nothing about model stability. Variance in this decomposition is measured **at one input, across training sets**. ## Why they are worth separating Because they have different signatures and different cures. A high-bias model is wrong the same way every time: it will be wrong on new data in exactly the way it is wrong on the data you have, so more rows of the same kind will not rescue it - the model needs more capacity, better features, or a weaker penalty. A high-variance model is wrong in a fresh direction each time: its error is luck-driven, so it responds to constraints on the fitting procedure and to more data. Naming which one you are looking at is the first move in almost every model-debugging conversation. The two do tend to pull in opposite directions - the changes that let a model track the data more closely also let it track the noise more closely - but the tradeoff is a tendency, not an identity. Better features, or simply more informative data, can lower both at once. A candidate who claims that lowering one *must* raise the other is over-reading the name. ## The third term The decomposition has a third piece, the irreducible noise. Even a model that nailed the true function `f(x)` exactly would still miss the observed label, because the label carries variation the features do not explain. That term is a property of the target and the feature set, not of your model, and it is why an expected squared error of zero is not a coherent goal.
- Can a model have low bias and low variance at the same time, or must one always cost the other?Both can be low at once. The tradeoff is a tendency, not an identity: the knobs that let a model track the data more closely usually raise variance while lowering bias, but adding a genuinely informative feature, or collecting more data from the same population, can shrink both terms together. The tradeoff bites only when you are moving along a fixed budget of flexibility with the same features.
- Is the bias term the same thing as the average residual on the training data?No. The training residual measures how far the fitted model sits from the labels it was fitted on, which flexible models can drive near zero while still being systematically wrong elsewhere. Bias compares the true value at an input to the average prediction over many refits on fresh training samples - a quantity you never see from a single fit on a single dataset.
- Where does a model's bias at a given input actually come from?From everything fixed before the data arrives: the shape of function the model is allowed to express, the features it can see, and any penalty pulling its coefficients toward zero. If the true relationship curves and the model can only draw straight lines, the average prediction cannot follow the curve, and the shortfall is bias at every input where the curvature matters.
Think of a rifle sighted at a target. Bias is how far the average shot lands from the bullseye - the sights are off in a fixed direction. Variance is how widely the shots scatter around their own centre.
saying these in an interview costs you the question
- Calls variance the spread of predictions across the test set
- Defines bias as the average residual on the training data
- Confuses this bias with sampling bias or fairness bias
- Claims lowering one term must always raise the other
- Tries to measure variance from a single fitted model