skip to content

Bias-Variance and Generalization

You will learn why models fail — too simple or too flexible — and how to tell which from learning curves and split scores. 'Explain the bias-variance tradeoff' opens more ML screens than any other.

on this pageshow

explore

questions

28

What is model capacity, and why do underfitting and overfitting sit at its two ends?

level: juniorimportance: must knowfreq 88%

answer

  1. size of the space of possible fits
  2. two opposite ways of being wrong
  3. degree 1 versus degree 15
  4. both errors high, or a big gap

basics

~20 s

Capacity is how large and flexible a model's space of possible fitted functions is. Too little capacity and the model cannot represent the real pattern, which is underfitting. Too much and it reproduces the training sample's noise, which is overfitting.

solid answer

~50 s

Capacity is the size of the hypothesis space a learning procedure can search: how many distinct functions it could end up representing. Raising the polynomial degree, raising a tree's maximum depth, or adding features all enlarge that space. At the low end the space contains nothing close to the true relationship, so the model is wrong in the same direction everywhere and training and test error are both high and close together: underfitting. At the high end the space is rich enough to thread the particular sample you happened to draw, including its noise, so training error collapses while test error climbs: overfitting. Fitting a straight line to 60 apartment-rent listings against square metres underfits if rent bends with size; a degree-15 polynomial through the same 60 points passes near every listing and swings wildly between them. The dial in between is what you are actually choosing.

go deeper

for a junior

Be ready to define underfitting and overfitting in one sentence each and to name a dial that moves between them, such as polynomial degree or tree depth. Knowing the training-versus-held-out error signature of each is expected.

for a middle

Explain capacity as the size of the hypothesis space rather than as a vague notion of complexity, and say why a low-capacity model is wrong consistently while a high-capacity one is wrong differently on every sample.

for a senior

Show that you check per-segment error rather than a single aggregate before declaring which end you are on, and that you separate the declared capacity limit from what was actually realised on the data you had.

for a principal

Own the framing that capacity is chosen relative to the data volume and noise level of a specific problem, not inherited as a house default, and be able to explain that tradeoff to people who only see the headline accuracy.

## What capacity actually is A learning algorithm does not consider every conceivable function. It searches a **hypothesis space** - the set of functions the model class can express once its parameters are fitted. Linear regression on one feature can express every straight line and nothing else. A degree-15 polynomial can express every straight line *and* every wiggly curve up to degree 15. A decision tree limited to depth 3 can express at most 8 piecewise-constant regions; a tree allowed depth 30 can carve the input space into far more. **Capacity** is the size or richness of that space. It is the single most useful dial in supervised learning because almost every generalization failure is a statement about it being set wrong. A subtlety worth carrying: capacity is a property of the model *class*, not of the fitted model and not of the data. Fitting a degree-15 polynomial to 60 listings does not make the class smaller; you have merely picked one member of a very large family. This is why capacity is normally discussed before you look at any result. ## The two ends of the dial **Underfitting (the low-capacity end).** The hypothesis space simply does not contain a function close to the true relationship. No amount of optimisation helps, because the best available member is still a poor approximation. The signature is that training error is high *and* test error is high, and the two are close to each other. A depth-3 tree on 5,000 rental listings gives you eight rent predictions for the entire market; if rent genuinely depends on size, district, floor and age, eight buckets cannot express it, and the model is systematically wrong in the same direction for whole groups of listings. **Overfitting (the high-capacity end).** The space is rich enough that the procedure can find a function which follows the particular sample you drew, including the part of that sample that is noise - a landlord who priced oddly, a listing with a typo in the floor area. The signature is a large gap: training error very low, test error much higher. A degree-15 polynomial fitted to 60 listings can be made to pass essentially through every point; between the points it does whatever the algebra demands, which can mean predicting negative rent for a 45-square-metre flat because two nearby listings pulled the curve apart. ## Which dials change capacity - **Polynomial degree** - the classic textbook dial: degree 1 versus degree 15 on the same 60 listings. - **Tree depth and number of leaves** - a depth-30 tree versus a depth-3 stump on the same 5,000 listings. - **The number of input features** - each added feature enlarges the space of expressible functions. - **The number of free parameters in general** - useful as a rough intuition, though it is not a reliable measure; some one-parameter families can express astonishingly many labelings. ## Why more is not better The reason the high end fails is that you fit on a *sample*. Any finite sample carries structure that belongs to the sample rather than the population. A low-capacity model cannot chase that structure because it has nowhere to put it; a high-capacity model can, and optimisation will happily do it, because a training loss cannot tell signal from noise. That is why the two ends fail in qualitatively different ways: the low end is wrong in a stable, repeatable way, and the high end is wrong in a different way on every resampled training set. ## How to tell which end you are on Compare training error with held-out error and look at both the level and the gap: - both high, gap small - low-capacity end; - training low, held-out much higher - high-capacity end; - both low - you are somewhere sensible, assuming the held-out data really is held out. A caution: both errors being high is only *evidence* of underfitting. If the features carry no information about the target, or if the labels are intrinsically noisy, no capacity setting fixes it, and adding capacity in that situation just moves you to the other failure mode. ## Effective capacity What the model class *could* express and what it *does* express on a given dataset are not the same. Declaring a maximum depth of 30 on 5,000 listings does not produce a fully grown depth-30 tree - branches run out of examples long before that. The declared limit is an upper bound on capacity; what you actually realise depends on the data you fitted. People sometimes call this the effective capacity, and it is why the same setting can be reckless on a small dataset and conservative on a large one.

  • Training and held-out error are both high and almost equal. Which end of the dial are you on?
    The low-capacity end - the model class cannot express the pattern, so it is wrong in the same way on data it has seen and data it has not. The one alternative worth ruling out is that the features carry no signal about the target, or the labels are so noisy that no model class does better; in that case adding capacity buys nothing.
  • Does adding input features increase capacity?
    Yes. Each extra feature widens the set of functions the model can express, so the hypothesis space grows even if the number of rows does not. That is why a wide table with few rows overfits easily: capacity has gone up while the evidence available to pin down the fit has not.
  • Can one model underfit and overfit at the same time?
    Yes, in different regions of the input space. A tree can be starved of splits in a dense region of ordinary flats and simultaneously grow single-listing leaves out in the luxury tail. Aggregate training and test numbers hide this, which is why per-segment error is worth looking at before concluding the dial is set wrongly overall.

Think of fitting as drawing a line through pins on a board. A rigid steel ruler makes a straight line no matter where the pins are and misses most of them; a floppy wire threads every pin exactly, including the ones that were pushed in slightly wrong.

saying these in an interview costs you the question

  • Says overfitting means the model is too accurate
  • Equates capacity with the number of parameters
  • Claims more training data cures underfitting
  • Diagnoses overfitting from training error alone
  • Assumes any complex model must overfit

context

open as a page

Why does a held-out score only estimate future performance if the data is i.i.d.?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A held-out score is an average of errors over a sample of rows. It estimates future error only if future rows are independent draws from that same distribution. When the population or the relationship moves, the number stops applying.

open as a page

In the bias-variance decomposition of prediction error, what do bias and variance mean?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Bias is the gap between the true value and the average prediction a model makes when it is retrained on different training samples. Variance is how much that prediction moves from one training sample to the next.

open as a page

Why is a model's error on its own training rows an optimistically biased estimate of future error?

level: juniorimportance: must knowfreq 85%

basics

~20 s

Training error is optimistic because the fitting procedure tuned the model to those exact rows, absorbing their random noise as if it were signal. On fresh rows the noise does not repeat, so error rises.

open as a page

What does a validation curve plot, and how do you read it to pick a hyperparameter value?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A validation curve plots training score and held-out validation score against one hyperparameter that controls model capacity. Pick the value where validation peaks; a widening gap between the two lines past that point is the overfitting signal.

open as a page

Why does training error fall monotonically as capacity grows while test error is U-shaped?

level: middleimportance: must knowfreq 74%

basics

~20 s

Enlarging a nested hypothesis space keeps every earlier fit available, so training error can only fall. Test error instead trades falling systematic error against rising sensitivity to the sample, so it bottoms out and climbs.

open as a page

How do covariate shift, label shift and concept drift differ in what they break?

level: middleimportance: must knowfreq 54%

basics

~20 s

Covariate shift moves the input distribution p(x) while p(y|x) holds. Label shift moves the class balance p(y) while p(x|y) holds. Concept drift moves p(y|x) itself, which is the only one that makes the model wrong.

open as a page

On a learning curve over training-set size, what do converged curves versus a persistent gap tell you?

level: middleimportance: must knowfreq 72%

basics

~20 s

Curves that meet at a high error mean high bias: the model is too simple, and more rows will not help. A gap that stays wide at every training size means high variance, where more rows still help.

open as a page

A model hits 2% training error but 14% validation error - what do you try first?

level: middleimportance: must knowfreq 78%

basics

~20 s

A 12-point train-validation gap is high variance, not high bias. Rank the fixes by cost: more training rows, then a stronger penalty or fewer features, then a lower-capacity model. Adding capacity would make the gap worse.

open as a page

Expected squared error splits into bias squared, variance and noise - what is the expectation over?

level: middleimportance: must knowfreq 62%

basics

~20 s

The expectation runs over two random things: the training set the model was fitted on, drawn afresh from the same population, and the noise in the new label being predicted. The input point and the fitting procedure are held fixed.

open as a page

Why does a model's held-out error flatten above zero no matter how many training rows you add?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Part of the error is irreducible: the label is not fully determined by the features, so near-identical inputs carry different labels. Data cannot remove that. The rest of the floor is the model's own bias, which rows also cannot fix.

open as a page

Why must every learning algorithm carry an inductive bias to generalize?

level: juniorimportance: should knowfreq 38%

basics

~20 s

An inductive bias is the assumptions a learner uses to choose among hypotheses that fit the training data equally well. Data alone never picks one, so without a bias there is no basis for predicting a new row.

open as a page

What does the no-free-lunch theorem claim about comparing learning algorithms?

level: middleimportance: should knowfreq 42%

basics

~20 s

The no-free-lunch theorem says that averaged uniformly over every possible target function, all learning algorithms have the same expected error on unseen inputs. No algorithm dominates in general; an advantage only exists relative to a restricted class of problems.

open as a page

What determines how large the gap between training error and true error will be?

level: middleimportance: should knowfreq 55%

basics

~20 s

Three things set the gap: the fitting procedure's effective degrees of freedom, the noise variance of the target, and the number of rows. The expected gap scales as degrees of freedom times noise, over rows.

open as a page

A k-NN validation curve over k = 1 to 200 shows zero training error at k = 1. Which end is high capacity?

level: middleimportance: should knowfreq 45%

basics

~20 s

The small-k end. At k = 1 every training point is its own nearest neighbour, so training error is zero by construction. Capacity falls as k grows, so this curve's complexity axis runs right to left.

open as a page

How does training-set size change the model capacity you should choose?

level: seniorimportance: should knowfreq 45%

basics

~10 s

More training data lets you afford more capacity. Sensitivity to the sample shrinks as the sample grows while a class's systematic error does not, so the best capacity moves upward with dataset size.

open as a page

How do you construct a learning curve over training-set size so its shape is trustworthy?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Vary only the training-set size. Score every point against one fixed held-out set, draw several random subsamples per size and average them, space sizes logarithmically, and refit preprocessing inside each subsample so nothing leaks.

open as a page

A speech model shows 9% training and 10% validation word error rate, humans 6% - what next?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Most of the error is avoidable bias: three points sit between the model's 9% training error and the 6% human benchmark, against a one-point train-validation gap. Attack bias first - more capacity, richer features, weaker regularization - not more data.

open as a page

Duplicate assays of one water sample differ by 0.3 mg/L - what does that imply for a model predicting the assay?

level: seniorimportance: should knowfreq 45%

basics

~20 s

That disagreement estimates the irreducible-noise term of the error decomposition. No model of the recorded features can beat it on average, so it sets the floor your error should be judged against rather than against zero.

open as a page

A trading rule's backtest Sharpe is 2.5 in-sample but 0.3 the following year. What happened?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The in-sample Sharpe was never a forecast: it is the best of many rules scored on the very history that selected them. A maximum over noisy statistics is biased upward, so the winner's number was inflated from the start.

open as a page

A 2019 credit scorecard now scores a much wider 2021 applicant pool — do you still trust its held-out AUC?

level: principalimportance: should knowfreq 34%

basics

~20 s

Not as a statement about 2021. The 2019 held-out AUC describes the 2019 applicant mix. Re-estimate it under the new mix where the two populations overlap, and cap exposure where 2019 gave no coverage at all.

open as a page

What does the VC dimension of a classifier class measure, and why is a line in the plane 3?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

VC dimension is the largest number of points a classifier class can label in every possible way. Some 3 non-collinear points admit all 8 labelings by a line, but no 4 points do, so the answer is 3.

open as a page

If 20,000 reviews come from just 200 sellers, how many independent observations do you have?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Far fewer than 20,000. Reviews from one seller are correlated, so they repeat information rather than adding it. With a design effect near 20 the sample behaves like about a thousand independent rows, so every error bar widens.

open as a page

Why does the bias-variance decomposition not add up cleanly under 0-1 classification loss?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Squared error is quadratic, so expanding it leaves exactly three additive non-negative terms. Counting misclassifications is a step function of the prediction error, so no such algebra exists, and instability across refits can even lower the error rate.

open as a page

Your validation curve is still improving at the largest hyperparameter value you swept — what do you do?

level: seniorimportance: nice to knowfreq 32%

basics

~10 s

An edge value is a property of your grid, not of the model: the sweep ended before the optimum. Extend the range and refit until validation clearly flattens or turns over, then choose.

open as a page

Given a learning curve over training-set size, how do you decide whether 20,000 more labels are worth buying?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Extrapolate the curve one or two doublings at most, price the projected error drop in hours and money, and ask whether it changes a decision the product makes. Buy a small tranche first to test the projection.

open as a page

A colleague says boosted trees are always the best model on tabular data — how do you respond?

level: principalimportance: nice to knowfreq 30%

basics

~10 s

No algorithm is best on every problem, so the claim is empirical — a statement about the problems that team has seen. Ask what evidence supports the default and which observable conditions void it.

open as a page

Expert annotators agree only 92% of the time - how much headroom does your model really have?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Pairwise agreement of 92% is not an 8% error floor. If annotators err independently, each is wrong on roughly 4% of items. Adjudicate a sample to split irreducible ambiguity from fixable guideline drift before funding more modelling.

open as a page