skip to content

What is model capacity, and why do underfitting and overfitting sit at its two ends?

level: juniorimportance: must knowfreq 88%

answer

  1. size of the space of possible fits
  2. two opposite ways of being wrong
  3. degree 1 versus degree 15
  4. both errors high, or a big gap

basics

~20 s

Capacity is how large and flexible a model's space of possible fitted functions is. Too little capacity and the model cannot represent the real pattern, which is underfitting. Too much and it reproduces the training sample's noise, which is overfitting.

solid answer

~50 s

Capacity is the size of the hypothesis space a learning procedure can search: how many distinct functions it could end up representing. Raising the polynomial degree, raising a tree's maximum depth, or adding features all enlarge that space. At the low end the space contains nothing close to the true relationship, so the model is wrong in the same direction everywhere and training and test error are both high and close together: underfitting. At the high end the space is rich enough to thread the particular sample you happened to draw, including its noise, so training error collapses while test error climbs: overfitting. Fitting a straight line to 60 apartment-rent listings against square metres underfits if rent bends with size; a degree-15 polynomial through the same 60 points passes near every listing and swings wildly between them. The dial in between is what you are actually choosing.

go deeper

for a junior

Be ready to define underfitting and overfitting in one sentence each and to name a dial that moves between them, such as polynomial degree or tree depth. Knowing the training-versus-held-out error signature of each is expected.

for a middle

Explain capacity as the size of the hypothesis space rather than as a vague notion of complexity, and say why a low-capacity model is wrong consistently while a high-capacity one is wrong differently on every sample.

for a senior

Show that you check per-segment error rather than a single aggregate before declaring which end you are on, and that you separate the declared capacity limit from what was actually realised on the data you had.

for a principal

Own the framing that capacity is chosen relative to the data volume and noise level of a specific problem, not inherited as a house default, and be able to explain that tradeoff to people who only see the headline accuracy.

## What capacity actually is A learning algorithm does not consider every conceivable function. It searches a **hypothesis space** - the set of functions the model class can express once its parameters are fitted. Linear regression on one feature can express every straight line and nothing else. A degree-15 polynomial can express every straight line *and* every wiggly curve up to degree 15. A decision tree limited to depth 3 can express at most 8 piecewise-constant regions; a tree allowed depth 30 can carve the input space into far more. **Capacity** is the size or richness of that space. It is the single most useful dial in supervised learning because almost every generalization failure is a statement about it being set wrong. A subtlety worth carrying: capacity is a property of the model *class*, not of the fitted model and not of the data. Fitting a degree-15 polynomial to 60 listings does not make the class smaller; you have merely picked one member of a very large family. This is why capacity is normally discussed before you look at any result. ## The two ends of the dial **Underfitting (the low-capacity end).** The hypothesis space simply does not contain a function close to the true relationship. No amount of optimisation helps, because the best available member is still a poor approximation. The signature is that training error is high *and* test error is high, and the two are close to each other. A depth-3 tree on 5,000 rental listings gives you eight rent predictions for the entire market; if rent genuinely depends on size, district, floor and age, eight buckets cannot express it, and the model is systematically wrong in the same direction for whole groups of listings. **Overfitting (the high-capacity end).** The space is rich enough that the procedure can find a function which follows the particular sample you drew, including the part of that sample that is noise - a landlord who priced oddly, a listing with a typo in the floor area. The signature is a large gap: training error very low, test error much higher. A degree-15 polynomial fitted to 60 listings can be made to pass essentially through every point; between the points it does whatever the algebra demands, which can mean predicting negative rent for a 45-square-metre flat because two nearby listings pulled the curve apart. ## Which dials change capacity - **Polynomial degree** - the classic textbook dial: degree 1 versus degree 15 on the same 60 listings. - **Tree depth and number of leaves** - a depth-30 tree versus a depth-3 stump on the same 5,000 listings. - **The number of input features** - each added feature enlarges the space of expressible functions. - **The number of free parameters in general** - useful as a rough intuition, though it is not a reliable measure; some one-parameter families can express astonishingly many labelings. ## Why more is not better The reason the high end fails is that you fit on a *sample*. Any finite sample carries structure that belongs to the sample rather than the population. A low-capacity model cannot chase that structure because it has nowhere to put it; a high-capacity model can, and optimisation will happily do it, because a training loss cannot tell signal from noise. That is why the two ends fail in qualitatively different ways: the low end is wrong in a stable, repeatable way, and the high end is wrong in a different way on every resampled training set. ## How to tell which end you are on Compare training error with held-out error and look at both the level and the gap: - both high, gap small - low-capacity end; - training low, held-out much higher - high-capacity end; - both low - you are somewhere sensible, assuming the held-out data really is held out. A caution: both errors being high is only *evidence* of underfitting. If the features carry no information about the target, or if the labels are intrinsically noisy, no capacity setting fixes it, and adding capacity in that situation just moves you to the other failure mode. ## Effective capacity What the model class *could* express and what it *does* express on a given dataset are not the same. Declaring a maximum depth of 30 on 5,000 listings does not produce a fully grown depth-30 tree - branches run out of examples long before that. The declared limit is an upper bound on capacity; what you actually realise depends on the data you fitted. People sometimes call this the effective capacity, and it is why the same setting can be reckless on a small dataset and conservative on a large one.

  • Training and held-out error are both high and almost equal. Which end of the dial are you on?
    The low-capacity end - the model class cannot express the pattern, so it is wrong in the same way on data it has seen and data it has not. The one alternative worth ruling out is that the features carry no signal about the target, or the labels are so noisy that no model class does better; in that case adding capacity buys nothing.
  • Does adding input features increase capacity?
    Yes. Each extra feature widens the set of functions the model can express, so the hypothesis space grows even if the number of rows does not. That is why a wide table with few rows overfits easily: capacity has gone up while the evidence available to pin down the fit has not.
  • Can one model underfit and overfit at the same time?
    Yes, in different regions of the input space. A tree can be starved of splits in a dense region of ordinary flats and simultaneously grow single-listing leaves out in the luxury tail. Aggregate training and test numbers hide this, which is why per-segment error is worth looking at before concluding the dial is set wrongly overall.

Think of fitting as drawing a line through pins on a board. A rigid steel ruler makes a straight line no matter where the pins are and misses most of them; a floppy wire threads every pin exactly, including the ones that were pushed in slightly wrong.

saying these in an interview costs you the question

  • Says overfitting means the model is too accurate
  • Equates capacity with the number of parameters
  • Claims more training data cures underfitting
  • Diagnoses overfitting from training error alone
  • Assumes any complex model must overfit

context