How does training-set size change the model capacity you should choose?
answer
- instability shrinks, systematic error does not
- the turning point slides rightward
- how many listings land in each leaf
- a setting tuned against a sample size
basics
~10 sMore training data lets you afford more capacity. Sensitivity to the sample shrinks as the sample grows while a class's systematic error does not, so the best capacity moves upward with dataset size.
solid answer
~40 sThe capacity that minimises held-out error depends on how much data you have. A larger sample reduces how much the fit moves when you redraw the training set, and that instability is the force pushing held-out error up at high capacity; the systematic error of the class is unaffected by sample size. So the minimum of the curve slides toward higher capacity and the whole curve drops. A tree grown to depth 30 on 5,000 rental listings is memorising - its deep leaves hold one or two listings each - while the same depth on millions of listings can leave thousands of listings per leaf and be well supported. So a depth or degree setting is not transferable between projects or data refreshes: it was tuned against a sample size.
go deeper
Remember the direction: more training data supports a more flexible model, and small datasets favour simple ones. Being able to state that without the mechanism is enough here.
Explain why only the sample-sensitivity part of the error shrinks with more rows while the systematic part does not, and use that to say why the turning point of the held-out curve moves.
Show that you distinguish a declared capacity limit from what the data actually realises, and that you revisit settings when a dataset grows or a market changes rather than carrying them forward untested.
Own the tradeoff when someone proposes buying or labelling more data: be able to say what it will and will not buy, and set the expectation that capacity settings are dataset-scoped decisions with an owner and a review point.
## The question behind the question Interviewers ask this to find out whether a candidate treats capacity settings as fixed properties of an algorithm or as choices made relative to the data at hand. The second view is the correct one, and it changes how you run projects. ## Why the optimum moves with n Held-out error at a given capacity is pushed up by two things: how wrong the class is systematically, and how much the fitted function moves when you redraw the training sample. - **The systematic part does not depend on n.** The best straight line for the rent-versus-size relationship is what it is; collecting more listings does not make straight lines able to bend. If the class cannot express the pattern, no sample size fixes that. - **The instability part shrinks as n grows.** Each fitted parameter, split or leaf is estimated from data, and more data pins each of them down more tightly. Roughly, the fit's sensitivity to the sample falls as the sample grows. Since only the second term shrinks, and it is the term that penalises high capacity, the balance point moves right. With more data the U-shaped curve both drops and finds its minimum at a larger hypothesis space. This is the formal version of the practitioner's rule that big models want big data. ## The concrete picture Take tree depth on rental listings. At 5,000 listings, a tree allowed depth 30 runs out of data long before depth 30: branches split down to leaves holding one or two listings, and the prediction in each leaf is essentially one landlord's asking price. Every resampled training set produces a visibly different tree, and held-out error is poor. Now imagine several million listings from the same market. The same depth-30 limit now leaves leaves that still contain many listings each, so the value in each leaf is an average over real evidence rather than an echo of one row. The tree is expressing genuinely finer structure - a specific district, a specific size band, a specific floor - that a depth-3 stump would flatten away. Same capacity setting, completely different behaviour, because the amount of evidence per region of the input space changed. The same story runs on the polynomial dial. A degree-15 curve fitted to 60 listings is one coefficient away from interpolating them; the same degree fitted to 60,000 listings is a mild, well-determined curve. ## Declared capacity versus realised capacity This is where senior candidates separate themselves. A maximum depth, a maximum number of leaves or a polynomial degree is an *upper bound* on the hypothesis space. What the fit actually realises depends on how much data reaches each part of the model. That means the effect of a capacity setting is not linear in the setting - raising a depth limit past the point where branches are already starved changes nothing, while the same increase on a much larger dataset changes a great deal. ## Other things that move the optimum Sample size is the biggest lever but not the only one: - **Label noise.** Noisier targets push the best capacity *down*: extra flexibility spends itself reproducing noise, so the turning point comes earlier. A market where identical flats are listed at wildly different prices supports less capacity than one where pricing is disciplined, at the same sample size. - **Signal strength and feature quality.** If the features genuinely determine the target, more capacity has more real structure to find. If they barely relate to it, capacity buys instability only. - **How the data is distributed across the input space.** A million listings that are all one-bedroom flats in one district do not support a fine-grained model of the whole market. What matters is data *per region the model wants to distinguish*, not the headline row count. ## What to do with this in practice Three habits follow directly. First, do not port capacity settings between projects, and treat a setting inherited from a tutorial or an older version of the dataset as unvalidated. It was tuned against a different sample size, a different noise level and a different feature set. Second, when the dataset grows materially - a backfill, a new market, a year of extra history - the previously chosen capacity is now probably too low, not merely still valid. Growth is a reason to revisit the dial upward. Third, when someone proposes collecting more data, be clear about what it will buy. If the current model is unstable - a big gap between error on fitted data and error on unseen data - more data helps directly and also unlocks more capacity. If both errors are high and close together, the class cannot express the pattern, and more rows of the same kind will not move either number; the constraint is the hypothesis space, not the sample.
- Does collecting more data reduce the systematic error of a model class?Not directly. Systematic error is a property of what the class can express, so more listings do not make straight lines bend. The indirect effect is the real one: more data makes a richer class safe to use, and the richer class has less systematic error. So data buys you the option of lower bias rather than lower bias itself.
- Held-out error does not improve at all when you double the training data. What does that tell you?That you are at the low-capacity end - the error is dominated by what the class cannot express, and extra rows do not change that. A model whose error is driven by instability would visibly improve with more data. It can also mean the new rows duplicate the old ones or come from a different regime.
- How does label noise change the capacity you should pick?Noisier labels push the best capacity down. Extra flexibility gets spent reproducing noise, so the turning point of the held-out curve arrives at a smaller hypothesis space. Two datasets of identical size can therefore support very different depths, which is another reason the setting is not transferable between problems.
Capacity is like the number of price bands a market report divides a city into. With 200 sales you can only justify a few bands; with two million you can report street by street, and the same street-level detail that was noise before is now measurement.
saying these in an interview costs you the question
- Treats a depth or degree setting as a fixed algorithm property
- Says more data always improves any model
- Claims more data reduces bias directly
- Ignores label noise when choosing capacity
- Counts total rows, not evidence per region