A model hits 2% training error but 14% validation error - what do you try first?
answer
- look at the gap, not the level
- training error is already low
- which fix costs you fit?
- rows before shrinking the model
- adding capacity is the wrong direction
basics
~20 sA 12-point train-validation gap is high variance, not high bias. Rank the fixes by cost: more training rows, then a stronger penalty or fewer features, then a lower-capacity model. Adding capacity would make the gap worse.
solid answer
~50 sTraining error of 2% says the model has enough capacity to fit the signal, so nothing here is a bias problem; the 12-point jump on held-out data is variance. Before spending money I sanity-check the split - leakage, a duplicated row landing on both sides, or a validation set drawn from a different period can all fake this pattern. Then I rank remedies by cost and lead time: gather or generate more training rows first, because that cuts variance without giving up the fit I already have; next tighten regularization or prune weak features; only then step down to a lower-capacity model, since that trades away some of the 2% I paid for. Anything that *adds* capacity - a deeper tree, more features, a looser penalty - is the wrong direction and should be off the list.
go deeper
Recall which number signals which problem: a low training error with a much higher validation error means overfitting. Be ready to name two or three fixes without being asked to justify their order.
Explain why each remedy attacks variance and what it costs in fit, and why more representative data is the only one that costs no fit at all. Interviewers expect a ranked list, not a bag of options.
Show that you check the split before you spend: leakage, time-based mismatch, validation-set size and tuning pressure all fake this pattern. Then give a plan with expected effect, cost and a stopping rule.
Own the framing that this is a budget allocation, not a modelling puzzle. Present the options with lead times attached so the organisation can choose between buying data, buying compute and buying engineering time.
## Reading the pair of numbers Two error numbers are given: **training error** (how well the model reproduces the data it was fitted on) and **validation error** (how well it does on held-out data it never saw during fitting). The *level* of the training error tells you about bias; the *gap* between the two tells you about variance. - Training error 2% - the model is expressive enough to represent the pattern in the data. Whatever is wrong, it is not that the hypothesis class is too small. - Gap of 12 points - the fit does not transfer. The model has memorised idiosyncrasies of the particular sample it was trained on. That is the textbook signature of **high variance** (overfitting): re-draw the training set and you would get a noticeably different model. The common junior mistake is to look at the 14% and say "error is high, so the model underfits". Underfitting shows up as a *high training error*, not a high validation error alone. ## Step zero: is the gap real? Before you prescribe anything, confirm the diagnosis is not an artefact: - **Leakage in reverse** - if near-duplicate rows exist and the split was random, the training side gets an unfair advantage on data that is effectively in both sets; this usually *shrinks* the apparent gap, so its absence is worth checking too. - **Split mismatch** - a validation set drawn from a later time window, a different region, or a different customer mix is measuring distribution shift, not variance. The remedy for shift is a representative split and reweighting, not regularization. - **A tiny validation set** - with a few hundred rows, a 14% estimate carries several points of sampling noise. Widen the estimate with cross-validation before acting on it. - **Tuning pressure** - if you have already selected dozens of configurations against this same validation set, the number is optimistic for the winner and the true gap may be even larger. Hold back a test set you touch once. ## The ranked plan With the diagnosis confirmed, order the remedies by how much variance they buy per unit of cost and lead time, and by what they cost you in bias: 1. **More training rows.** The unique remedy that reduces variance *without* reducing the model's ability to represent the signal. Every other option on this list buys variance reduction by giving up some fit. Sources are not only "buy labels": pooling adjacent segments, lengthening the history window, and relabelling rows you already own are cheaper than a new collection contract. 2. **Constrain the model without changing its family.** Strengthen the penalty on the weights, cut low-signal or highly correlated features, cap tree depth or leaf size, stop the fit earlier. These are hours of work, not weeks, and each is reversible if it costs you too much training accuracy. 3. **Average many fits.** Training several models on resampled data and averaging their predictions cuts variance; it costs compute and interpretability rather than accuracy. 4. **Step down to a simpler model family.** Cheapest to run, but the bluntest: you are deliberately raising bias to lower variance, and with 2% training error you have some room to spend - but not unlimited room, since the floor is wherever irreducible noise sits. The reason more data leads the list is asymmetry of risk. If you regularize too hard you push training error up and may land in a *worse* place than you started; if you add representative data you may see no improvement, but you rarely go backwards. ## What to rule out Everything that increases capacity is disqualified by the diagnosis: a deeper tree, a richer feature set, a looser penalty, longer optimization. Each buys a lower training error you do not need and widens the gap you are trying to close. If a stakeholder proposes "a bigger model", the answer is that the model already fits the training data at 2% - the problem is that the fit does not generalise. ## Knowing when to stop A remedy plan needs a stopping rule as much as it needs a first step. Decide up front what validation error would make the model shippable and what the achievable floor plausibly is. If two rounds of the plan move validation error from 14% to 9% and the target is 8%, that is a different conversation from moving it to 13.5%. Report the plan as *ranked options with expected effect and cost*, not as a single action, so whoever holds the budget can choose between buying data and buying engineering time.
- Under what conditions would more training rows fail to close that gap?When the extra rows are not representative of the validation distribution, when they are near-duplicates of rows you already have, or when the labels themselves are noisy - then you are adding variance in the target rather than information. It also fails once you are near the model's own asymptote, where the curve flattens and further rows buy fractions of a point.
- How would your ranking change if training error were 2% and validation error were 4%?A two-point gap is a mild variance signal, and the whole plan changes: it is no longer worth buying data. I would ask instead whether 4% is close to the achievable floor, and if it is, stop tuning and spend the effort on deployment, monitoring, or the slices where errors are most costly.
- Your only cheap option is to shrink the model. How much training error are you willing to give up?As much as it takes to keep validation error falling, and no more. I sweep the constraint and watch both curves: while validation error drops I accept the rising training error, and I stop at the point where validation error flattens or turns up. The training number is not a target - it is the price tag.
A student who aces every past paper but fails the real exam has not learned too little - they have memorised the answers. You give them a wider set of problems before you tell them to study a simpler syllabus.
saying these in an interview costs you the question
- Calls high validation error underfitting without looking at the gap
- Proposes a bigger or deeper model to close a variance gap
- Adds features when the training error is already near zero
- Never checks the split for leakage or distribution mismatch
- Treats more data as free and always the answer
- Keeps tuning against the same validation set and quotes the winning score