A k-NN validation curve over k = 1 to 200 shows zero training error at k = 1. Which end is high capacity?
answer
- the number goes up, complexity goes down
- each point is its own neighbour
- the jagged boundary lives at small k
- zero training error is free, not good
basics
~20 sThe small-k end. At k = 1 every training point is its own nearest neighbour, so training error is zero by construction. Capacity falls as k grows, so this curve's complexity axis runs right to left.
solid answer
~50 sHigh capacity sits at k = 1 and capacity decreases as k rises, so the overfitting region of this curve is on the left, not the right. The zero training error at k = 1 is an artefact rather than a result: the nearest neighbour of any training row is that row itself, so it always predicts its own label — assuming no two identical feature vectors carry conflicting labels. On handwritten-digit pixels that means k = 1 memorises the training set perfectly while the validation branch is what actually moves. The general lesson is that a hyperparameter's numeric direction does not tell you its capacity direction: depth and polynomial degree add capacity as they grow, while the neighbour count, the minimum samples required in a leaf and the minimum samples required to split all remove it. Before reading any validation curve, decide which end is the flexible model.
go deeper
Remember that at k = 1 a point is its own nearest neighbour, so a perfect training score there is automatic. Be able to say that more neighbours means a smoother, simpler model.
Explain the mechanism both ways: why the boundary is jagged at small k and rigid at large k, and how that flips which end of the plot is the overfitting region. Expect to be asked for other dials that run backwards.
Demonstrate the habit of labelling the capacity direction before interpreting any sweep, and of treating a perfect training score as an artefact to explain rather than a result to report.
Own the reporting convention: sweeps in a team's model reports should be plotted so that capacity always increases the same way, otherwise reviewers misread each other's curves and argue about the wrong end.
## Why k = 1 gives exactly zero training error A k-nearest-neighbour classifier stores the training rows and predicts a new point by looking up its k closest stored rows. When you score it on the very rows it stored, the closest row to each training point is that point itself, at distance zero. With k = 1 the prediction is therefore that point's own label, for every training point. Training error is zero — not because the model is good, but because the scoring set and the memory are the same set. The one caveat worth stating precisely: this holds as long as no two identical feature vectors carry different labels. If the same pixel pattern appears twice in the digit data with two different digit labels, one of them must be wrong at k = 1, and training error is a small positive number instead of exactly zero. ## Capacity runs backwards on this axis At k = 1 the decision boundary is maximally jagged: it wraps individually around every training point, and moving one point moves the boundary. That is the definition of a high-variance, high-capacity fit. As k grows, each prediction is decided by a larger and larger neighbourhood, the boundary smooths, and the fit becomes more rigid. At k = 200 on a modest dataset the model is close to predicting the majority class nearly everywhere — maximum bias, minimum variance. So the plot you get looks like a mirror image of a tree-depth curve. Training error rises with k instead of falling; the train-validation gap is widest on the left; validation error is U-shaped with its minimum somewhere in the middle, and the overfitting region is at small k. ## Why this trips people up Candidates internalise "bigger hyperparameter, more complex model" from depth and polynomial degree, then read every curve left to right as underfit-to-overfit. Applied to a neighbour sweep, that reading is exactly inverted, and the same inversion applies to several common dials: - **Adds capacity as the number grows:** maximum tree depth, polynomial degree, number of leaf nodes allowed, number of hidden units. - **Removes capacity as the number grows:** number of neighbours, minimum samples required in a leaf, minimum samples required to split a node, pruning strength. The safe habit is to name the two ends before interpreting: "left is flexible, right is rigid" or the reverse, written on the plot. Some practitioners plot such dials reversed, or plot an explicit capacity proxy on the x-axis, precisely so every validation curve in a report reads the same way. ## What zero training error is worth Nothing, on its own. A perfect training score is available for free from any model with enough capacity to memorise: a fully grown tree with one row per leaf reaches it too. It carries information only in contrast with the validation branch — the size of that gap is what tells you how much of the fit is memorisation. A candidate who reports "my model gets 100% on training" as a positive result is telling the interviewer they do not distinguish fitting from generalising. ## Reading the neighbour sweep in practice On a digit-pixel sweep the useful part of the curve is the validation branch between roughly k = 1 and the point where it clearly turns upward. Small k values are worth including precisely because they anchor the high-capacity end and make the gap visible; very large k values are worth including because they show you the floor the model degrades to. If validation is still improving at the largest k you swept, the sweep stopped too early and the grid needs extending — the same rule as for any other dial. One boundary worth respecting: the curve tells you which k generalises best on this data. How the k neighbours are combined into a prediction — the voting scheme itself — is a property of the algorithm rather than of the curve, and it is a separate discussion from reading the sweep.
- Name two other hyperparameters whose capacity direction is the reverse of their numeric direction.The minimum number of samples required in a leaf and the minimum number required to split a node: as both grow, the tree is forced to stop earlier and the fitted function gets smoother, so larger values mean a simpler model. Pruning strength behaves the same way. Depth, polynomial degree and the number of leaves allowed all run the other direction, adding capacity as they grow.
- When would training error at k = 1 not be exactly zero?When the data contains identical feature vectors with conflicting labels. The nearest neighbour of such a point at distance zero may be its contradictory twin, so at least one of the pair is predicted wrongly no matter what. It is a useful smell test: a non-zero training error at k = 1 usually means duplicated or contradictory rows worth investigating before tuning anything.
- Does a perfect training score anywhere on a validation curve tell you anything useful?Only in contrast with the validation line. Any model with enough capacity to memorise reaches it — a fully grown tree does too — so on its own it measures memory, not skill. What carries information is the gap: how far validation sits below a perfect training score is a direct read on how much of the fit is noise.
saying these in an interview costs you the question
- Assumes a larger hyperparameter value always means more capacity
- Reports zero training error at k = 1 as a good result
- Reads the overfitting region as the right-hand end of every curve
- Says a larger neighbour count always generalises better