Why do kNN and SVM need feature scaling while decision trees do not?
answer
- does the model measure a distance
- or ask one question per feature
- sum of squared differences across columns
- splits depend only on value ordering
- any increasing rescale keeps that ordering
basics
~20 skNN and SVM compare examples by distance, so a feature measured in large units dominates every comparison. A tree splits one feature at a time at a threshold; rescaling preserves the ordering of values, so exactly the same splits remain available.
solid answer
~50 sThe dividing line is whether the model combines features into a single geometric quantity or looks at them one at a time. k-nearest-neighbours sums squared differences across all features, so with fasting glucose in mg/dL, BMI, and annual household income in dollars, the income term dwarfs anything glucose or BMI contributes and the model becomes a nearest-neighbour model on income alone. An RBF-kernel SVM inherits that problem, since its kernel is a function of squared Euclidean distance: on wine-chemistry columns where proline runs in the hundreds and pH sits near 3, proline decides the similarity. The same reasoning covers k-means, PCA, and penalised linear models. A decision tree instead asks `is feature j <= t`, and any strictly increasing rescale keeps the ordering of that column identical, so the candidate splits, the impurity gains and the resulting tree are unchanged. Tree ensembles, boosted ones included, inherit that invariance.
go deeper
Memorise the two lists and one reason each: distance-based and kernel methods need scaling because they add up per-feature differences; trees do not because they test one feature at a time.
Derive the lists rather than reciting them. Walk through a concrete distance calculation showing a large-unit feature swamping the others, and explain why an increasing rescale leaves a tree's candidate splits untouched.
Be the person who says scaling is not free and not universal. Know that an unpenalised linear fit is invariant in its predictions, that boosted trees need nothing, and that the real risk is applying a ritual without checking what it bought.
Frame it as pipeline design: which transforms are model-specific rather than data-specific, where they belong so a model swap does not silently change behaviour, and how the fitted statistics are stored and served alongside the model.
## The one distinction that decides it Ask a single question about the model: **does it combine several features into one number whose value depends on their units?** If yes, it is scale-sensitive. If it examines features one at a time through comparisons, it is scale-free. ## The scale-sensitive families **Distance-based learners.** k-nearest-neighbours ranks candidate neighbours by squared Euclidean distance, a sum of per-feature squared differences. Consider a diabetes-risk model with three features: fasting glucose in mg/dL (roughly 70-200), BMI (roughly 18-45), and annual household income in dollars (roughly 20,000-200,000). A 30,000-dollar income gap contributes a squared term of nine hundred million; a clinically enormous 50 mg/dL glucose gap contributes 2,500. The glucose and BMI terms are numerically invisible. The model still trains, still returns predictions, and is silently a nearest-neighbour model on income. Standardising all three puts their contributions on the same footing. k-means shares the mechanism, since its assignment step is the same distance comparison. **Kernel methods.** An RBF-kernel SVM computes similarity as a decaying function of squared Euclidean distance between examples, so it inherits the distance problem exactly. On wine-chemistry features where proline runs in the hundreds of units and pH sits near 3, the proline differences swamp the pH differences inside the kernel. Worse, a single kernel width has to serve every feature at once — there is no per-feature knob to compensate with — so no hyperparameter search rescues an unscaled fit. Linear SVMs are affected too, because the margin is measured in the same geometry. **Variance-based decompositions.** PCA finds the directions of greatest variance, and variance is measured in the square of whatever unit the column uses. On a 20-channel semiconductor process panel measured in pascals, millivolts and degrees Celsius, the channel with the numerically largest units carries the largest variance and therefore drags the first component toward itself. Rescale the channels and the component ordering changes — the geometry was never intrinsic, it was a property of the measurement units. **Penalised linear models.** A single penalty strength is applied to every coefficient, but a coefficient's magnitude depends on its feature's units: express a length in metres instead of kilometres and its coefficient shrinks by a thousand. So the penalty falls unevenly across features unless they are first put on a common scale. **Gradient-based optimisation.** Even where scaling does not change the *solution*, features on wildly different scales make the loss surface long and narrow in some directions, so a gradient method takes many small zig-zagging steps to converge. This is a speed and stability argument rather than a correctness one. ## Why trees are exempt A decision tree grows by searching, feature by feature, for a threshold `t` and a split of the form `feature j <= t`. Which rows fall to the left child depends only on **how the values of that one feature are ordered**, not on their magnitudes and not on any other feature's units. Standardisation and min-max ranging are strictly increasing maps: they never reorder values. So the set of distinct ways to partition the rows on that column is identical before and after scaling, the class proportions in each candidate child are identical, the impurity gains are identical, and the algorithm picks the same partition. The chosen threshold is simply reported on the new scale — a split at glucose `<= 126` becomes a split at the standardised equivalent of 126. Everything built from axis-aligned trees inherits this: bagged forests, random forests, and gradient-boosted tree algorithms such as XGBoost, LightGBM and CatBoost. A frequent confusion is that gradient boosting must need scaling "because it uses gradients" — but the gradient steps happen in function space, fitting each new tree to the residual signal, while the trees themselves still split on unscaled thresholds. ## The case that catches people out An unpenalised linear or logistic regression fitted to convergence is **not** scale-sensitive in its predictions. Multiply a feature by 1,000 and its fitted coefficient divides by 1,000; the fitted function, and therefore every prediction, is unchanged. Scaling matters there for two secondary reasons: gradient-based fitting converges faster on a well-conditioned problem, and coefficients on a common scale can be compared to each other, which raw-unit coefficients cannot. ## How to answer in an interview Do not recite a list. State the rule — geometry versus per-feature comparisons — then derive the list from it, and be explicit that scaling is not a universal ritual: on a tabular problem fed to a gradient-boosted tree, it buys nothing at all.
- Does an unpenalised linear regression need its features scaled?Not for correctness. Fitted to convergence, multiplying a feature by a constant simply divides its coefficient by the same constant, so the fitted function and every prediction are identical. Scaling still helps in two ways: gradient-based fitting converges faster when features share a scale, and coefficients on a common scale can be compared with each other, which raw-unit coefficients cannot be.
- Would standardising a feature change which threshold a decision tree picks?No — only how it is written. Standardisation is strictly increasing, so it never reorders a column's values. The same partitions of rows are available, the child class proportions are identical, the impurity gains are identical, and the same split wins. The threshold is then expressed in standardised units rather than the original ones.
- Is a gradient-boosted tree ensemble scale-sensitive?No. Candidates often assume it must be, because boosting is described in terms of gradients. But the gradient steps are taken in function space — each new tree is fitted to the current residual signal — while the base learners are still axis-aligned trees splitting on thresholds. Rescaling a feature leaves every split, and therefore the whole ensemble, unchanged.
saying these in an interview costs you the question
- Says every model needs its features scaled before training
- Claims scaling changes which splits a random forest chooses
- Cannot say what specifically breaks in an unscaled kNN model
- Thinks scaling changes an unpenalised linear model's predictions
- Assumes gradient boosting needs scaling because it uses gradients