Why can min-max scaling still leave one feature dominating a distance-based model?
answer
- everything sits in zero to one
- but distance eats differences, not bounds
- a heavy tail sets the divisor
- the bulk collapses near zero
- range equalised, spread not
basics
~20 sMin-max equalises each feature's range, not its spread. A heavy-tailed column whose maximum sits far above the bulk gets compressed near zero after scaling, so a well-spread bounded column ends up supplying almost all of the distance.
solid answer
~50 sTake a clustering run over two columns: a 0-to-10 satisfaction score and unbounded monthly page-views. Min-max puts both in `[0, 1]`, which feels like fairness, but the mapping is fixed by the two extremes. Satisfaction uses the whole interval, so its scaled values have a spread of maybe 0.3. Page-views, where most users sit in the hundreds and a few sit near half a million, collapse into the first thousandth of the interval, so their scaled spread is nearer 0.001. Squared differences between rows are therefore hundreds of times larger on satisfaction, and the clustering is effectively driven by satisfaction alone — the reverse of what the analyst intended. Fix it by equalising the quantity distance actually uses: standardise, or scale by the interquartile range if the tail also distorts the standard deviation. Then check the scaled columns' spreads are comparable.
go deeper
Take away one fact: putting two columns in the same 0-to-1 range does not make them count equally in a distance. Range and typical spread are different things.
Work the arithmetic out loud. Show what a heavy-tailed column's typical values become after being divided by an extreme maximum, and compare that with a bounded column that fills its interval.
Demonstrate the diagnosis. Explain how you would notice this in a real pipeline that raises no error, which statistic you compare after scaling, and why standardisation is the fix that targets the quantity distance actually uses.
Own the review standard. Silent preprocessing failures pass code review because the step is present and looks correct, so decide what evidence a pipeline must produce about its transforms before a model built on them ships.
## The assumption that fails "Everything is between 0 and 1, so no feature can dominate" sounds airtight, and it is wrong. Distance does not consume ranges; it consumes differences between rows. Two columns can share the interval `[0, 1]` and still produce differences that live on completely different scales. ## Working the example A clustering run has two features: - **Satisfaction score**, an integer from 0 to 10, with users spread fairly evenly across it. - **Monthly page-views**, unbounded, where the median user is in the low hundreds and a handful of power users reach 500,000. Min-max ranging maps each to `[0, 1]` using the training minimum and maximum. Satisfaction is a bounded column whose sample minimum and maximum are the true bounds, so scaled values land at 0, 0.1, 0.2, ..., 1.0. Two randomly chosen users typically differ by something like 0.3 on this column. Page-views are divided by roughly 500,000. The median user, at say 300 views, lands at 0.0006. A heavy user at 3,000 lands at 0.006. Two randomly chosen users typically differ by a few thousandths. Only the tiny handful of power users occupies the rest of the interval. In a squared-Euclidean sum, a typical satisfaction difference contributes about 0.09, and a typical page-views difference contributes about 0.00001. The page-views column, for every ordinary pair of rows, has effectively been switched off. The clusters that come out are satisfaction bands wearing a two-feature costume. ## The general statement Min-max equalises **range**: the distance from the smallest to the largest value. Distance-based learners are driven by **spread**: how far typical rows sit from one another. Those two quantities agree only when the columns have similar distribution shapes. When one column is bounded and evenly filled and the other is heavy-tailed, the same range hides wildly different spreads, and the evenly-filled column wins. The deeper point is that min-max's mapping is determined by two order statistics, the extremes. In a heavy-tailed column those extremes describe almost none of the data, so a transform anchored to them tells you almost nothing about where the data actually lives. ## What to do instead **Standardise.** Dividing by the standard deviation targets exactly the quantity distance uses, so every column contributes comparable typical differences by construction. This is the default fix. **Use median/IQR scaling when the tail also distorts the standard deviation.** If page-views are extreme enough that the standard deviation is itself set by the power users, scaling by the interquartile range spreads the ordinary rows out properly and still leaves the power users far away. **Reshape the column, if the tail is the real problem.** A heavy tail is a distribution-shape problem, and no linear rescale can fix a shape. A nonlinear transform is the separate tool for that, chosen on its own merits. **Reserve min-max for columns that are genuinely bounded.** When the bounds come from the domain — a 0-to-10 score, a percentage, a fixed-ceiling rating — min-max is well behaved: the mapping is a constant of the problem, not an accident of the sample, and no production value can fall outside it. It is the unbounded, heavy-tailed column where min-max misleads. ## The check that catches it After scaling, compute the standard deviation (or the interquartile range) of each scaled column and compare. If the numbers differ by an order of magnitude or more, the columns are not contributing comparably, whatever their ranges say. This costs one line and turns an invisible modelling failure into a visible one. A second habit is to sanity-check the result rather than the transform: cluster or fit, then look at whether the structure you found varies along more than one feature. If every cluster is defined by one column, the scaling step did not do the job you thought it did. ## Why this is worth knowing This failure is silent. Nothing errors, nothing looks wrong in a summary table, and every column dutifully reports a minimum of 0 and a maximum of 1. Someone reviewing the pipeline sees a scaling step and ticks it off. The only way to catch it is to know that equal ranges are not equal influence, and to check the quantity the model actually consumes.
- How would you check quickly that a scaling step actually equalised feature influence?Compute the standard deviation, or the interquartile range, of each scaled column and compare them. Equal ranges with spreads that differ by an order of magnitude mean unequal contributions to any Euclidean distance. It is one line of arithmetic and it converts a silent modelling failure into something visible before the model is ever fitted.
- Does z-score standardisation guarantee that every feature contributes equally?It equalises variance, which is the right quantity, but it is not a guarantee of equal influence. A heavy tail still means a few rows dominate specific pairwise distances. Two near-duplicate columns still give that one underlying concept double weight. And equal statistical influence is not the same as equal predictive usefulness — that is a modelling judgement, not a scaling one.
- When is min-max ranging the right choice for a feature?When the bounds come from the domain rather than from the sample: a 0-to-10 score, a percentage, a rating with a fixed ceiling. Then the minimum and maximum are constants of the problem, no production value can fall outside the interval, and the scaled column fills its range properly. It is also right when a downstream step genuinely requires inputs in a fixed range.
Two runners are told to race the same 100-metre track, so the contest looks fair. But one is confined to the first ten centimetres and the other uses the whole distance. Same track length, wildly different amounts of movement.
saying these in an interview costs you the question
- Assumes a shared 0-to-1 range means equal influence
- Reaches for min-max by default on unbounded columns
- Treats the sample maximum as if it were a domain bound
- Never inspects the spread of the scaled columns
- Says scaling can fix a heavy-tailed distribution's shape