skip to content

Scaling and Standardization

Z-score standardisation, min-max ranging and median/IQR robust scaling, and why kNN, SVMs, k-means and gradient descent care while trees do not. Interviewers use it to test scale-sensitivity.

on this pageshow

questions

4

What is the difference between z-score standardisation and min-max scaling?

level: juniorimportance: must knowfreq 84%

answer

  1. two ways to make columns comparable
  2. one uses centre and spread
  3. the other uses the two endpoints
  4. unbounded output versus bounded output
  5. both are straight-line rescales

basics

~20 s

Z-score standardisation subtracts the mean and divides by the standard deviation, producing mean 0, standard deviation 1, and no fixed bounds. Min-max scaling rescales values into a fixed range such as 0 to 1 using the training minimum and maximum.

solid answer

~50 s

Both put numeric columns on a comparable scale, but they use different statistics and give different guarantees. Z-score standardisation computes `z = (x - mean) / sd`, so the column ends up centred at 0 with a standard deviation of 1 and no upper or lower bound; it is driven by the whole distribution. Min-max computes `x' = (x - min) / (max - min)`, so training values land in `[0, 1]`, but the mapping is decided by just two extreme order statistics, and unseen values can fall outside the range. Max-abs scaling, a close relative, only divides, so exact zeros stay zero and a sparse column stays sparse. All three are straight-line transforms: they change units, not shape, so skew and outliers survive intact. I default to standardisation, and reach for min-max when the column is genuinely bounded or a downstream step needs bounded inputs.

go deeper

for a junior

Be ready to write both formulas from memory and say what the column looks like afterwards: mean 0 and standard deviation 1 for one, values between 0 and 1 for the other.

for a middle

Explain which statistics each transform depends on, and why depending on two extreme order statistics makes min-max fragile while depending on the whole distribution makes standardisation the safer default.

for a senior

Show that you think about new data: what happens when a production value falls outside the training range, when a bound is a domain constant rather than a sample extreme, and when preserving zeros matters more than centring.

for a principal

Own the convention. Pick one default transform per column type across the team, write down when the exception applies, and make sure the fitted statistics are versioned with the model rather than recomputed wherever the column happens to be used.

## Why rescale at all A table of numeric features usually arrives in whatever units the source system happened to use: a glucose reading in mg/dL, a body-mass index around 25, a household income in dollars. Several model families compare or combine feature values directly, and when they do, the column with the biggest numbers wins by accident of measurement rather than by relevance. Scaling is the family of transforms that removes that accident. ## Z-score standardisation The formula is ``` z = (x - mean) / sd ``` where `mean` and `sd` are the column's mean and standard deviation, computed once on the training data and then reused unchanged for validation, test and production rows. After the transform the column has mean 0 and standard deviation 1. The output is unbounded: a value five standard deviations above the mean becomes 5.0, and nothing clips it. Both statistics are computed from every row, so every row has some influence on the result. That is usually what you want, and it is also the weakness: the mean and especially the standard deviation are pulled by extreme values, so one enormous observation can inflate `sd` and squeeze the rest of the column into a narrow band around zero. ## Min-max ranging The formula is ``` x' = (x - min) / (max - min) ``` using the training minimum and maximum. Training values are mapped onto `[0, 1]`; the minimum becomes exactly 0 and the maximum exactly 1. Any target interval works — subtracting and rescaling again gives `[-1, 1]` — but the character is the same: the mapping is determined entirely by two order statistics, the two most extreme rows in the column. Every other row is only positioned relative to them. Two consequences follow. First, a single freak maximum decides the whole mapping, and every ordinary value gets pushed toward the bottom of the interval. Second, the bounded-output guarantee only holds for data inside the training range. A new row above the training maximum maps above 1, and a new row below the training minimum maps below 0. Nothing breaks numerically, but if a downstream component genuinely requires inputs in `[0, 1]`, you have to decide whether to clip — and clipping throws away the distinction between merely large and enormous. ## Max-abs scaling A third linear option divides by the largest absolute value in the column: ``` x' = x / max(|x|) ``` Values land in `[-1, 1]`, and because nothing is subtracted, an exact zero maps to an exact zero. That matters for wide, mostly-zero count columns: subtracting a mean would turn every stored zero into a non-zero number and destroy the sparsity that made the representation cheap in the first place. ## What none of them do All three are affine transforms — multiply by a constant, add a constant. Affine transforms move and stretch the number line; they do not bend it. So: - Standardising does **not** make a column normally distributed. A right-skewed column standardised to mean 0 and sd 1 is exactly as skewed afterwards. - Scaling does not remove outliers. An extreme row is still extreme, just expressed in new units. - The Pearson correlation between two columns is unchanged by rescaling either of them (with a positive multiplier), and so is the ordering of values within a column. Changing the *shape* of a distribution requires a nonlinear transform, which is a different tool for a different problem. ## Choosing between them Standardisation is the sensible default for the scale-sensitive families — distance-based learners, kernel methods, variance-based decompositions and penalised linear models — because those methods care about spread, and standardisation is what equalises spread. Min-max earns its place when the column is genuinely bounded and you know the bounds from the domain rather than from the sample: a 0-to-10 satisfaction score, a percentage, a rating with a fixed ceiling. Then `min` and `max` are constants of the problem, not sample accidents, and out-of-range values cannot appear. It is also the right answer when a downstream component requires bounded inputs. Max-abs is the sparse-data choice, for the zero-preservation reason above. ## A note on vocabulary "Normalisation" is used loosely in practice: some people mean min-max ranging, some mean z-score standardisation, some mean scaling a row vector to unit length. In an interview, say which formula you mean rather than relying on the word — and if the interviewer uses it, ask.

  • Does standardising a column make it normally distributed?
    No. Subtracting the mean and dividing by the standard deviation is a straight-line rescale, so it moves and stretches the distribution without bending it. A right-skewed column is exactly as skewed after standardisation as before, and its extreme values are still extreme. Changing the shape needs a nonlinear transform, which is a separate decision from putting columns on a comparable scale.
  • What happens when a production value exceeds the training maximum used by min-max scaling?
    It maps above 1. Arithmetic still works, but the bounded-range guarantee is gone, so any downstream step that assumed inputs in `[0, 1]` is now out of contract. You either accept out-of-range values, or clip them — and clipping means every value above the training maximum becomes indistinguishable from every other. That is a reason to prefer domain bounds over sample extremes when using min-max.
  • Why is max-abs scaling preferred for a wide, mostly-zero count column?
    Because it only divides, never subtracts. Dividing by the largest absolute value leaves exact zeros as exact zeros, so a sparse column stays sparse. Subtracting a mean would give every stored zero a non-zero value, turning a cheap sparse representation into a dense one and, on a wide feature set, blowing up memory for no modelling benefit.

Standardisation is like reporting a student's score as how many standard deviations from the class average it sits. Min-max is like reporting it as a percentage of the range between the worst and best score in that particular class.

saying these in an interview costs you the question

  • Says standardisation makes a column normally distributed
  • Claims min-max is always safer because outputs are bounded
  • Recomputes the minimum and maximum separately on new data
  • Thinks rescaling changes the correlation between two columns
  • Uses normalisation and standardisation as if they were one thing

context

open as a page

Why do kNN and SVM need feature scaling while decision trees do not?

level: middleimportance: must knowfreq 76%

basics

~20 s

kNN and SVM compare examples by distance, so a feature measured in large units dominates every comparison. A tree splits one feature at a time at a threshold; rescaling preserves the ordering of values, so exactly the same splits remain available.

open as a page

When would you use median/IQR robust scaling instead of z-score standardisation?

level: middleimportance: should knowfreq 46%

basics

~20 s

Use it when a column carries extreme values. Robust scaling subtracts the median and divides by the interquartile range, statistics that a handful of extreme points barely move, so the bulk of the data still lands on a usable scale.

open as a page

Why can min-max scaling still leave one feature dominating a distance-based model?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

Min-max equalises each feature's range, not its spread. A heavy-tailed column whose maximum sits far above the bulk gets compressed near zero after scaling, so a well-spread bounded column ends up supplying almost all of the distance.

open as a page