skip to content

When would you use median/IQR robust scaling instead of z-score standardisation?

level: middleimportance: should knowfreq 46%

answer

  1. the problem is a stretched divisor
  2. extreme rows inflate the standard deviation
  3. swap in order statistics
  4. median at zero, quartiles one apart
  5. beware a column with zero IQR

basics

~20 s

Use it when a column carries extreme values. Robust scaling subtracts the median and divides by the interquartile range, statistics that a handful of extreme points barely move, so the bulk of the data still lands on a usable scale.

solid answer

~50 s

Standardisation divides by the standard deviation, which squares deviations and is therefore inflated by extreme values. One row a thousand times larger than the rest can push it so high that every ordinary row is squeezed into a band near zero, and a distance-based model then sees almost no difference between typical examples. Robust scaling replaces the two statistics with order statistics: `x' = (x - median) / IQR`, where `IQR = Q3 - Q1`. Neither statistic moves much when a few points sit far out, so the middle half of the data lands in a band of width exactly 1 and typical rows stay distinguishable. It does not bound the output and it does not remove or cap the extreme rows, which simply keep large values in IQR units. Watch for a column where over half the values are identical: its IQR is zero and the division is undefined.

go deeper

for a junior

Know the formula and the two statistics it uses: subtract the median, divide by the interquartile range, which is the 75th percentile minus the 25th.

for a middle

Explain the mechanism. Say why squaring deviations lets one extreme row inflate the standard deviation, and why quartiles are unmoved by the same row, then show what that does to the scaled bulk of the data.

for a senior

Bring the operational edges: the zero-IQR column that produces infinities, the diagnostic of comparing standard deviation against interquartile range before choosing, and the fact that scaling is not outlier treatment.

for a principal

Decide the policy. Which columns get which transform, whether the choice is made per column by a documented rule or fixed for the whole table, and who owns the fallback when a column's spread statistic degenerates.

## The failure it fixes Standardisation divides by the standard deviation, which is computed from squared deviations around the mean. Squaring means the largest deviations dominate the sum, so one or two extreme rows can raise the standard deviation far above what the typical spread of the column looks like. When that happens, the standardised column has a peculiar shape: the vast majority of rows sit in a tiny interval such as -0.05 to 0.05, while a small number sit at 20 or 40. For a model that cares about differences between rows, that is close to having thrown the feature away. Two typical examples now differ by hundredths of a unit on this column while differing by whole units on every well-behaved column, so the feature stops contributing to the comparison even though it is perfectly informative for ordinary rows. ## The transform Robust scaling swaps both statistics for order statistics: ``` x' = (x - median) / IQR where IQR = Q3 - Q1 ``` `Q1` and `Q3` are the 25th and 75th percentiles. Substituting the two quartiles into the formula shows the geometry directly: `Q1` maps to `(Q1 - median) / IQR` and `Q3` to `(Q3 - median) / IQR`, and the two differ by exactly 1. So the middle half of the data always occupies a unit-width interval, and the median sits at 0. The key property is breakdown behaviour. The median and the quartiles are determined by the *position* of values, not their magnitude: you can push the largest 10% of the column to arbitrarily large numbers and the median will not move at all, and the quartiles barely. The mean and standard deviation have no such protection — a single value can drag either as far as you like. ## What it deliberately does not do **It does not bound the output.** Robust scaling is not min-max with a different centre. A row far above `Q3` stays far above; after scaling it might sit at 45. The scaled column is unbounded in both directions, exactly like a standardised one. **It does not treat the extremes.** Nothing is dropped, clipped or winsorised. The extreme rows are still present with their original relative position; the transform only chose a centre and a unit that ignore them. Whether those rows should be capped, removed, or kept as genuine signal is a separate decision made on the data's meaning, not something a scaler decides for you. **It does not change shape.** Like every other affine rescale, it moves and stretches without bending. A skewed column stays skewed. ## When it is the wrong choice **A zero interquartile range.** If more than half of a column's values are identical — very common for count columns that are mostly zero, or for a flag-like numeric field — then `Q1 = Q3`, `IQR = 0`, and the division is undefined. Any implementation of the idea needs a rule for that case: leave the column unscaled, drop it as near-constant, or use a wider pair of percentiles such as the 5th and 95th. Discovering this in production because a batch of scaled values came back as infinities is a bad way to learn it. **A well-behaved column.** If the column has no heavy tail, the standard deviation and the IQR carry the same information and standardisation is the more conventional, more widely understood choice. Robust scaling is not strictly better; it is a targeted answer to a specific problem. **When the extremes are the signal.** In fraud or fault detection the enormous values are often the cases you most want the model to separate. Rescaling so that the ordinary rows spread out is usually still fine — the extremes remain far away — but be sure you are not reasoning as though robust scaling somehow protects the model from them. ## Diagnosing which you need Before choosing, compare the two summaries of the same column: mean against median, and standard deviation against IQR. When the mean sits far from the median, or the standard deviation is several times the IQR, extreme values are steering the mean-based statistics and robust scaling is the safer transform. After scaling, look at the distribution of the output — if nearly every row landed inside a hair's width of zero, the divisor was set by the tail, and that is the symptom you were trying to avoid.

  • What breaks if a column's interquartile range is zero?
    The divisor is zero and the transform is undefined. It happens whenever at least half the values are identical — a mostly-zero count column is the classic case. You need an explicit rule: leave the column unscaled, drop it as effectively constant, or widen the percentile pair to something like the 5th and 95th. Without a rule you get infinities in the scaled column.
  • Does robust scaling remove outliers from the column?
    No. It chooses a centre and a unit that are insensitive to extreme values, but every row keeps its relative position — an extreme row simply becomes a large number of IQR units from the median. Whether those rows should be capped, dropped or kept is a separate judgement about what they mean in the data, and the scaler makes none of it for you.
  • How can you tell from the data that z-score standardisation went badly on a column?
    Look at the scaled output. If nearly every row sits within a hundredth of zero while a handful sit at 30 or 40, the standard deviation was set by the tail rather than by the bulk, and the feature has effectively stopped contributing to any distance-based comparison. The raw-column tell is a standard deviation several times larger than the interquartile range.

saying these in an interview costs you the question

  • Thinks robust scaling deletes or caps extreme values
  • Believes the output is bounded like min-max
  • Assumes the interquartile range is never zero
  • Applies robust scaling everywhere without inspecting the column
  • Says it makes a skewed column symmetric

context