Why can a plain z-score miss an outlier that a MAD-based modified z-score flags?
answer
- the suspect helps compute its own threshold
- an inflated denominator hides the point
- swap the centre and the scale for rank-based ones
- median of absolute deviations from the median
basics
~20 sThe plain z-score divides by the sample standard deviation, which the outlier itself inflates, so the point hides behind the spread it created. The modified z-score uses the median and the median absolute deviation, neither of which the outlier can move.
solid answer
~50 sA z-score is `(x - mean) / s`. The suspect point is inside both the mean and `s`, so it pulls the centre toward itself and inflates the denominator — the effect is called masking, and with two or more extreme values it gets much worse because they inflate `s` together. There is even a hard ceiling: with the sample standard deviation, no observation in a sample of size `n` can have an absolute z-score above `(n - 1) / sqrt(n)`, so in a sample of 10 nothing can ever exceed about 2.85 and a threshold of 3 flags nothing at all. The modified z-score replaces both pieces: `0.6745 * (x - median) / MAD`, where `MAD` is the median of the absolute deviations from the median. Both parts have a 50% breakdown point, so contamination cannot set its own threshold. The usual cutoff is 3.5.
go deeper
Know that a z-score is the distance from the mean measured in standard deviations, and that a rule of thumb flags values beyond about three. Be able to compute one.
Explain masking mechanically: the suspect point inflates the standard deviation it is divided by. State the modified z-score formula and what MAD means.
Demonstrate that you have hit the practical failure modes — zero MAD on rounded or count data, clusters of extremes masking each other, and cutoffs that are conventions rather than error rates.
Own the framing that any threshold computed from possibly-contaminated data is circular, and be ready to argue when a univariate rule should be replaced entirely rather than hardened.
## The plain z-score and its self-defeating denominator The standard score of an observation is `z = (x - mean) / s`, where `s` is the sample standard deviation. A rule such as flag anything with `|z| > 3` is the first outlier detector most people learn, and it has a structural flaw: the point under suspicion contributes to both the mean and `s`. A large value drags the mean toward itself, shrinking the numerator, and simultaneously inflates `s`, enlarging the denominator. Both effects push `|z|` down. The outlier is, in effect, allowed to widen the goalposts it must clear. ## Masking and swamping When one extreme value prevents itself or a companion from being flagged, the effect is called masking. It is worst with clusters: two or three extreme values inflate `s` together, so none of them individually looks far from the centre in standard-deviation units. The mirror-image failure is swamping, where the presence of extreme values inflates the spread and shifts the centre enough that perfectly ordinary observations on the opposite side start looking unusual. Both follow from the same cause — a threshold computed from contaminated data. There is also a hard arithmetic ceiling. For a sample of size `n` with `s` computed using the `n - 1` denominator, the largest achievable absolute z-score for any observation is `(n - 1) / sqrt(n)`. For `n = 10` that is about 2.85, so a cutoff of 3 can never fire regardless of how absurd one value is. For `n = 20` the ceiling is about 4.25. Any small-sample z-score rule is partly measuring sample size rather than unusualness. ## The robust replacement The modified z-score swaps in two rank-based quantities: - centre: the median - scale: `MAD = median(|x_i - median|)`, the median of the absolute deviations from the median The statistic is `M = 0.6745 * (x - median) / MAD`, and the conventional cutoff is `|M| > 3.5`. The constant `0.6745` is the standard normal 0.75 quantile; equivalently, `1.4826 * MAD` is a consistent estimator of the standard deviation for normal data, since `1 / 0.6745` is about `1.4826`. Scaling this way makes the modified z-score comparable in magnitude to an ordinary z-score when the data really is normal and clean, which is what makes a threshold near 3.5 meaningful rather than arbitrary. The robustness comes from breakdown points. The sample standard deviation has a breakdown point of 0% — one value moves it without limit. The `MAD` has a breakdown point of 50%, the same as the median, so nearly half the sample would have to be corrupt before the scale estimate itself is compromised. The suspect point no longer sets its own threshold, which is precisely why the modified z-score catches what the plain z-score hides. ## Worked shape of the failure Imagine a small column of nearly identical measurements plus one value orders of magnitude larger. The mean is dragged well above the bulk and `s` is enormous, so the extreme value scores perhaps 2.5 in z units and passes a cutoff of 3. The median sits squarely in the bulk and the `MAD` reflects the bulk's tiny spread, so the same point's modified z-score is huge and is flagged immediately. Nothing about the data changed; only which statistics were allowed to be contaminated. ## Where the MAD breaks The `MAD` is zero whenever more than half the observations share the same value. This happens routinely with coarse counters, rounded readings, zero-inflated columns and low-cardinality integers. A zero denominator makes the modified z-score undefined or infinite for every distinct value, which usually surfaces as an implausible flood of flagged points. When that occurs, switch to a scale estimate built from a wider slice of the distribution — an interquartile-based measure, for instance — or accept that a univariate outlier rule is the wrong tool for that column. Two further caveats apply to both statistics. They are univariate, so a point that is unremarkable on every variable separately but impossible in combination will not be flagged by either. And neither is a hypothesis test: 3 and 3.5 are conventions, not error-rate guarantees, and on genuinely heavy-tailed data a robust rule will still flag real observations by the fistful. ## What interviewers listen for The words masking and breakdown point, the correct definition of `MAD` (median of absolute deviations from the median — not the mean of absolute deviations from the mean), awareness that a scaling constant exists so the cutoff is comparable, and the practical knowledge that a zero `MAD` is a common real-world failure.
- What exactly is the MAD, and what is the constant 1.4826 for?`MAD` is the median of the absolute deviations from the median: take the median, take `|x_i - median|` for every point, then take the median of those. It is a rank-based scale estimate with a 50% breakdown point. Multiplying it by `1.4826` makes it a consistent estimator of the standard deviation when the data is normal, so a robust scale can be compared against thresholds calibrated in familiar standard-deviation units.
- What happens to the modified z-score when the MAD is zero?It becomes undefined or infinite, and every value differing from the median gets flagged. A zero `MAD` means more than half the observations share the same value, which is common in rounded readings, count columns and zero-inflated data. The fix is a scale estimate that reads a wider slice of the distribution, such as an interquartile-based measure, or dropping the univariate rule for that column entirely.
- Can a modified z-score rule still miss an outlier?Yes, in two ways. It is univariate, so a record whose individual values are all ordinary but whose combination is impossible passes untouched. And if more than half the sample is contaminated, the median and `MAD` describe the contamination rather than the clean data, so the genuine values become the flagged ones. Robustness raises the breakdown point to 50%; it does not eliminate the failure mode.
Judging a sprinter against the average of a field that includes them is like grading on a curve the cheater also sits in — their result shifts the curve enough to make itself look ordinary.
saying these in an interview costs you the question
- Defines MAD as the mean of absolute deviations from the mean
- Treats a z-score above 3 as a proven error
- Ignores that the outlier inflates its own denominator
- Applies a z-score cutoff of 3 to a sample of ten values
- Assumes a robust rule never flags genuine observations