skip to content

When does Mahalanobis distance flag a point that Euclidean distance calls ordinary?

level: seniorimportance: should knowfreq 38%

answer

  1. circles versus ellipses around the mean
  2. the covariance matrix enters inverted
  3. Euclidean distance after whitening the data
  4. unremarkable alone, impossible together
  5. tall and light when height tracks weight

basics

~20 s

Whenever features are correlated and a point breaks the correlation. Mahalanobis distance divides out the covariance, so a point that is unremarkable on each feature alone but sits off the joint ridge gets a large distance. Euclidean distance sees nothing unusual.

solid answer

~50 s

Mahalanobis distance is `d(x) = sqrt((x - mu)^T * S^-1 * (x - mu))`, where `mu` is the mean vector and `S` the covariance matrix. Multiplying by `S^-1` is the same as rotating and rescaling the space so the features become uncorrelated with unit variance, then measuring ordinary Euclidean distance there. Its contours are ellipses aligned with the data cloud rather than circles. So take height and weight, strongly positively correlated: someone very tall and very light is within a couple of standard deviations on each feature separately, and Euclidean distance places them close to the middle, but they sit far off the height-weight ridge and Mahalanobis distance flags them hard. When `S` is the identity the two measures coincide exactly. The cost is that `S` must be estimated, which needs enough rows relative to columns and is itself sensitive to the points you are trying to find.

go deeper

for a junior

Recall the shape of the idea: Mahalanobis measures distance in units of the data's own spread and correlation, so its contours are ellipses rather than circles around the mean.

for a middle

Explain the formula term by term and why the covariance appears inverted. Be ready to say what happens when the covariance is the identity and when it is merely diagonal.

for a senior

Show operational judgment: how many rows you need before the covariance estimate is trustworthy, what collinear features do to the inverse, and why contaminated data can mask the very points you are looking for.

for a principal

Own the tradeoff between a covariance-aware measure nobody can interpret and a plain scaled distance everyone can. Be ready to argue when the extra sensitivity justifies the estimation burden and the explanation cost.

## The definition For a point `x` in `p` dimensions, a mean vector `mu` and a covariance matrix `S`, the Mahalanobis distance is `d(x) = sqrt( (x - mu)^T * S^-1 * (x - mu) )` The same formula measures the distance between two points by replacing `mu` with the second point. Everything hinges on the inverse covariance matrix sandwiched in the middle. ## What the inverse covariance actually does Write `S^-1 = S^(-1/2) * S^(-1/2)`. Then the formula is just `|| S^(-1/2) * (x - mu) ||_2`: the ordinary Euclidean norm applied **after** a linear transformation that whitens the data, meaning it rotates the cloud onto its own principal directions and rescales each of them to unit variance. Mahalanobis distance is therefore not a new kind of distance at all; it is Euclidean distance measured in coordinates where the data is round. Two consequences follow directly. - If `S` is the identity matrix, the transformation does nothing and Mahalanobis distance **is** Euclidean distance. - If `S` is diagonal but with unequal variances, the transformation only rescales each axis, so Mahalanobis reduces to a Euclidean distance on per-feature scaled values. The genuinely new behaviour appears only when `S` has nonzero off-diagonal entries, that is, when features are correlated. Its contours are ellipsoids whose axes follow the shape of the data cloud, where Euclidean contours are spheres that ignore it. ## The height and weight case Suppose adult height and weight in some population have a correlation of about 0.8. Consider a person whose height is 2 standard deviations above the mean and whose weight is 1.25 standard deviations below it. On each feature alone this person is unusual but not alarming. Even measuring Euclidean distance on the scaled values gives `sqrt(2^2 + 1.25^2)` which is about 2.36, a perfectly common reading in a large sample. Now put the correlation back in. With correlation `r`, the squared Mahalanobis distance in scaled coordinates is `(z1^2 - 2*r*z1*z2 + z2^2) / (1 - r^2)`. With `z1 = 2`, `z2 = -1.25` and `r = 0.8`, the cross term contributes `+4` instead of subtracting, the numerator is 9.5625 and the denominator is 0.36, giving a squared distance of about 26.6 and a distance of about 5.2. The combination tall-and-light is rare precisely because tall people are usually heavy, and only the covariance-aware measure knows that. ## Reading the number If the data really is multivariate normal and `mu` and `S` are the true parameters, the squared Mahalanobis distance follows a chi-squared distribution with `p` degrees of freedom, `p` being the number of features. That gives a principled cutoff rather than an eyeballed one. In two dimensions the chi-squared with 2 degrees of freedom has the tidy tail `P(D^2 > c) = exp(-c / 2)`, so the value 26.6 above corresponds to roughly one in six hundred thousand. Note the conditions: true parameters, and multivariate normality. With estimated parameters the reference distribution is only approximate, and with skewed or multimodal data the cutoff can be badly wrong even though the distance itself is still measuring the right geometry. ## The operational difficulties - **Estimating `S` needs data.** A `p` by `p` covariance matrix has `p * (p + 1) / 2` free parameters. With few rows relative to columns the estimate is noisy, and if rows are fewer than columns the sample covariance is singular and cannot be inverted at all. Regularising or shrinking the covariance toward a diagonal target is the standard remedy. - **Collinear features break the inverse.** Two nearly duplicated features make `S` nearly singular, and the inverse then amplifies noise in that direction into enormous distances. Drop or combine such features first. - **The estimate is contaminated by what you are hunting.** If unusual points are included when `S` is computed, they inflate the covariance in their own direction and shrink their own distance. Several together can hide each other completely. A robust covariance estimate, computed from a trimmed subset of the data, avoids this. - **Interpretability.** A Euclidean distance is in the units of the data. A Mahalanobis distance is unitless, which is a strength for combining features on different scales and a weakness when a stakeholder asks what the number means. ## The reverse case The asymmetry runs both ways. A point that is extreme on one feature and equally extreme on a strongly correlated partner, tall **and** heavy, has a large Euclidean distance from the mean but a much smaller Mahalanobis distance, because it lies along the ridge where the data naturally extends. Whether that is the right answer depends on the question: for consistency with the joint pattern, Mahalanobis is right; for sheer magnitude, it is not.

  • What happens to Mahalanobis distance when two features are almost perfectly correlated?
    The covariance matrix becomes nearly singular, and inverting it amplifies the tiny remaining variance in the collapsed direction. Small measurement noise perpendicular to the ridge then produces huge distances, so the measure becomes unstable and unusable. The fix is to drop one feature, combine them, or shrink the covariance estimate toward a better-conditioned target.
  • Why can several unusual points hide each other from Mahalanobis distance?
    Because the covariance matrix is estimated from the same data that contains them. A cluster of unusual points inflates the covariance along their own direction, which is exactly the direction that would otherwise make them stand out, so each ends up with a modest distance. Estimating the covariance from a trimmed or robust subset restores the sensitivity.
  • Is Mahalanobis distance a proper metric?
    Yes, provided the covariance matrix is fixed and positive definite. It is the Euclidean norm applied after an invertible linear transformation, so it inherits symmetry, non-negativity, zero only for identical points, and the triangle inequality. Note that this holds for the distance itself, not its square.

Judging how unusual a person's shoes are: size 13 alone is fine and a 5-foot frame alone is fine, but the pair together is startling because the two normally travel together.

saying these in an interview costs you the question

  • Confuses Mahalanobis distance with per-feature rescaling only
  • Uses the covariance matrix without inverting it
  • Forgets that the covariance must be estimated from data
  • Applies it when features are nearly collinear
  • Assumes the chi-squared cutoff holds for any distribution

context