skip to content

Why is Levene's test usually preferred over Bartlett's test for checking equal variances?

level: middleimportance: should knowfreq 38%

answer

  1. one of the two trusts normality
  2. shape problems masquerade as spread problems
  3. test the spread, not the values
  4. absolute deviations from a group centre
  5. median centring is the robust variant

basics

~20 s

Bartlett's test assumes normality and mistakes heavy tails or skew for unequal variance, so it rejects far too often on real data. Levene's test compares absolute deviations from each group's centre and stays reliable when the data are not normal.

solid answer

~50 s

Both test the null that several groups share a common variance, but they earn that answer differently. Bartlett's statistic is derived under the assumption of normality and is unusually sensitive to departures from it: on heavy-tailed or skewed data its false-positive rate climbs well above the nominal level, so a rejection may reflect unequal spread or merely non-normal shape. Levene's test first converts each observation into its absolute deviation from its group's centre, then compares those deviations across groups in an ANOVA-style way. Because it works on deviation magnitudes rather than on the raw values, it is far less dependent on the original shape. Centring on the median rather than the mean - the Brown-Forsythe variant - is more robust again for skewed data and is what most people mean in practice. Bartlett is the better choice only when you genuinely believe the data are normal, where it has more power.

go deeper

for a junior

Know that both tests answer the same question - do these groups share a variance - and that Levene's is the safer default because it does not lean on the data being normal.

for a middle

Explain the mechanism: Bartlett compares logged variances against a chi-square derived under normality, while Levene transforms values into absolute deviations from a group centre and compares those.

for a senior

Demonstrate that you know when the check changes a decision. Discuss the variance-ratio-plus-imbalance combination that actually distorts error rates, and why median centring is the version you reach for.

for a principal

Take a position on gating. Argue whether an automated pipeline should branch on a variance test at all, given that the branch itself distorts the reported error rate and its sensitivity is driven by sample size.

## The shared null hypothesis Both procedures test the same null: that two or more groups are drawn from populations with the same variance, sometimes called homogeneity of variance or homoscedasticity. The alternative is that at least one group's variance differs. The reason anyone runs such a test is that several classical procedures pool variances across groups, and pooling only makes sense if there is one variance to pool. ## Bartlett's test and its fragility Bartlett's statistic compares the logarithm of the pooled variance against a weighted sum of the logarithms of the individual group variances, with a correction factor, and refers the result to a chi-square distribution with one fewer degree of freedom than the number of groups. That reference distribution is derived assuming the underlying data are normal, and the derivation is not robust: the statistic's null distribution depends on the kurtosis of the population. If the data have heavier tails than a normal - which real measurements very often do - the true distribution of the statistic is wider than the chi-square being used, and the test rejects far more often than its stated significance level. The practical failure mode is that you feed it skewed or outlier-prone data with genuinely equal variances, get a small p-value, and conclude the spreads differ when what you have actually detected is non-normality. That is a particularly bad kind of error, because the assumption you were trying to check has been confounded with a different assumption entirely. ## What Levene's test does instead Levene's test sidesteps the shape dependence with a transformation. For each observation, compute the absolute deviation from its own group's centre. Now the question 'do these groups have different spreads?' has been turned into 'do these transformed values have different means?', which is answered with an ordinary between-versus-within-group variance comparison on the deviations. Because absolute deviations are a much better behaved quantity than raw values, the test's error rate holds up reasonably across a wide range of shapes. The choice of centre matters. The original formulation uses the group mean; the widely used variant substitutes the group median, which is the Brown-Forsythe modification. The median version is noticeably more robust on skewed or heavy-tailed data, because the median is not dragged by the same extreme values that inflate the deviations, and it is the default in most practical usage - when someone says 'I ran Levene's test' they usually mean the median-centred version. A third variant uses a trimmed mean, which sits between the two. ## Cost of the robustness Robustness is not free. When the data really are normal, Bartlett's test has more power to detect a given variance ratio than Levene's, because it uses the stronger assumption. That is the whole tradeoff: Bartlett is sharper when its assumption holds and dangerously miscalibrated when it does not, and Levene gives up a little sharpness for a verdict you can trust without first proving the data are normal. Since you rarely believe normality exactly, the robust choice is the sensible default. ## The deeper question: should you gate on it? A subtler point, and a good one to raise unprompted: using a variance test as an automatic gate before choosing your main test is itself questionable. The two-step procedure distorts the error rate of the final comparison, because the choice of test depends on the same data being analysed. Worse, the gate is badly calibrated in exactly the wrong way. With small samples - where unequal variance does the most damage - the variance test has almost no power and waves the data through. With very large samples it fires on variance differences too small to affect anything. Both tests, like all hypothesis tests, become more sensitive as n grows, so the verdict tracks sample size as much as it tracks the thing you care about. The modern practice is therefore to compare spreads descriptively - the ratio of the sample standard deviations, a side-by-side plot of the groups - to decide by design and prior knowledge whether the groups plausibly share a variance, and to prefer a procedure that does not require the assumption in the first place. That leaves Levene's test as a useful description and a reporting convention rather than an automatic switch. Being able to say that, rather than reciting a decision rule, is what distinguishes a strong answer.

  • Should a variance test be used as a gate before choosing which mean-comparison test to run?
    Generally no. A two-step procedure distorts the error rate of the final comparison, because the test choice depends on the same data. The gate is also miscalibrated in the wrong direction: it has almost no power at small samples, where unequal variance hurts most, and it fires on trivial differences at large ones. Most practitioners now default to a procedure that does not assume equal variances and treat the check as description.
  • When do unequal variances actually damage a comparison of two means?
    Mainly when the group sizes are also unequal. With balanced groups the pooled procedure tolerates a fairly large variance ratio. When the smaller group has the larger variance, the true error rate runs above nominal; when the larger group has the larger variance, the test turns conservative and loses power. The variance ratio alone is not the alarm - the ratio combined with imbalance is.
  • What does Levene's test reduce to when there are only two groups?
    The same computation: each value becomes its absolute distance from its own group's centre, and the two sets of distances are compared. With two groups that is effectively a comparison of mean absolute deviations, so the result is easy to sanity-check against the two sample standard deviations. If the standard deviations are close but the test rejects, suspect that a handful of extreme values in one group is driving it.

saying these in an interview costs you the question

  • Runs Bartlett's test on visibly skewed data and trusts the p-value
  • Treats a significant variance test as proof that spreads differ
  • Uses a variance test as a gate, then reports the second p-value as clean
  • Thinks Levene's test requires normally distributed data
  • Ignores that both tests grow more sensitive as the sample grows

context