Why is Welch's t-test usually a safer default than Student's pooled t-test?
answer
- one test makes an extra assumption
- pooling averages two variances by sample size
- unbalanced groups are where it breaks
- the denominator keeps s1^2/n1 and s2^2/n2 apart
- expect fractional degrees of freedom
basics
~20 sStudent's two-sample t-test pools the two groups into one variance estimate, which is only valid when their variances are equal. Welch's test keeps them separate and adjusts the degrees of freedom, so it stays accurate under unequal variances.
solid answer
~50 sStudent's test assumes both groups share one variance, so it pools them as `s_p^2 = ((n1-1)*s1^2 + (n2-1)*s2^2) / (n1+n2-2)` and uses `n1 + n2 - 2` degrees of freedom. That weighting is by sample size, which is exactly wrong when the smaller group is the noisier one. With `n = 10` and a large variance against `n = 200` with a small variance, the 199 degrees of freedom from the big group dominate the pooled estimate, the standard error comes out too small, the t statistic too large, and the false-positive rate well above 5%. Welch's test never pools: the denominator is `sqrt(s1^2/n1 + s2^2/n2)`, and the Welch-Satterthwaite degrees of freedom come out fractional — here near 9 or 10 rather than 208. With equal variances and equal group sizes the two agree closely, so defaulting to Welch costs almost nothing.
go deeper
Be ready to say that Student's two-sample t-test assumes both groups have the same variance and Welch's does not, and that Welch is the safer default.
Explain the mechanics: how the pooled variance weights each group by its degrees of freedom, why Welch keeps the two terms separate, and where the fractional degrees of freedom come from.
Diagnose it in a real comparison: spot an unbalanced design with a small noisy arm, predict the direction of the error, and explain why the reported degrees of freedom collapsed toward the small group.
Own the standard your teams use. Decide whether to mandate Welch across all two-group comparisons for uniformity of reported error rates, and how to handle historical results computed with the pooled test.
## Two ways to build the denominator Both two-sample tests have the same numerator, the difference in sample means `xbar1 - xbar2`. They differ in how they estimate the variability of that difference. **Student's pooled t-test** assumes a single common variance `sigma^2` in both groups and estimates it by a degrees-of-freedom-weighted average: ``` s_p^2 = ((n1 - 1) * s1^2 + (n2 - 1) * s2^2) / (n1 + n2 - 2) t = (xbar1 - xbar2) / sqrt(s_p^2 * (1/n1 + 1/n2)) df = n1 + n2 - 2 ``` **Welch's t-test** makes no such assumption and never mixes the two variance estimates: ``` t = (xbar1 - xbar2) / sqrt(s1^2/n1 + s2^2/n2) df = (s1^2/n1 + s2^2/n2)^2 / ( (s1^2/n1)^2/(n1 - 1) + (s2^2/n2)^2/(n2 - 1) ) ``` That second expression is the **Welch-Satterthwaite** approximation. It is not required to be an integer. ## Where pooling goes wrong Pooling weights each group's variance by its degrees of freedom, so the bigger sample dominates. That is the right thing to do *if* both groups genuinely share a variance — you are combining two estimates of the same quantity. If they do not, the pooled number is an average of two different things, and whether it is too big or too small depends on how variance lines up with sample size: - **Larger variance in the smaller group** (the dangerous case). Take one arm with `n1 = 10` and a large variance against another with `n2 = 200` and a small variance. The pooled estimate is pulled toward the low-variance arm because that arm carries 199 of the 208 degrees of freedom. The standard error of the difference is then understated, the t statistic is inflated, and the claimed degrees of freedom of 208 make the critical value smaller still. The test is **anti-conservative**: its true false-positive rate can run well above the nominal 5%. - **Larger variance in the bigger group.** The bias runs the other way and the test becomes conservative — it under-rejects and loses power. - **Equal sample sizes.** Pooling is fairly forgiving of unequal variances when `n1 = n2`, which is why the pooled test survives so long in balanced experiments. The trap is that unbalanced designs are common — a small pilot arm against a large control, a rare segment against everyone else — and unbalanced designs are exactly where pooling misbehaves. ## What Welch's degrees of freedom do Welch's df always lands between `min(n1, n2) - 1` and `n1 + n2 - 2`. Which end it approaches depends on which group contributes more to `s1^2/n1 + s2^2/n2`: - If one group's `s^2/n` term dominates the sum, the effective degrees of freedom collapse toward that group's own `n - 1`. In the 10-versus-200 example, the small noisy arm's term dominates and the df land near 9 or 10, not 208. - If the two terms are comparable and the samples balanced, the df approach `n1 + n2 - 2` and Welch nearly coincides with Student. So Welch does two things at once: it fixes the standard error, and it charges you honestly for the fact that most of your information about the difference came from a small, noisy group. Reporting a fractional value such as 9.4 degrees of freedom is expected, not a bug; the fraction is what the Satterthwaite approximation produces when it matches moments between the true distribution and a t distribution. ## What does defaulting to Welch cost? Very little. When the variances are genuinely equal and the groups balanced, Welch's degrees of freedom sit just below the pooled value and the two p-values differ in the third decimal. Simulation studies of this comparison consistently show Welch holding its nominal error rate across variance ratios where Student's does not, while giving up only a sliver of power in the equal-variance case that Student's assumes. That asymmetry — large protection, tiny cost — is the whole argument for making Welch the default. Several statistical environments have made it the default two-sample test for precisely this reason. ## Related mechanics worth knowing - Welch's test still assumes the two samples are **independent** and that each group's sampling distribution of the mean is approximately normal. It relaxes the equal-variance assumption only. - With `n1 = n2` and equal variances, the Welch and pooled statistics are algebraically the same number; only the degrees of freedom differ slightly. - The confidence interval for the difference follows the same split: `(xbar1 - xbar2) +/- t_crit * sqrt(s1^2/n1 + s2^2/n2)`, with `t_crit` taken at the Welch degrees of freedom. ## How to answer State the assumption that separates them (equal variances), give the formula difference in one line (pool versus keep separate), name the failure case concretely — a small high-variance arm against a large low-variance arm inflates false positives — and close with the practical rule: default to Welch, because when Student's assumption holds the two agree, and when it does not, only one of them is still correct.
- Why does Welch's test report fractional degrees of freedom such as 9.4?The exact distribution of the Welch statistic is not a t distribution, so the Satterthwaite approximation finds the t distribution that matches it best by matching moments. The matching value is a continuous quantity, not a count of independent pieces of information, so there is no reason for it to be a whole number. Rounding down is a mildly conservative simplification.
- With n = 10 and a large variance against n = 200 with a small variance, which way does the pooled test err?It errs toward false positives. The pooled variance is dominated by the 199 degrees of freedom from the low-variance arm, so it understates the true variability of the small arm's mean. The standard error of the difference comes out too small and the 208 claimed degrees of freedom make the critical value too lenient, so the test rejects more often than its nominal 5%.
- When do Welch's and Student's tests give effectively the same answer?When the two sample variances are similar and the group sizes are equal. With equal n the two statistics are algebraically identical and only the degrees of freedom differ, by a small amount that barely moves the critical value. That is exactly why defaulting to Welch is nearly free: in the situation Student's test assumes, it agrees with Student's test.
Pooling is like quoting one average commute time for a whole city: harmless if everyone lives the same distance out, badly misleading if the loudest outlier is also the smallest neighbourhood.
saying these in an interview costs you the question
- Says Welch is only for small samples
- Thinks pooling is safe whenever both groups look roughly similar
- Treats fractional degrees of freedom as a software bug
- Claims unequal variances always make a test conservative
- Picks whichever of the two gives the smaller p-value