Which assumptions must hold for a two-sample t-test comparing group means to be valid?
answer
- three of them, ranked by danger
- one comes from the design, not the data
- the CLT covers only one of the three
- normality of the mean, not raw values
- pooling adds an equal-spread condition
basics
~20 sA two-sample t-test assumes observations are independent within and between groups, that the sampling distribution of each group mean is approximately normal, and, for the pooled version, that the two populations share a common variance.
solid answer
~50 sThree things, in order of danger. First, independence: every row contributes fresh information, within a group and across the two groups. That comes from the design, not from a test, and no amount of extra data repairs it. Second, approximate normality of the sampling distribution of the mean, not of the raw values: with normal data it holds exactly, and with a decent sample size the central limit theorem makes it hold approximately unless the data are extremely skewed or heavy-tailed. Third, the pooled version assumes the two populations have the same variance; I check that by comparing group standard deviations or with a spread test such as Levene's. In practice I plot each group on a normal Q-Q plot, compare the spreads, and spend most of my attention on independence, because that is the assumption with no rescue.
go deeper
Be ready to name the three assumptions cleanly - independence, approximate normality of the group mean, equal variances for the pooled version - and to say which check pairs with which assumption.
Explain why the normality requirement concerns the sampling distribution of the mean rather than the raw values, and be specific about what the central limit theorem does and does not buy you at a given sample size.
Show that you rank assumptions by consequence. Describe a dataset where independence quietly fails and explain why collecting more rows makes the resulting error worse rather than better.
Own the standard your team applies: when an assumption check is worth blocking a decision on, how it gets documented, and how you stop analysts from shopping for the test that gives the answer they hoped for.
## What the test actually needs A two-sample t-test asks whether two population means differ. It builds a statistic of the form `t = (mean1 - mean2) / SE`, where `SE` is the estimated standard error of the difference, and then compares that statistic to a t distribution. Everything the test assumes exists to make that last step true: that the numerator is approximately normal around the true difference, and that the denominator is an independent estimate of its spread with the stated degrees of freedom. ## Assumption 1 — independence Observations must be independent both within each group and between the groups. This is a property of how the data were produced, and it cannot be checked by looking at the numbers alone. You establish it by asking what a single row represents and whether any unit contributes more than one row: repeated measurements on the same subject, multiple sessions from one user, pupils inside a classroom, orders inside a household, consecutive readings from one sensor. When rows are correlated, the standard error is computed as if there were far more information than there really is, so it comes out too small; intervals are too narrow and p-values too optimistic. Crucially, collecting more correlated rows makes this worse, not better. Independence is also why random sampling (or random assignment) matters: it is the design feature that buys the assumption. ## Assumption 2 — normality, of the right thing The common misstatement is that the raw data must be normal. What the test needs is that the sampling distribution of each group mean is approximately normal. Normal raw data is a sufficient condition for that, not a necessary one. The central limit theorem says that as sample size grows, the distribution of a sample mean approaches normality for any population with finite variance. How fast depends on the shape: mildly skewed data converge quickly, and by a few dozen observations per group the approximation is usually good; strongly skewed or heavy-tailed data need far more. Two situations therefore deserve care. Small samples, where the theorem has not had a chance to work, so raw-data shape matters directly. And heavy tails or outliers at any size, because a single extreme value inflates the sample standard deviation, which is what the denominator of the statistic is built from — the error rate may be fine while power quietly collapses. You check this with a normal Q-Q plot of each group: sorted values against the values a normal distribution would place at those positions. A systematic bend tells you the direction and size of the departure. A formal normality test tells you far less, because its verdict tracks sample size more than it tracks any departure that matters. ## Assumption 3 — equal variances The classical pooled version combines both groups' variances into one estimate, which is only sensible if the populations share a variance. Compare the two sample standard deviations, or use a spread test — Levene's, which works on absolute deviations from a group centre, is the robust choice. Unequal variance matters most when the group sizes are also unequal: with balanced groups the pooled procedure tolerates a fairly large variance ratio, whereas an imbalanced design where the smaller group has the larger variance pushes the true error rate above nominal. A procedure that does not pool the variances removes this assumption entirely, which is why many practitioners default to one. ## A fourth, quieter requirement The mean has to be a meaningful summary. On a strongly skewed outcome the mean is not what a stakeholder pictures, and on an ordinal scale it may not be defined at all. Passing every assumption check does not make the mean the right question. ## How to talk about it in an interview Name the three, then rank them. Say that normality applies to the sampling distribution of the mean and that the central limit theorem covers it at reasonable sample sizes; say that equal variance is only a requirement of the pooled version; say that independence is structural, undetectable from the numbers alone, and the one that more data makes worse. Add that assumption checks should be decided before you see the outcome — choosing the test after seeing which one gives the smaller p-value invalidates the error rate you are reporting.
- With 200 observations per group and moderate right skew, is non-normality still a problem?Usually not. The test needs the sampling distribution of the mean to be near-normal, and at n = 200 the central limit theorem delivers that for moderate skew. I would still look for extreme outliers or a very heavy tail, because a handful of extreme values inflates the sample standard deviation and drains power even when the error rate stays near nominal. Skew that severe is a reason to question whether the mean is the right summary, not just the test.
- Does the t-test assume the raw observations are normally distributed?No. It assumes the sampling distribution of the group mean is approximately normal. Normal raw data is a sufficient condition, not a necessary one; the population's shape and the sample size together decide whether the approximation holds. That distinction is why a normality test run on the raw sample answers a slightly different question from the one the t-test actually cares about.
- How would you check the independence assumption?By reading how the data were collected, not by running a test. I ask what a single row is and whether any unit contributes more than one: repeated measurements per subject, several sessions per user, pupils inside a classroom, orders inside a household. If units repeat, the rows are correlated and the row count overstates the information available, so the reported standard error will be too small.
saying these in an interview costs you the question
- Claims the raw data themselves must be normally distributed
- Runs a normality test and stops there, never asking about independence
- Says a large sample fixes every assumption, dependence included
- Treats equal variance as required by every version of the t-test
- Checks assumptions after seeing the p-value, then picks the friendlier test