skip to content

T-Tests and Z-Tests

One-sample, two-sample and paired tests of means: Student's version assumes equal variances, Welch's does not, and a known variance lets you use z. Paired versus unpaired trips people up.

on this pageshow

questions

5

When should you use a paired t-test instead of a two-sample t-test?

level: juniorimportance: must knowfreq 76%

answer

  1. start from how the data were collected
  2. are the two columns the same units?
  3. reduce two columns to one of differences
  4. degrees of freedom come from pairs
  5. positive correlation cancels between-unit spread

basics

~20 s

Use a paired t-test when each value in one group is naturally matched to one in the other, such as the same patient measured before and after treatment. It tests the mean of the within-pair differences.

solid answer

~50 s

The deciding question is whether the two columns are linked observation by observation. If 30 patients each have a blood pressure reading before and after a drug, the two columns are the same 30 people, so I compute the 30 within-patient differences and run a one-sample t-test on them: `t = dbar / (s_d / sqrt(n))` with `n - 1 = 29` degrees of freedom. A two-sample t-test instead assumes the two groups are independent samples of different units, and its degrees of freedom come from the total sample size. Running the unpaired version on paired data violates independence and, because before and after readings on the same person are positively correlated, usually inflates the standard error and throws away power. Pairing is a property of how the data were collected, not something you can apply afterwards to two unrelated groups.

go deeper

for a junior

Be ready to name the trigger in one sentence: the same units measured twice, so you analyse the differences. Know that the degrees of freedom come from the number of pairs minus one.

for a middle

Explain the mechanics: collapse to one column of differences and run a one-sample t on it, and show why the variance of a difference shrinks when the two measurements are positively correlated.

for a senior

Show judgment about designs you have run: recognise a within-subject or matched-pairs structure in a messy dataset, and explain what an unpaired analysis of it costs in power and validity.

for a principal

Own the design decision before data collection. Argue when a within-subject design is worth its carryover and dropout risks versus a simpler parallel-group design, and what that choice implies for required sample size.

## The design decides the test A t-test compares means, but which t-test you run is fixed by how the data were collected, not by what you would like to conclude. The **paired** t-test applies when every observation in one condition is tied to exactly one observation in the other: the same subject measured twice, the same machine run under two settings, twins split across two arms, or units matched on covariates before assignment. The **two-sample** (unpaired) t-test applies when the two groups contain different, independent units. ## What the paired test actually does The paired test is not a special two-group procedure at all. It collapses the two columns into one: ``` d_i = after_i - before_i for i = 1 .. n pairs ``` and then runs an ordinary **one-sample t-test** on those differences against a null value of zero: ``` t = dbar / (s_d / sqrt(n)), df = n - 1 ``` where `dbar` is the mean of the differences and `s_d` is their sample standard deviation. With 30 patients measured before and after, you have 30 differences and 29 degrees of freedom — not 58 or 60. Every subject contributes one number, so the sample size for the test is the number of *pairs*. ## Why pairing helps Suppose the before reading has variance `sigma_1^2`, the after reading has variance `sigma_2^2`, and the two are correlated with coefficient `rho`. The variance of the difference is ``` Var(after - before) = sigma_1^2 + sigma_2^2 - 2 * rho * sigma_1 * sigma_2 ``` When `rho > 0` — which is the normal situation, because a patient with high blood pressure before tends to have relatively high blood pressure after — the `-2 * rho * sigma_1 * sigma_2` term shrinks the variance the test has to fight against. All the stable, person-specific level cancels inside each difference, and what remains is the change. That is why a paired design can detect a small average effect with a modest number of subjects, while an unpaired comparison of the same two columns would be swamped by the fact that people simply differ from one another. The unpaired test on paired data does not use that cancellation: it estimates the variability of the before column and the after column separately, treating the large between-person spread as noise. The result is a larger standard error, a smaller t statistic, and a real loss of power. It is also formally invalid, because the two-sample test assumes the two groups are independent and they are not. ## The cost of pairing Pairing is not free. You spend degrees of freedom: 30 pairs give 29 df, whereas 30 subjects in each of two independent arms would give 58 df. When the correlation between the two measurements is near zero, the variance reduction does not materialise and the paired analysis is slightly *less* powerful than the unpaired one on the same number of observations. In practice, repeated measurements on the same unit are strongly correlated, so pairing wins comfortably — but the claim "paired is always better" is wrong, and an interviewer may probe it. ## Common setups and which test they call for - Same subjects measured before and after an intervention: **paired**. - Each subject tries both variants (a crossover or within-subject design): **paired**. - Subjects matched into pairs on age and baseline severity, then one of each pair randomised: **paired** on the matched pairs. - Two independent groups of different people, treatment versus control: **two-sample**. - Two groups of different sizes: **must** be two-sample — unequal group sizes make pairing impossible, and that alone tells you the design was not paired. ## Direction of the difference Once you pair, decide the sign convention and keep it: `after - before` makes a positive `dbar` mean "the value went up". The t statistic flips sign if you subtract the other way, but the two-sided p-value is identical. Only the interpretation of a one-sided alternative changes, so state the direction in the hypothesis before you look at the data. ## What to say in an interview Lead with the design question — "are these the same units measured twice?" — then name the mechanic: reduce to one column of differences, one-sample t on that column, `n - 1` degrees of freedom where `n` is the number of pairs, and the payoff is that stable between-unit variation cancels. Mentioning that the two-sample test on paired data breaks the independence assumption and usually costs power shows you understand *why* the rule exists rather than having memorised it.

  • With 30 patients measured before and after, how many degrees of freedom does the paired t-test have, and why?
    29. The paired test reduces the data to 30 within-patient differences and runs a one-sample t-test on them, so the degrees of freedom are the number of pairs minus one. The 60 raw measurements are not 60 independent observations, and df = n1 + n2 - 2 = 58 would be the answer for two independent groups of 30, which is a different design.
  • What happens if you run an unpaired two-sample t-test on data that were actually paired?
    It violates the independence assumption, and when the paired measurements are positively correlated it typically overstates the standard error of the difference. The point estimate of the mean difference is unchanged, but the t statistic shrinks and the p-value grows, so you lose power and may miss a real effect. The error is usually conservative, which is why it survives code review so often.
  • Is a paired design always more powerful than an unpaired one with the same number of observations?
    No. Pairing trades degrees of freedom for variance reduction: 30 pairs give 29 df while two independent groups of 30 give 58. The trade pays off only when the paired measurements are positively correlated. If the correlation is near zero, the paired analysis has the same variance and fewer degrees of freedom, so it is marginally weaker.

Weighing yourself before and after a diet on the same scale tells you more than comparing your weight to a stranger's: the pairing cancels everything about you that never changed.

saying these in an interview costs you the question

  • Treats before and after readings on the same person as independent
  • Uses df = n1 + n2 - 2 for a paired design
  • Thinks pairing just means the two groups have equal sample sizes
  • Claims you can pair two independent groups after collecting the data
  • Asserts paired is always more powerful regardless of correlation

context

open as a page

When is a z-test valid for testing a mean instead of a t-test?

level: middleimportance: must knowfreq 68%

basics

~20 s

A z-test for a mean is valid only when the population standard deviation is known rather than estimated. If you plug in the sample standard deviation, the extra uncertainty makes the statistic follow a t distribution instead.

open as a page

Why is Welch's t-test usually a safer default than Student's pooled t-test?

level: middleimportance: should knowfreq 58%

basics

~20 s

Student's two-sample t-test pools the two groups into one variance estimate, which is only valid when their variances are equal. Welch's test keeps them separate and adjusts the degrees of freedom, so it stays accurate under unequal variances.

open as a page

In a two-proportion z-test comparing 42/300 with 61/300, when is the standard error pooled?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Pool for the hypothesis test, not for the confidence interval. The test assumes the null that both rates are equal, so one combined proportion estimates their shared variance; an interval must allow them to differ.

open as a page

How does the t distribution with 4 degrees of freedom differ from the standard normal?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

Both are symmetric bell curves centred at zero, but t with 4 degrees of freedom has heavier tails and variance 2 rather than 1. Its two-sided 95% cutoff is about 2.78 against the normal's 1.96.

open as a page