skip to content

When is a z-test valid for testing a mean instead of a t-test?

level: middleimportance: must knowfreq 68%

answer

  1. look at the denominator, not the numerator
  2. who supplied the standard deviation?
  3. estimating it adds a second source of noise
  4. a null on a proportion pins its variance
  5. the gap disappears as n grows

basics

~20 s

A z-test for a mean is valid only when the population standard deviation is known rather than estimated. If you plug in the sample standard deviation, the extra uncertainty makes the statistic follow a t distribution instead.

solid answer

~50 s

Both statistics have the same shape, `(estimate - null value) / standard error`; the difference is where the standard error comes from. The z-test uses a population standard deviation you already know, so the denominator is a fixed constant and the statistic is standard normal under the null. The t-test uses the sample standard deviation `s`, which is itself a random quantity, and that extra estimation noise is what gives the t distribution with `n - 1` degrees of freedom its heavier tails. In real work the population standard deviation is almost never known, so t is the default for a mean. The honest exception is a proportion: a null such as `p = 0.5` pins the variance at `p0 * (1 - p0)`, so nothing is estimated and a z-test genuinely applies. With a few hundred observations the two agree anyway, so this matters most for small samples.

go deeper

for a junior

Be ready to say that a t-test is what you use in practice because the population standard deviation is unknown, and that the t distribution carries degrees of freedom while the normal does not.

for a middle

Explain the mechanism: substituting the sample standard deviation makes the denominator random, which fattens the tails, and the correction fades as degrees of freedom grow.

for a senior

Show you know when the distinction stops mattering and when it bites. Be able to say why a proportion test is legitimately a z-test and why a small-sample analysis with a normal cutoff is over-rejecting.

for a principal

Own the convention your organisation uses. Decide whether reporting z everywhere for large samples is acceptable simplification or a habit that will quietly mislead when someone reruns the analysis on a small segment.

## Same skeleton, different denominator Every test in this family is built the same way: ``` statistic = (observed estimate - value under the null) / (standard error of the estimate) ``` Testing a battery's mean life against a manufacturer's claim of 40 hours, the numerator is `xbar - 40`. What separates a z-test from a t-test is entirely the denominator. - **z-test for a mean:** the denominator is `sigma / sqrt(n)`, where `sigma` is the *population* standard deviation, treated as a known constant. Under the null, and given normal data (or a large enough sample), the statistic follows the standard normal distribution. - **t-test for a mean:** the denominator is `s / sqrt(n)`, where `s` is the *sample* standard deviation computed from the same data. Under the null, the statistic follows a t distribution with `n - 1` degrees of freedom. ## Why estimating the denominator changes the distribution If `sigma` is known, only the numerator varies from sample to sample, and the ratio is normal. If you substitute `s`, the denominator wobbles too. Sometimes `s` happens to come out small, and a modest deviation in the numerator is then divided by a small number, producing an unusually large statistic. That double source of randomness makes extreme values more common than the normal distribution predicts — the t distribution's heavier tails are exactly the price of not knowing `sigma`. Using a normal critical value while estimating `sigma` from a small sample therefore rejects too often: the test's real false-positive rate exceeds its nominal level. ## When is sigma actually known? Almost never for a mean measured on a fresh sample. Textbook cases where the claim is defensible: - A measuring instrument with a manufacturer-certified precision, where the measurement error is a documented property of the device rather than something you estimate. - A long-running, stable process with years of historical data, where the process standard deviation is known far more precisely than anything the current small sample could tell you. - A simulation or a designed sampling scheme where you set the noise level yourself. Even then, "known" is a modelling assumption: the historical standard deviation is still an estimate, just one based on so much data that its own uncertainty is negligible. ## The case where z is genuinely right: proportions For a proportion, the variance is a function of the mean. A binary outcome with success probability `p` has variance `p * (1 - p)`, so under a *specific* null hypothesis such as `p = 0.5`, the standard error under the null is completely determined: ``` SE_0 = sqrt(p0 * (1 - p0) / n) ``` Nothing is estimated, so a normal reference distribution is appropriate and the test is legitimately a z-test. This is why textbooks introduce z-tests for proportions and t-tests for means: it is not an arbitrary convention, it reflects which parameters the null pins down. The caveat is that the normal approximation to the binomial needs enough data — the usual rough guideline is at least about 10 expected successes and 10 expected failures. ## How much does it matter in practice? Less than the emphasis in a first course suggests, once `n` is large. The two-sided 5% critical value is 1.96 for the normal. For a t distribution it is about 2.78 at 4 degrees of freedom, about 2.09 at 20, about 2.01 at 50, and about 1.97 at 200. So: - With a handful of observations, using 1.96 instead of the correct t value materially inflates the false-positive rate. - With 200 observations, the difference between 1.972 and 1.960 changes essentially no decision, which is why analysts of large datasets often speak loosely of "the z-score" even when they computed a t statistic. Because the t-test converges to the z-test as `n` grows, and never the other way round, **t is the safe default**. You lose nothing by using it when `n` is large, and you are protected when it is small. ## A worked framing A manufacturer claims a mean battery life of 40 hours. You test 12 units and observe a sample mean of 37.5 hours with a sample standard deviation of 4 hours. Because `s` came from those 12 units, the statistic is ``` t = (37.5 - 40) / (4 / sqrt(12)) = -2.5 / 1.1547 = -2.165, df = 11 ``` and it is compared against the t distribution with 11 degrees of freedom, not the normal. Had the manufacturer instead published a certified process standard deviation of 4 hours, the same arithmetic would give a z of -2.165 compared against the normal — same number, different reference distribution, and a smaller p-value. ## What interviewers listen for Two things. First, that you locate the difference in the denominator — known `sigma` versus estimated `s` — rather than in the sample size, which is a consequence rather than the rule. Second, that you know why the rule exists: the estimated denominator adds randomness, the t distribution absorbs it through its degrees of freedom, and the correction vanishes as `n` grows.

  • Practitioners often say to use z when n is above 30 — is that rule right?
    It is a rough shortcut, not the rule. The real condition is whether the standard deviation is known or estimated; large n matters only because the t distribution converges to the normal, so the two answers stop differing. Since t is correct at every sample size and z is only ever an approximation to it, there is no reason to switch — just always use t for a mean.
  • What goes wrong if you use a normal critical value with a small sample and an estimated standard deviation?
    You reject too often. The 1.96 cutoff corresponds to a genuine 5% tail only for the normal; the t distribution with few degrees of freedom puts more mass beyond that point, so the actual false-positive rate is higher than the nominal 5%. At 4 degrees of freedom the correct two-sided cutoff is about 2.78, so 1.96 is far too permissive.
  • Why is a test of a proportion naturally a z-test rather than a t-test?
    Because the variance of a binary outcome is a function of its mean: `p * (1 - p)`. Under a specific null such as `p = 0.5`, the null value fixes the standard error at `sqrt(p0 * (1 - p0) / n)` with nothing estimated from the data, so no extra estimation uncertainty needs absorbing and the normal reference distribution applies.

saying these in an interview costs you the question

  • Says the choice depends only on whether n exceeds 30
  • Claims the z-test assumes normal data while the t-test does not
  • Uses 1.96 as the cutoff for a small-sample t-test
  • Cannot say what quantity has to be known for z to apply
  • Thinks the t-test estimates the mean while the z-test does not

context