When is a z-test valid for testing a mean instead of a t-test?
answer
- look at the denominator, not the numerator
- who supplied the standard deviation?
- estimating it adds a second source of noise
- a null on a proportion pins its variance
- the gap disappears as n grows
basics
~20 sA z-test for a mean is valid only when the population standard deviation is known rather than estimated. If you plug in the sample standard deviation, the extra uncertainty makes the statistic follow a t distribution instead.
solid answer
~50 sBoth statistics have the same shape, `(estimate - null value) / standard error`; the difference is where the standard error comes from. The z-test uses a population standard deviation you already know, so the denominator is a fixed constant and the statistic is standard normal under the null. The t-test uses the sample standard deviation `s`, which is itself a random quantity, and that extra estimation noise is what gives the t distribution with `n - 1` degrees of freedom its heavier tails. In real work the population standard deviation is almost never known, so t is the default for a mean. The honest exception is a proportion: a null such as `p = 0.5` pins the variance at `p0 * (1 - p0)`, so nothing is estimated and a z-test genuinely applies. With a few hundred observations the two agree anyway, so this matters most for small samples.
go deeper
Be ready to say that a t-test is what you use in practice because the population standard deviation is unknown, and that the t distribution carries degrees of freedom while the normal does not.
Explain the mechanism: substituting the sample standard deviation makes the denominator random, which fattens the tails, and the correction fades as degrees of freedom grow.
Show you know when the distinction stops mattering and when it bites. Be able to say why a proportion test is legitimately a z-test and why a small-sample analysis with a normal cutoff is over-rejecting.
Own the convention your organisation uses. Decide whether reporting z everywhere for large samples is acceptable simplification or a habit that will quietly mislead when someone reruns the analysis on a small segment.
## Same skeleton, different denominator Every test in this family is built the same way: ``` statistic = (observed estimate - value under the null) / (standard error of the estimate) ``` Testing a battery's mean life against a manufacturer's claim of 40 hours, the numerator is `xbar - 40`. What separates a z-test from a t-test is entirely the denominator. - **z-test for a mean:** the denominator is `sigma / sqrt(n)`, where `sigma` is the *population* standard deviation, treated as a known constant. Under the null, and given normal data (or a large enough sample), the statistic follows the standard normal distribution. - **t-test for a mean:** the denominator is `s / sqrt(n)`, where `s` is the *sample* standard deviation computed from the same data. Under the null, the statistic follows a t distribution with `n - 1` degrees of freedom. ## Why estimating the denominator changes the distribution If `sigma` is known, only the numerator varies from sample to sample, and the ratio is normal. If you substitute `s`, the denominator wobbles too. Sometimes `s` happens to come out small, and a modest deviation in the numerator is then divided by a small number, producing an unusually large statistic. That double source of randomness makes extreme values more common than the normal distribution predicts — the t distribution's heavier tails are exactly the price of not knowing `sigma`. Using a normal critical value while estimating `sigma` from a small sample therefore rejects too often: the test's real false-positive rate exceeds its nominal level. ## When is sigma actually known? Almost never for a mean measured on a fresh sample. Textbook cases where the claim is defensible: - A measuring instrument with a manufacturer-certified precision, where the measurement error is a documented property of the device rather than something you estimate. - A long-running, stable process with years of historical data, where the process standard deviation is known far more precisely than anything the current small sample could tell you. - A simulation or a designed sampling scheme where you set the noise level yourself. Even then, "known" is a modelling assumption: the historical standard deviation is still an estimate, just one based on so much data that its own uncertainty is negligible. ## The case where z is genuinely right: proportions For a proportion, the variance is a function of the mean. A binary outcome with success probability `p` has variance `p * (1 - p)`, so under a *specific* null hypothesis such as `p = 0.5`, the standard error under the null is completely determined: ``` SE_0 = sqrt(p0 * (1 - p0) / n) ``` Nothing is estimated, so a normal reference distribution is appropriate and the test is legitimately a z-test. This is why textbooks introduce z-tests for proportions and t-tests for means: it is not an arbitrary convention, it reflects which parameters the null pins down. The caveat is that the normal approximation to the binomial needs enough data — the usual rough guideline is at least about 10 expected successes and 10 expected failures. ## How much does it matter in practice? Less than the emphasis in a first course suggests, once `n` is large. The two-sided 5% critical value is 1.96 for the normal. For a t distribution it is about 2.78 at 4 degrees of freedom, about 2.09 at 20, about 2.01 at 50, and about 1.97 at 200. So: - With a handful of observations, using 1.96 instead of the correct t value materially inflates the false-positive rate. - With 200 observations, the difference between 1.972 and 1.960 changes essentially no decision, which is why analysts of large datasets often speak loosely of "the z-score" even when they computed a t statistic. Because the t-test converges to the z-test as `n` grows, and never the other way round, **t is the safe default**. You lose nothing by using it when `n` is large, and you are protected when it is small. ## A worked framing A manufacturer claims a mean battery life of 40 hours. You test 12 units and observe a sample mean of 37.5 hours with a sample standard deviation of 4 hours. Because `s` came from those 12 units, the statistic is ``` t = (37.5 - 40) / (4 / sqrt(12)) = -2.5 / 1.1547 = -2.165, df = 11 ``` and it is compared against the t distribution with 11 degrees of freedom, not the normal. Had the manufacturer instead published a certified process standard deviation of 4 hours, the same arithmetic would give a z of -2.165 compared against the normal — same number, different reference distribution, and a smaller p-value. ## What interviewers listen for Two things. First, that you locate the difference in the denominator — known `sigma` versus estimated `s` — rather than in the sample size, which is a consequence rather than the rule. Second, that you know why the rule exists: the estimated denominator adds randomness, the t distribution absorbs it through its degrees of freedom, and the correction vanishes as `n` grows.
- Practitioners often say to use z when n is above 30 — is that rule right?It is a rough shortcut, not the rule. The real condition is whether the standard deviation is known or estimated; large n matters only because the t distribution converges to the normal, so the two answers stop differing. Since t is correct at every sample size and z is only ever an approximation to it, there is no reason to switch — just always use t for a mean.
- What goes wrong if you use a normal critical value with a small sample and an estimated standard deviation?You reject too often. The 1.96 cutoff corresponds to a genuine 5% tail only for the normal; the t distribution with few degrees of freedom puts more mass beyond that point, so the actual false-positive rate is higher than the nominal 5%. At 4 degrees of freedom the correct two-sided cutoff is about 2.78, so 1.96 is far too permissive.
- Why is a test of a proportion naturally a z-test rather than a t-test?Because the variance of a binary outcome is a function of its mean: `p * (1 - p)`. Under a specific null such as `p = 0.5`, the null value fixes the standard error at `sqrt(p0 * (1 - p0) / n)` with nothing estimated from the data, so no extra estimation uncertainty needs absorbing and the normal reference distribution applies.
saying these in an interview costs you the question
- Says the choice depends only on whether n exceeds 30
- Claims the z-test assumes normal data while the t-test does not
- Uses 1.96 as the cutoff for a small-sample t-test
- Cannot say what quantity has to be known for z to apply
- Thinks the t-test estimates the mean while the z-test does not