How do you standardise a sample mean into a z-score using the Central Limit Theorem?
answer
- recentre, then rescale
- units of the average's own wobble
- the denominator shrinks with sample size
- square root of n, not n
- subtract mu, divide by sigma over sqrt(n)
basics
~20 sSubtract the population mean from the sample mean, then divide by sigma / sqrt(n), where sigma is the population standard deviation and n the sample size. The Central Limit Theorem says that ratio is approximately standard normal.
solid answer
~50 sThe standardised form of the CLT is `z = (Xbar - mu) / (sigma / sqrt(n))`, where `Xbar` is the observed sample average, `mu` the population mean, `sigma` the population standard deviation and `n` the sample size. The theorem says this z is approximately standard normal for large n, so any probability question about the sample average becomes a lookup on the standard normal curve. Worked example: a population has `mu = 100` and `sigma = 20`, and you take n = 100 draws. The scale for the average is `20 / sqrt(100) = 2`, so an observed average of 104 gives `z = (104 - 100) / 2 = 2`, and the chance of seeing an average that high or higher is about 0.023. The two mistakes to avoid are dividing by `sigma` instead of `sigma / sqrt(n)`, and dividing by `n` instead of `sqrt(n)`.
go deeper
Memorise the formula and the order of operations: subtract the population mean first, then divide by sigma / sqrt(n). Practise until the square root never goes missing.
Be ready to derive the denominator, not just recite it: variances of independent draws add, so the average has variance sigma^2 / n. Then walk a numeric example end to end.
Demonstrate judgment about when the normal lookup is trustworthy at all, and say out loud which assumptions — independence, known sigma, adequate n for the skew — the number is resting on.
Own the failure mode at scale: standardised scores get automated into dashboards and alerts. Be able to argue where that automation is safe and where it manufactures false alarms.
## The transformation The Central Limit Theorem says that for i.i.d. draws with population mean `mu` and finite population standard deviation `sigma`, the sample average `Xbar` of n draws is approximately normal with mean `mu` and standard deviation `sigma / sqrt(n)`. Standardising means shifting and rescaling that quantity so it lands on the standard normal curve, the one with mean 0 and standard deviation 1: `z = (Xbar - mu) / (sigma / sqrt(n))` Equivalently, `z = sqrt(n) * (Xbar - mu) / sigma`. Both forms are the same algebra; the second makes the `sqrt(n)` growth explicit. ## Why this shape and not another Subtracting `mu` recentres the average on zero, so z measures a distance from the population mean rather than an absolute level. Dividing by `sigma / sqrt(n)` converts that distance into *units of the average's own variability*. A z of 2 means the observed average sits two of its own typical wobbles above the population mean, regardless of whether the metric is measured in seconds, dollars or clicks. That unit-free property is the point: it lets one table of normal probabilities answer every question. The `sqrt(n)` in the denominator is where candidates slip. The variance of an average of n independent draws is `sigma^2 / n` — variances of independent variables add, and dividing the sum by n divides the variance by `n^2`, leaving `sigma^2 / n`. Taking a square root gives `sigma / sqrt(n)` as the standard deviation, not `sigma / n`. Dividing by n instead of `sqrt(n)` inflates z by a factor of `sqrt(n)` and produces absurdly extreme probabilities. ## Worked example Suppose a population has mean `mu = 100` and standard deviation `sigma = 20`, and you draw n = 100 independent observations. 1. Scale of the average: `sigma / sqrt(n) = 20 / 10 = 2`. 2. Observed average `Xbar = 104`. 3. `z = (104 - 100) / 2 = 2.0`. 4. Upper-tail probability: `P(Z > 2) ~ 0.0228`, roughly a 2.3% chance. So, if the population really is centred at 100, seeing an average of 104 or higher from 100 draws happens about one time in forty. Notice how much the sample size does: a *single* observation of 104 would be only `(104 - 100) / 20 = 0.2` standard deviations out, entirely unremarkable. Averaging 100 draws sharpens the same 4-unit gap into a two-sigma event. ## Direction and tails Be careful which tail you want. `P(Xbar > 104)` maps to `P(Z > 2)`. `P(Xbar < 96)` maps to `P(Z < -2)`, the same 0.0228 by symmetry. A two-sided question — how unusual is an average at least 4 units away from 100 in either direction — is `P(|Z| > 2) ~ 0.0455`. Getting the direction of the inequality right after the transformation is half the marks on this question. ## Sanity checks worth saying out loud **Sign.** If `Xbar` is above `mu`, z is positive; below, negative. If your arithmetic produces a negative z for an average that exceeds the mean, you swapped the subtraction. **Magnitude.** z-values from real samples are usually within a few units of zero. A z of 40 almost always means you divided by `sigma / n` or forgot the square root entirely. **Units.** `Xbar`, `mu` and `sigma` must all be in the same units for the ratio to be dimensionless. ## The assumptions you are leaning on The standardisation itself is just algebra; the *normal probability lookup* is what requires the CLT, and the CLT requires independent identically distributed draws with finite variance and an n large enough for the approximation to bite. If the parent distribution is strongly skewed, n = 100 may be nowhere near enough and the tail probability you read off the normal curve can be badly wrong — usually wrong in the tail you care about most. This worked form also assumes `sigma` is a known population value. In practice you often only have an estimate computed from the same sample, and the correct reference curve then changes; that adjustment is a separate topic, but flagging that you know the difference between a known population `sigma` and an estimated one is exactly the nuance an interviewer listens for.
- What goes wrong if you divide by n instead of sqrt(n)?The z-score is inflated by a factor of `sqrt(n)`, so with n = 100 every result looks ten times more extreme than it is. Perfectly ordinary averages come out as impossible tail events. The correct scale comes from `Var(Xbar) = sigma^2 / n`, whose square root is `sigma / sqrt(n)`.
- Does a large z-score prove the sample came from a different population?No. A large z says the observed average would be unusual *if* the assumed mean and standard deviation were the true ones. It could equally signal a broken assumption: dependent draws, a non-representative sample, a mis-specified sigma, or a parent distribution too skewed for the normal approximation at that n.
- How does the required z change if the sample size quadruples?The denominator halves, because `sqrt(4n) = 2 * sqrt(n)`, so the same gap between `Xbar` and `mu` produces a z-score twice as large. Quadrupling the data doubles the sensitivity to a fixed-size difference — the familiar reason sample-size increases show diminishing returns.
saying these in an interview costs you the question
- Divides by sigma instead of sigma over sqrt(n)
- Divides by n rather than the square root of n
- Reverses the subtraction and reports the wrong sign
- Reads a one-tailed probability when the question is two-tailed
- Applies the normal lookup to strongly skewed data at small n