Why must you quadruple the sample size to halve the standard error of a mean?
answer
- diminishing returns on extra data
- variance of the mean divides by n
- take the square root of that variance
- shrinking by k costs k-squared observations
basics
~20 sThe standard error of a mean equals the sample standard deviation divided by the square root of the sample size, so precision improves with the square root of n. Halving it therefore requires four times the data.
solid answer
~50 sBecause precision is bought in square roots. For independent observations the variance of the sample mean is `sigma^2 / n`, so its standard deviation - the standard error - is `sigma / sqrt(n)`. Only `sqrt(n)` is in the denominator, so to divide the standard error by 2 you must multiply `n` by 4; to divide it by 10 you need 100 times the data. Concretely, exam scores with a spread of 15 points give a standard error of 1.5 at `n = 100`, 0.75 at `n = 400`, and 0.375 at `n = 1,600`. The practical consequence is brutal diminishing returns: the first thousand observations buy most of the precision you will ever get, and the last increments cost quadratically. It also means extra data is the wrong lever for a biased sample, since bias does not shrink with `n` at all.
code
python · 15 linesimport random, statistics
SIGMA = 15.0
def observed_se(n, trials=4000):
means = [statistics.fmean([random.gauss(100.0, SIGMA) for _ in range(n)])
for _ in range(trials)]
return statistics.stdev(means)
for n in (25, 100, 400):
print(n, round(observed_se(n), 3), round(SIGMA / n ** 0.5, 3))
# 25 2.98 3.0
# 100 1.51 1.5
# 400 0.76 0.75go deeper
Memorise the shape: standard error is the spread over the square root of the sample size, so four times the data gives twice the precision. Be able to plug numbers into it without stumbling over the square root.
Derive it. Show that variances of independent observations add, so the mean has variance sigma squared over n, and explain that the square root is what converts a linear variance gain into a square-root precision gain.
Bring the operational consequences: diminishing returns on data collection, clustered observations that break independence and inflate effective error, and the fact that bias is untouched by sample size. Suggest variance-reduction levers instead of brute-force sampling.
Own the tradeoff between measurement cost and decision value. Argue when precision has stopped being worth its quadratic price, and set expectations that a request for ten times tighter numbers is a hundredfold increase in data, not a scheduling detail.
## The rule For `n` independent observations from a population with standard deviation `sigma`, the standard error of the sample mean is ``` SE = sigma / sqrt(n) ``` Estimated from data it is `s / sqrt(n)`. Precision therefore scales as `1 / sqrt(n)`: **to reduce the standard error by a factor of k, you need k-squared times as many observations.** | target improvement | data multiplier | |---|---| | 1.4x tighter | 2x | | 2x tighter | 4x | | 3x tighter | 9x | | 10x tighter | 100x | ## Why the square root, not n The derivation is two lines and worth being able to produce on a whiteboard. Write the sample mean as `(X1 + ... + Xn) / n`. For independent variables, variances of a sum add: ``` Var(X1 + ... + Xn) = n * sigma^2 Var(mean) = (1 / n^2) * n * sigma^2 = sigma^2 / n SE = sqrt(sigma^2 / n) = sigma / sqrt(n) ``` The **variance** falls linearly in `n`. The standard error is the square root of the variance, and taking that square root is what turns a linear gain into a square-root gain. Candidates who say 'doubling the sample halves the standard error' are quoting the variance rule and forgetting the square root. Intuitively, averaging cancels independent noise: some observations sit above the population mean, some below, and their errors partly annihilate. The cancellation is only partial - random errors add in quadrature, not in absolute value - which is exactly where the square root comes from. ## A worked example A school measures exam scores with a population spread of about 15 points. - `n = 100`: `SE = 15 / 10 = 1.5` points. - `n = 400`: `SE = 15 / 20 = 0.75` points. - `n = 1,600`: `SE = 15 / 40 = 0.375` points. Going from 100 to 400 students - tripling the data collection effort - buys 0.75 points of precision. Going from 400 to 1,600 buys only 0.375 more. The curve flattens fast, and the third quadrupling costs twelve hundred more students for a gain most decisions cannot even use. ## Seeing it rather than believing it The formula is a claim about the spread of the sample mean across hypothetical repeated samples, so the honest check is to repeat the sampling. Draw many samples of size `n`, record each sample's mean, and take the standard deviation of those means: it should track `sigma / sqrt(n)` closely, and it should visibly halve each time `n` quadruples. Doing this once removes most of the mystery from the formula. ## What the rule assumes `sigma / sqrt(n)` is not a law of nature; it follows from assumptions that fail in recognisable ways. 1. **Independence.** If observations come in correlated groups - several sessions from the same user, several employees from the same office - the errors no longer cancel as fast. The effective sample size can be a small fraction of the nominal one, and a naive standard error is too small, sometimes by a large multiple. Adding more rows from the same handful of clusters buys much less than the square-root rule promises. 2. **Identically distributed data from the population you care about.** If the sampling frame excludes part of the population, more data narrows the standard error around **the wrong number**. Precision improves while accuracy does not. 3. **Finite variance.** The derivation uses `sigma^2`; if the population has extremely heavy tails so that the variance is not finite, the formula has nothing to compute. 4. **A single mean, not a difference or a ratio.** Other statistics have their own standard error formulas, though most inherit some version of the square-root decay. ## Bias versus variance of the estimate The most valuable point to make in an interview: **the square-root rule governs random error only.** A systematic error - a broken instrument, a survey that only reaches happy customers, a mis-specified metric - is unaffected by sample size. With enough data such a study reports a very precise wrong answer, and its tight error bars actively mislead. When you see a request for 'more data' aimed at a suspicious result, the first question is whether the problem is noise or bias. ## Turning the rule around Used forwards, the rule sizes studies: pick the precision you need, solve `n = (sigma / SE_target)^2`. Wanting a standard error of 0.5 points on the exam-score example gives `n = (15 / 0.5)^2 = 900`. Used backwards, it is a sanity check on someone else's claim - a reported standard error implies a sample size and a spread, and if that trio is implausible, something in the analysis is wrong.
- Your stakeholder wants the standard error of an average cut by a factor of ten - what does that cost?One hundred times the data. Since `SE = sigma / sqrt(n)`, reducing it tenfold requires `sqrt(n)` to grow tenfold, so `n` grows by a factor of 100. That is usually the moment to ask whether cheaper levers exist: reducing `sigma` through better measurement, stratifying, or controlling for a covariate, all of which shrink the numerator instead of inflating the denominator.
- When does collecting more data stop reducing the standard error as one over root n?When the observations are not independent. Repeated measurements on the same users, or respondents clustered within a few sites, make errors correlate, so extra rows carry less new information and the effective sample size lags the nominal one. The rule also gives no protection against bias: a systematic error stays exactly the same size no matter how much data you gather.
- How would you use the square-root rule to size a study rather than to critique one?Invert it. Fix the standard error you can act on, estimate the spread from pilot data or history, and solve `n = (sigma / SE_target)^2`. With a spread of 15 points and a target standard error of 0.5, that is 900 observations. Sanity-check the answer against budget before committing, since the cost grows quadratically in the precision demanded.
Precision is priced like the area of a square photograph: to double the sharpness of the picture you have to buy four times the pixels.
saying these in an interview costs you the question
- Says doubling the sample size halves the standard error
- Treats the standard error as falling linearly in n
- Believes more data can fix a biased sampling frame
- Applies the square-root rule to clustered or repeated measurements
- Confuses the shrinking standard error with a shrinking data spread