skip to content

Sampling Distributions

How a statistic moves from sample to sample: the sampling distribution of the mean, the standard error that measures its spread, and what makes an estimator unbiased. Every p-value rests on this.

on this pageshow

explore

questions

21

What does it mean for a sample statistic to be an unbiased estimator of a population parameter?

level: juniorimportance: must knowfreq 70%

answer

  1. a property of the recipe, not one number
  2. think across repeated samples
  3. look at the centre of the sampling distribution
  4. expected value versus the true parameter
  5. zero on average, for every theta

basics

~20 s

An estimator is unbiased when its expected value across all possible samples equals the true parameter. Its estimates are centred on the target: too high as often, and by as much, as they are too low.

solid answer

~50 s

An estimator is a rule you apply to a sample, so it is itself a random quantity with a distribution over repeated samples. It is unbiased if the mean of that distribution equals the parameter it targets: `E[theta_hat] = theta`, for every possible value of theta. The sample mean is unbiased for the population mean, because `E[xbar] = (1/n) * sum E[X_i] = mu` for any sample size and any distribution with a finite mean. Two clarifications matter in an interview. First, unbiasedness is a property of the *procedure*, not of the one number you computed — a single estimate that misses by a mile is not evidence of bias. Second, unbiased does not mean accurate: an estimator can be centred on the truth and still swing wildly from sample to sample, so variance has to be judged separately.

go deeper

for a junior

Be ready to state the definition in one line and prove it for the sample mean using linearity of expectation. Say clearly that bias is about the average over repeated samples, not about any single estimate.

for a middle

Explain why unbiasedness survives shifting and rescaling but breaks under square roots and other non-linear transforms, and give a concrete unbiased-but-useless estimator to show the property is weaker than it sounds.

for a senior

Show that you judge an estimator on where it is centred and how much it moves, and that in real reporting you often accept a little bias for a large drop in variance rather than defending unbiasedness for its own sake.

for a principal

Own the framing question: what loss are we actually minimising across all the numbers the organisation publishes? Decide when a centred-but-noisy metric is worse for decisions than a slightly shrunken, stable one.

## The setup A **parameter** is a fixed but unknown number describing a whole population — the mean height of adults in a country, the true click-through rate of a page. An **estimator** is a recipe that turns a sample into a guess at that parameter: take the sample mean, take the sample median, take the largest value observed. The number that recipe produces on one particular sample is an **estimate**. Keeping the two words apart is half the battle; interviewers listen for it. Because the sample is drawn at random, the estimator is a random variable. Draw a different sample and you get a different number. The distribution of those numbers over all the samples you might have drawn is the estimator's **sampling distribution**, and every property discussed here is a property of that distribution. ## The definition Write `theta` for the parameter and `theta_hat` for the estimator. The **bias** is ``` Bias(theta_hat) = E[theta_hat] - theta ``` and the estimator is **unbiased** when this is zero — not just for one convenient value of theta, but for every value the parameter could take. In words: if you could repeat the whole study endlessly and average all the estimates, that long-run average would land exactly on the truth. ## The worked case: the sample mean Let `X_1, ..., X_n` be a random sample from a population with mean `mu`. Each observation has `E[X_i] = mu`, and expectation is linear, so ``` E[xbar] = E[(X_1 + ... + X_n)/n] = (1/n) * (mu + ... + mu) = mu ``` This holds for **any** sample size — n = 3 is as unbiased as n = 3000 — and for **any** population shape with a finite mean, skewed or not. Unbiasedness of the sample mean needs no normality assumption at all, which surprises many candidates. Similarly, the sample proportion is unbiased for the population proportion, since a proportion is just the mean of zeros and ones. ## What unbiasedness is not **It is not a statement about your one sample.** You cannot look at a single estimate and declare the estimator biased. If a sample of 30 heights averages 171 cm when the population mean is 168 cm, that is ordinary sampling variation. Bias lives in the average over the infinity of samples you did not draw. **It is not accuracy.** Consider estimating the population mean by simply reporting the first observation you happen to see. Its expected value is `mu`, so it is perfectly unbiased — and it is a terrible estimator, because it is as noisy as a single data point no matter how much data you collected. Being centred on the target says nothing about how tightly the estimates cluster around it. Spread is a separate property and is judged separately. **It is not preserved by non-linear transformation.** This trips up strong candidates. The sample variance with the n-1 divisor is unbiased for the population variance, but its square root — the sample standard deviation — is **biased low** for the population standard deviation. Taking a square root is a concave operation, and averaging then transforming is not the same as transforming then averaging. Unbiasedness survives linear rescaling and shifting; it generally does not survive squaring, taking roots, inverting, or exponentiating. **It is not the same as sampling bias.** Statistical bias is a mathematical property of an estimator given that the sample was drawn as assumed. Sampling bias is a data-collection failure — surveying only people who answer the phone at 2 p.m. The sample mean is a perfectly unbiased estimator of the mean of the population you actually sampled from; if that population is not the one you care about, no amount of estimator theory rescues you. ## Why interviewers care Unbiasedness is the first vocabulary check in any estimation discussion, and the honest follow-up is always "so is unbiased always what you want?" The mature answer is no: unbiasedness is one desirable property among several, and it is routinely traded away for a large reduction in variance. What an interviewer wants to hear is that you know the definition precisely, can prove it for the sample mean in one line, and do not confuse being centred on the truth with being close to it.

  • If the sample variance with the n-1 divisor is unbiased for the population variance, is the sample standard deviation unbiased for the population standard deviation?
    No. The square root is a concave function, so the average of the square roots is smaller than the square root of the average: the sample standard deviation is biased low. Unbiasedness is preserved by linear transformations such as shifting and rescaling, but not by non-linear ones like square roots, reciprocals or exponentials. In practice the downward bias is tiny except at very small sample sizes.
  • Can you name an unbiased estimator of the population mean that no one would ever use?
    Report the first observation and ignore the rest. Its expected value is exactly the population mean, so it is unbiased at every sample size, yet its spread never shrinks — collecting a thousand points buys nothing. It is the cleanest demonstration that unbiasedness alone is a weak requirement: it constrains where the estimates are centred, not how tightly they cluster.
  • Your survey oversamples one region and the sample mean misses the national average badly. Is the sample mean a biased estimator here?
    Not in the statistical sense. The sample mean is unbiased for the mean of the population the sample was actually drawn from, which is your oversampled one. What failed is the sampling frame, not the estimator. Fixing it means reweighting to the target population or repairing the design; swapping in a different estimator of the same wrong quantity changes nothing.

A bathroom scale that reads a random amount high or low but averages out to your true weight is unbiased. It is still useless for tracking a two-pound change, because unbiased says nothing about how much it wobbles.

saying these in an interview costs you the question

  • Says unbiased means each estimate is close to the true value
  • Claims one sample's error proves the estimator is biased
  • Thinks unbiased implies best or minimum variance
  • Assumes the sample standard deviation is unbiased for sigma
  • Confuses statistical bias with biased data collection
  • Believes the sample mean is only unbiased for normal populations

context

open as a page

How does a likelihood differ from a probability in maximum likelihood estimation?

level: juniorimportance: must knowfreq 80%

basics

~20 s

A probability treats the parameter as fixed and asks how likely the data are; a likelihood fixes the observed data and reads the same formula as a function of the parameter. Likelihood values do not sum to one over parameters.

open as a page

What does the standard error of the mean measure that the sample standard deviation does not?

level: juniorimportance: must knowfreq 82%

basics

~20 s

The standard error of the mean measures how precisely a sample mean estimates the population mean; the sample standard deviation measures how spread out individual observations are. Only the standard error shrinks as the sample grows.

open as a page

Why is the sum of squared deviations divided by n a biased estimator of the population variance?

level: middleimportance: must knowfreq 74%

basics

~20 s

Deviations are measured from the sample mean, which is pulled toward the data and makes those deviations as small as they can possibly be. Dividing their squares by n therefore underestimates the population variance by the factor (n-1)/n.

open as a page

What is Fisher information, and how does it relate to the score function of a log-likelihood?

level: middleimportance: must knowfreq 58%

basics

~20 s

The score is the derivative of the log-likelihood with respect to the parameter; Fisher information is the variance of the score at the true parameter, equivalently minus the expected second derivative. It measures how sharply data pin down the parameter.

open as a page

How do you derive the maximum likelihood estimate of a coin's heads probability from 7 heads in 10 flips?

level: middleimportance: must knowfreq 68%

basics

~10 s

Write the likelihood p^7 * (1-p)^3, take logs to get 7log(p) + 3log(1-p), differentiate and set the result to zero. That gives 7/p = 3/(1-p), so p-hat = 0.7, the sample proportion.

open as a page

Why must you quadruple the sample size to halve the standard error of a mean?

level: middleimportance: must knowfreq 70%

basics

~20 s

The standard error of a mean equals the sample standard deviation divided by the square root of the sample size, so precision improves with the square root of n. Halving it therefore requires four times the data.

open as a page

How do you get a standard error for a maximum-likelihood estimate from the log-likelihood?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Take the Hessian of the log-likelihood at the optimum, negate it to get the observed information matrix, invert it, and read standard errors as the square roots of its diagonal. Curvature at the peak is what precision means here.

open as a page

Given two unbiased estimators of the same parameter, how do you choose between them?

level: middleimportance: should knowfreq 45%

basics

~20 s

Prefer the one with smaller variance — that is efficiency. Relative efficiency is the ratio of their variances at a given sample size. Which one wins depends on the population shape, so state the assumption you are making about the data.

open as a page

How does consistency of an estimator differ from unbiasedness?

level: middleimportance: should knowfreq 56%

basics

~20 s

Unbiasedness is a finite-sample property: the estimator is centred on the parameter at every sample size. Consistency is an asymptotic one: the estimator converges to the parameter as the sample grows. Neither implies the other.

open as a page

What does the Cramer-Rao lower bound say about the variance of an unbiased estimator?

level: middleimportance: should knowfreq 45%

basics

~20 s

It sets a floor on precision: under regularity conditions any unbiased estimator of a parameter has variance at least the reciprocal of the Fisher information in the sample. No unbiased estimator can beat that floor, and some attain it exactly.

open as a page

What are the maximum likelihood estimates of the mean and variance of a Normal sample?

level: middleimportance: should knowfreq 52%

basics

~20 s

The mean estimate is the sample average xbar. The variance estimate is the average squared deviation from xbar, that is sum of (xi - xbar)^2 divided by n. Maximising the likelihood returns the divisor n, not n-1.

open as a page

How do you compute the standard error of the difference between two independent sample means?

level: middleimportance: should knowfreq 54%

basics

~10 s

Variances add, standard errors do not. The standard error of the difference between two independent sample means is sqrt(SE1^2 + SE2^2), which expands to sqrt(s1^2/n1 + s2^2/n2).

open as a page

Why does an election poll of 1,000 respondents report a margin of error near 3 points?

level: middleimportance: should knowfreq 58%

basics

~20 s

The standard error of a sample proportion is sqrt(p(1-p)/n), which at p = 0.5 and n = 1,000 is about 1.6 percentage points. A reported margin of error is conventionally about two standard errors, so roughly 3 points.

open as a page

Why can a biased estimator have lower mean squared error than an unbiased one?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Mean squared error is variance plus squared bias. Accepting a little bias can cut variance far more than the squared bias adds, so total error drops. Shrinking a noisy small-sample average toward a pooled average is the standard example.

open as a page

What does Wilks' theorem say about the likelihood-ratio statistic for two nested models?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Twice the gap in maximised log-likelihoods between a full model and a nested restricted one converges, when the restriction is true, to a chi-square distribution whose degrees of freedom equal the number of parameters the restriction removes.

open as a page

Why is minimising squared error the same as maximum likelihood under Gaussian noise?

level: seniorimportance: should knowfreq 57%

basics

~20 s

Assume the errors are independent Gaussian with constant variance. The log-likelihood then equals a constant minus the sum of squared residuals divided by twice the variance, so maximising it and minimising squared error give the same fitted parameters.

open as a page

If the MLE of a coin's heads probability is 0.7, what is the MLE of the odds p/(1-p)?

level: middleimportance: nice to knowfreq 32%

basics

~10 s

It is 0.7/0.3, about 2.33. The invariance property of maximum likelihood says the estimate of any function of a parameter is that function applied to the parameter's estimate, so no new maximisation is needed.

open as a page

How would you estimate a factory's total output from a sample of observed serial numbers?

level: seniorimportance: nice to knowfreq 24%

basics

~10 s

The largest serial number observed always underestimates the true total, since it can never exceed it. Scale it up: with n serials seen and maximum m, the estimate m*(n+1)/n - 1 is unbiased.

open as a page

When does the finite population correction meaningfully shrink a survey's standard error?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

The finite population correction, sqrt((N-n)/(N-1)), matters only when the sample is a large fraction of the population. Below about a 5 percent sampling fraction it changes the standard error negligibly; at a census it drives it to zero.

open as a page

When would you refuse to trust the asymptotic standard errors and tests from a likelihood fit?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Distrust them when the sample is small relative to the number of parameters, when an estimate sits on a boundary of its parameter space, when the model is barely identified, or when the model is misspecified. Each breaks a precondition behind the asymptotics.

open as a page