What does the law of large numbers say about the sample mean as observations accumulate?
answer
- averages settle, individual draws do not
- independent draws from one distribution
- the expected value must be finite
- running average approaches the true mean
- convergence, never exact equality
basics
~20 sThe law of large numbers says that for independent draws from one distribution with a finite mean, the average of the observations converges to that distribution's true expected value as the number of draws grows.
solid answer
~40 sTake draws `X1, X2, ...` that are independent and identically distributed, from a distribution whose expected value `mu` is finite. The law of large numbers says the running average `Xbar_n = (X1 + ... + Xn) / n` converges to `mu` as `n` grows. That is what licenses estimating anything by averaging: roll a fair die many times and the running average of the faces closes in on 3.5. Note carefully what it does not claim. It says nothing about any individual draw, it never promises exact equality with `mu`, and it converges to the mean of whatever distribution you actually sampled — if your sampling is biased, more data walks you confidently to the wrong number.
go deeper
Be ready to state it in one clean sentence: independent draws from the same distribution with a finite mean, and the running average approaches that mean. Have a concrete example ready, such as the average die face closing on 3.5.
Explain the mechanism, not just the claim. An interviewer expects you to say why averaging cancels independent deviations, and to list the three conditions — independence, identical distribution, finite mean — and what breaks when each fails.
Show that you know the law converges to the mean of the distribution you actually sampled. The judgment interviewers look for is recognising that a biased collection process makes more data more confidently wrong, not more accurate.
Own the framing that the law justifies measurement strategy but sets no deadline. Be ready to argue how much data is enough for a given decision, given the skew and variability of the quantity, and when to buy precision versus fix the sampling frame instead.
## The statement Let `X1, X2, X3, ...` be independent random variables all drawn from the same distribution, and suppose that distribution has a finite expected value `mu = E[X]`. Define the sample mean after `n` draws: ``` Xbar_n = (X1 + X2 + ... + Xn) / n ``` The law of large numbers states that `Xbar_n` converges to `mu` as `n` grows without bound. Two things in that sentence deserve emphasis. **`mu` is a parameter; `Xbar_n` is an estimate.** `mu` is a fixed number attached to the distribution — unknown, but not random. `Xbar_n` is a random quantity: run the experiment again and you get a different value. The law is a statement about how the random thing behaves relative to the fixed thing. **Convergence, not arrival.** The law describes a limit. For any finite `n`, `Xbar_n` is almost certainly not exactly `mu` — for a continuous distribution the probability of exact equality is zero. What shrinks is the chance of being *far* off. ## The three conditions 1. **Independence.** Each draw must carry information the others do not. If every observation is driven by one shared shock, you effectively have far fewer observations than you counted, and the average need not settle at all. 2. **Identically distributed.** All draws come from the same distribution. If the underlying process drifts — the population you are sampling changes over the collection window — there may be no single `mu` to converge to. 3. **A finite mean.** The expected value must exist and be finite. Heavy-tailed quantities whose expectation is infinite have no `mu` for the average to approach; the running average keeps being dragged upward by rare enormous values instead of settling. Finite variance is a convenience, not a requirement: it makes the elementary proof easy, but the strong law for independent identically distributed draws needs only a finite mean. ## Why averaging works Each draw deviates from `mu` by some amount. Because the deviations are independent and have mean zero, positive and negative departures partially cancel when you add them up, and dividing by `n` shrinks whatever is left. The sum of the deviations does not shrink — it typically grows — but it grows far more slowly than `n` does, so the *per-observation* error is squeezed to nothing. This is the whole mechanism, and it is worth internalising because it explains the law's most-misunderstood consequence: averages settle down, totals do not. ## Probabilities are averages too A probability is the expected value of an indicator. Define `Y = 1` when an event happens and `Y = 0` when it does not; then `E[Y]` is exactly `P(event)`, and the sample mean of the `Y` values is the observed relative frequency. So the law of large numbers is also the bridge between probability as a mathematical object and probability as an observed long-run frequency. Every time you estimate a rate — a conversion rate, a defect rate, a win rate — by counting successes and dividing by trials, this law is the guarantee you are leaning on. ## What the law is not - **It is not a self-correcting force.** A run of unusual results is not pushed back by future results; it is diluted by them. Nothing in the law reaches back to fix what already happened. - **It is not about individual outcomes.** Knowing the average of a million rolls tells you nothing extra about roll number 1,000,001. - **It does not fix bias.** The average converges to the mean of the distribution you actually sampled. Sample only weekday users and you converge, with beautiful precision, to the weekday mean — the extra data increases your confidence without improving your accuracy about the target population. - **It gives no deadline.** The law is asymptotic: it says the limit is `mu`, not that `n = 1,000` is enough. How large is large enough depends on how variable and how skewed the underlying distribution is. ## In interviews The crispest way to answer is to state the setup (independent, identically distributed, finite mean), state the conclusion (the sample mean converges to the expected value), and then immediately name one thing it does not say. Candidates who only recite the first half sound like they memorised a formula; candidates who add "and it converges to the mean of whatever you sampled, so it does not rescue a biased sample" sound like they have used it.
- What must be finite for the law of large numbers to hold?The expected value. If the distribution's mean is infinite or undefined — a very heavy right tail, for instance — there is no target for the running average to approach, and it keeps being pulled by rare enormous values. Finite variance is not required for the strong law with independent identically distributed draws; it only makes the elementary proof convenient.
- Does the law of large numbers apply when the observations are dependent?Not in the form usually quoted, which assumes independence. Averages of dependent sequences can still converge under extra conditions when the dependence is weak and short-ranged, but if every observation shares one common driver, your effective number of independent observations barely grows with your row count and the average may not settle at all.
- Does the law promise the sample mean will eventually equal the population mean exactly?No. It promises convergence, not equality. For a continuous distribution the probability that the sample mean lands exactly on the population mean is zero at every sample size. The correct statement is that the probability of being further than any fixed distance from the mean goes to zero as the sample grows.
saying these in an interview costs you the question
- Says the sample mean eventually equals the population mean exactly
- Claims the law forces individual outcomes to even out
- Thinks the law requires a normal distribution
- Believes more data cures a biased sampling process
- Applies it to strongly dependent observations without comment
- Confuses the fixed parameter mu with the random sample mean