skip to content

What is the difference between the weak and strong laws of large numbers?

level: seniorimportance: nice to knowfreq 24%

answer

  1. one sample size versus one whole trajectory
  2. rare spikes may continue forever
  3. almost sure is the stronger mode
  4. finite mean suffices for both
  5. probability one that the sequence settles

basics

~20 s

The weak law is convergence in probability: at each large sample size, a big miss is unlikely. The strong law is almost sure convergence: with probability one the sequence of running averages settles at the population mean and stays.

solid answer

~50 s

The weak law states convergence in probability: for every tolerance `eps > 0`, `P(|Xbar_n - mu| > eps)` goes to zero as `n` grows. It is a statement about each fixed `n` separately — take a large sample and you are probably close — and it leaves room for a single infinite run of averages to keep producing rare excursions forever. The strong law states almost sure convergence: `P(lim Xbar_n = mu) = 1`. It is a statement about whole trajectories — with probability one, the sequence of running averages eventually enters any band around `mu` and never leaves. Almost sure convergence implies convergence in probability, so the strong law is the stronger claim. For independent identically distributed draws a finite expected value suffices for both; finite variance appears only in the elementary Chebyshev proof of the weak law.

go deeper

for a junior

Know that there are two versions and that both conclude the sample mean approaches the population mean. Being able to say the strong one is a stronger claim is enough at this level.

for a middle

Be able to state each mode in symbols and explain the difference in plain words: probability of a miss at each fixed sample size, versus probability one that the whole sequence settles.

for a senior

Demonstrate that you know which implies which and which assumptions each needs, and that you can produce a sequence converging in probability but not almost surely rather than only naming the modes.

for a principal

Own the practical judgment about when trajectory-level convergence matters — long-running simulations and iterative estimators watched as a single run — and when the weak law is the honest description of a one-shot dataset.

## Two modes of convergence Both laws conclude that the sample mean `Xbar_n` of independent identically distributed draws approaches the population mean `mu`. They differ in what kind of approaching they assert. **Weak law — convergence in probability.** For every `eps > 0`: ``` P(|Xbar_n - mu| > eps) -> 0 as n -> infinity ``` Read this one `n` at a time. Fix a large sample size, run the experiment, and the chance of missing `mu` by more than `eps` is small. Nothing is asserted about how one particular infinite sequence of averages behaves over time. **Strong law — almost sure convergence.** ``` P( lim_{n -> infinity} Xbar_n = mu ) = 1 ``` Read this one *trajectory* at a time. Imagine writing down the whole infinite sequence `Xbar_1, Xbar_2, Xbar_3, ...` for a single run. With probability one, that sequence converges to `mu` in the ordinary calculus sense: for any band around `mu`, the sequence enters it at some point and never leaves again. ## Why the distinction is not pedantic Convergence in probability permits a sequence to keep failing, provided the failures become rare enough. It says the *chance* of a spike at time `n` shrinks; it does not say the spikes stop. Almost sure convergence forbids infinitely many spikes: on all but a probability-zero set of runs, only finitely many terms are far from `mu`. The canonical example of the gap is a sequence of indicator variables built from sweeping windows on the interval from 0 to 1. The first window is the whole interval; then two half-length windows sweep across it; then four quarter-length windows; then eight, and so on. Each variable equals 1 when a uniformly chosen point lands in the current window and 0 otherwise. The probability of a 1 is the window width, which goes to zero, so the sequence converges to 0 *in probability*. Yet the windows keep sweeping across the entire interval forever, so every point is inside a window infinitely often: no individual trajectory ever settles at 0. This sequence converges in probability but not almost surely — the exact behaviour the weak law tolerates and the strong law rules out. ## Which implies which Almost sure convergence implies convergence in probability. The strong law is therefore the stronger statement, and it implies the weak law. The reverse is false in general: convergence in probability does not imply almost sure convergence, as the sweeping-window sequence shows. ## Assumptions For independent, identically distributed draws: - The **strong law** holds whenever the expected value is finite. No variance assumption is needed. - The **weak law** holds under a finite expected value too, and follows immediately from the strong law. - The **elementary proof** of the weak law adds an assumption of finite variance, because it goes through Chebyshev's inequality: the sample mean has variance `sigma^2 / n`, so `P(|Xbar_n - mu| >= eps) <= sigma^2 / (n * eps^2)`, which vanishes as `n` grows. That extra assumption is an artefact of the proof technique, not a requirement of the result. A frequent interview stumble is to claim the weak law needs finite variance and the strong law needs more. It is the other way round in spirit: the strong law is the stronger conclusion, reached for independent identically distributed data under the same mild finite-mean condition, and the variance appears only to make an easy proof available. ## Does the distinction matter in practice? For most applied work, no — you have one finite dataset, not an infinite trajectory, and the weak law is the statement that actually describes your situation: with this much data, a large miss is unlikely. The distinction earns its keep in three places. 1. **Simulation and iterative estimation.** When you run one long simulation and watch a running estimate, you are looking at a single trajectory. The strong law is what justifies believing that *this* run settles, rather than merely that any given checkpoint is probably fine. 2. **Theory.** Proofs of consistency for estimators and algorithms often need trajectory-level convergence, and citing the wrong mode is a genuine error. 3. **Diagnosing conditions.** Asking which law applies forces you to ask whether the mean is finite and whether the draws are independent — the questions that actually determine whether anything converges. ## How to answer it State both modes precisely, note that almost sure implies in probability so the strong law is stronger, give the intuition (each fixed `n` is probably fine, versus each whole run eventually settles), and mention that for independent identically distributed draws a finite mean is enough for both. If you can describe a sequence that converges in probability but not almost surely, you have demonstrated you understand the difference rather than reciting labels.

  • Can you describe a sequence that converges in probability but not almost surely?
    Take indicator variables for sweeping windows on the unit interval: the whole interval, then two halves, then four quarters, and so on, each indicating whether a fixed random point lies in the current window. The probability of a 1 shrinks to zero, so it converges in probability, but the windows keep sweeping forever, so every point is hit infinitely often and no trajectory settles.
  • Does the strong law of large numbers require a finite variance?
    No. For independent, identically distributed draws a finite expected value is enough for the strong law. Finite variance appears only in the elementary Chebyshev-based proof of the weak law, where it makes the argument short. Requiring it for the strong law is a common misstatement.
  • Which mode of convergence matters when you watch a running estimate in one long simulation?
    Almost sure convergence, because you are observing a single trajectory rather than resampling at a fixed sample size. The strong law is what justifies believing that this particular run eventually settles and stays settled, rather than only that any given checkpoint is probably close.

Convergence in probability says that at any given moment almost no streetlight in the city is flickering. Almost sure convergence says that each individual streetlight eventually stops flickering for good.

saying these in an interview costs you the question

  • Says the weak law needs finite variance and the strong law needs more
  • Claims convergence in probability implies almost sure convergence
  • Treats the two laws as interchangeable labels
  • Cannot state either mode of convergence precisely
  • Believes almost sure convergence means guaranteed for every outcome

context