skip to content

Why does the Poisson distribution have variance equal to its mean?

level: middleimportance: should knowfreq 52%

answer

  1. one parameter, not two
  2. spread is sqrt of the mean
  3. binomial variance with p tiny
  4. variance-to-mean ratio of one

basics

~20 s

A Poisson count has one parameter λ that fixes both its mean and its variance, so its standard deviation is sqrt(λ). Spread is not free to differ from level, and a variance-to-mean ratio far from 1 signals a wrong model.

solid answer

~40 s

The Poisson family has one parameter, λ, and both E[X] = λ and Var(X) = λ. You can see it directly: E[X] = λ and E[X(X-1)] = λ², so E[X²] = λ² + λ and Var(X) = λ² + λ - λ² = λ. You can also see it as the limit of a binomial: variance np(1-p) collapses to np as p becomes tiny, while the mean stays np. Practically this equidispersion is a **testable** property. Compute the sample mean and sample variance of your counts; a variance-to-mean ratio near 1 is consistent with Poisson, while a ratio of 4 or 5 says the counts are overdispersed — usually clustering or a rate that varies across the observed units. Then Poisson intervals and tail probabilities are far too narrow.

go deeper

for a junior

Remember the headline fact: for a Poisson count the mean and the variance are both lambda, so the standard deviation is the square root of the mean. Be able to state it and apply it to a number.

for a middle

Explain why it holds, either from the second factorial moment or from the binomial variance np(1-p) losing its (1-p) factor as p shrinks. Be able to use the variance-to-mean ratio as a quick fit check.

for a senior

Diagnose real count data with it. Recognise overdispersion from clustering or pooled heterogeneous rates, and state the consequence concretely: intervals too narrow, alert thresholds firing too often, extreme days wrongly flagged.

for a principal

Own the decision about what to do when the check fails: whether to disaggregate the pooled units, adopt a two-parameter count family, or accept the simpler model with documented caveats, and be explicit about which downstream decisions the understated spread would corrupt.

## The claim For X ~ Poisson(λ), both `E[X] = λ` and `Var(X) = λ`. The standard deviation is therefore `sqrt(λ)`. This property is called **equidispersion**, and it is a consequence of the family having only one parameter: unlike the normal family, where the centre and the spread are set independently, a Poisson count that averages 100 *must* have a standard deviation of 10. ## Derivation from the mass function Start with `P(X = k) = e^(-λ) λ^k / k!`. The mean: `E[X] = Σ_{k≥0} k · e^(-λ) λ^k / k!`. The k = 0 term vanishes, and for k ≥ 1 the k cancels one factorial factor, leaving `λ · Σ_{k≥1} e^(-λ) λ^(k-1)/(k-1)! = λ · 1 = λ`, because the remaining sum is the full Poisson mass function summing to 1. The same trick twice gives the second factorial moment: `E[X(X-1)] = Σ_{k≥2} k(k-1) · e^(-λ) λ^k / k! = λ² · Σ_{k≥2} e^(-λ) λ^(k-2)/(k-2)! = λ²`. So `E[X²] = E[X(X-1)] + E[X] = λ² + λ`, and `Var(X) = E[X²] - (E[X])² = λ² + λ - λ² = λ`. ## Derivation from the binomial limit A second, more intuitive route: a binomial count has mean np and variance np(1-p). Hold np = λ fixed while letting p shrink toward 0. The mean stays λ, and the variance np(1-p) → np = λ because the (1-p) factor tends to 1. The binomial's variance is always *below* its mean by the factor (1-p); the Poisson is the boundary case where that discount disappears. This also explains why a Poisson count never has variance below its mean while a binomial always does. ## What equidispersion buys you It gives a free diagnostic. The **variance-to-mean ratio**, sometimes called the index of dispersion, is 1 for a Poisson count. Compute the sample mean and the sample variance of your observed counts and compare: - **Ratio near 1** — consistent with Poisson; the model's spread is credible. - **Ratio well above 1 (overdispersion)** — the counts vary more than a single rate can explain. The usual causes are clustering (one event triggers others, so arrivals come in bursts) or heterogeneity (the underlying rate itself differs across the days, users or regions being pooled, adding rate-variation on top of the Poisson noise). Consequence: Poisson-based intervals are too narrow and extreme days look far more surprising than they are. - **Ratio well below 1 (underdispersion)** — the counts are *more regular* than random arrivals would be. This happens with scheduled or capacity-capped events: a queue that processes at most 10 items an hour cannot produce a Poisson tail. Here Poisson overstates the variability. With counts averaging 20 per day and a sample variance of 95, the ratio is about 4.75 — decisively not Poisson. The right reading is not "the mean is wrong" but "one rate does not describe these days". ## A caution on the diagnostic Sample variance is itself noisy. With only 10 or 15 observed windows, a ratio of 1.8 is unremarkable; you need either many windows or a large ratio before calling overdispersion. Also compare the mean and variance on the **same** window definition — mixing hourly counts with daily counts guarantees nonsense. ## Relative variability shrinks as λ grows Because the standard deviation is sqrt(λ) while the mean is λ, the coefficient of variation is `sqrt(λ)/λ = 1/sqrt(λ)`. A process averaging 4 events has a standard deviation of 2 — a 50 percent relative swing. One averaging 400 has a standard deviation of 20, only 5 percent. This is why low-count segments look wildly unstable while high-count aggregates look smooth, even when the underlying rate is identical. Ranking small regions by raw event counts is treacherous for exactly this reason: the noisiest-looking ones are usually just the smallest. ## When you need variance above the mean If counts really are overdispersed, a one-parameter family cannot represent them. A two-parameter discrete family such as the negative binomial has variance strictly greater than its mean and can absorb the extra spread, which is why it is the standard fallback whenever the variance-to-mean check fails. ## Common errors Saying the variance is λ² (that is the square of the mean, not the variance), claiming the standard deviation is λ, believing the mean and variance can be tuned separately, or observing variance ≫ mean and continuing to use Poisson because "counts are always Poisson".

  • Daily signups have a sample mean of 20 and a sample variance of 95 — what do you conclude?
    The variance-to-mean ratio is about 4.75, far above the 1 a Poisson model requires, so these counts are overdispersed. Likely causes are clustering or a signup rate that genuinely differs day to day. Any Poisson-based interval or tail probability computed here will be far too narrow.
  • What does a variance-to-mean ratio well below 1 suggest?
    Underdispersion — the counts are more regular than independent random arrivals would produce. Typical causes are scheduled events or a hard capacity cap that truncates busy periods. A Poisson model would then overstate variability and make your planning unnecessarily conservative.
  • How does relative variability change as the Poisson mean grows?
    The standard deviation is sqrt(λ) while the mean is λ, so the coefficient of variation is 1/sqrt(λ) and shrinks as λ grows. Small-count segments therefore look far noisier than large ones even at an identical underlying rate — a trap when ranking small regions by raw counts.

Most distributions let you set the centre and the width with separate dials. Poisson has a single dial: turn up the average and the spread follows automatically, at the square root of the setting.

saying these in an interview costs you the question

  • Says the variance is the mean squared
  • Claims the standard deviation equals lambda
  • Thinks mean and variance are separate Poisson parameters
  • Keeps using Poisson when variance far exceeds the mean
  • Calls overdispersion a sign the mean was mis-estimated

context