Page views per session have variance about eight times the mean. What breaks in a Poisson regression?
answer
- the variance assumption, not the link
- estimates survive, uncertainty does not
- deviance far above its degrees of freedom
- an extra dispersion parameter
- quadratic versus linear variance growth
basics
~20 sThat is overdispersion. The Poisson likelihood assumes the variance matches the mean, so it understates uncertainty: coefficient estimates stay roughly right, but standard errors come out far too small, intervals too narrow and significance too easy to claim.
solid answer
~50 sA variance-to-mean ratio near 8 means the variance assumption behind the Poisson likelihood is badly violated. The damage is specific and worth stating precisely: provided the mean model — log link, sensible predictors, correct offset — is roughly right, the coefficient estimates remain consistent, but the **standard errors** take the hit. They are computed from the assumed variance, so they come out several times too small, confidence intervals too narrow, and marginal predictors look convincingly significant. Typical diagnostics are a residual deviance far above its residual degrees of freedom and a fitted-versus-residual plot that fans out. The fixes are a quasi-Poisson fit, which keeps the same mean model and scales standard errors by the square root of an estimated dispersion, or a negative binomial model, which adds a dispersion parameter so variance grows as `mu + mu^2/theta` and is estimated by maximum likelihood.
go deeper
Know the phrase and the basic fact: the Poisson model assumes variance equals the mean, and real count data are often far more spread out than that, which makes the reported precision too good to be true.
Explain the mechanics of both the detection and the repair: deviance against residual degrees of freedom, variance scaled against fitted means, and how quasi-Poisson inflates standard errors while negative binomial adds a dispersion parameter to the likelihood.
Show that you separate the symptom from the cause. Say plainly that estimates stay consistent while standard errors collapse, then diagnose whether the excess variance is clustering, an omitted predictor, or genuine heterogeneity before choosing a remedy.
Own the consequences for decisions made on these models. Understated uncertainty in a count model quietly manufactures significant findings, so set the expectation that dispersion is checked and reported as a matter of course, not discovered when a result fails to replicate.
## What overdispersion is The Poisson likelihood carries a hard constraint: given the predictors, the conditional variance of the outcome equals its conditional mean. Overdispersion is the empirical finding that the data vary far more than that. A variance-to-mean ratio around 8 on page views per session is not a marginal violation; it says the data are roughly eight times as variable as the model believes. ## Where it comes from Overdispersion is almost never a defect of the counting itself. It usually signals structure the model has not captured: - **Unobserved heterogeneity.** Sessions belong to users with genuinely different browsing intensity. Even if each user's counts behaved perfectly, mixing users with different rates inflates variance across the pooled data. - **Clustering and repeated measurement.** Many sessions come from the same user, so observations are not independent; correlated observations behave like fewer effective observations with more spread. - **Omitted or mis-specified predictors.** A predictor that strongly separates high- from low-activity sessions, left out, shows up as extra variance. - **Contagion within a unit.** One page view makes the next more likely, breaking the independent-increments picture the Poisson likelihood rests on. - **A pile-up of zeros** from a distinct sub-population, which shows up in the same summary statistic even though it is a different problem. Diagnosing which one you have matters, because two of these — clustering and omitted predictors — have fixes better than simply widening the error bars. ## What actually breaks This is the part interviewers listen for, and the answer is narrower than most candidates expect. **The coefficients largely survive.** The Poisson estimating equations depend on the mean model, not on the variance assumption. If the log-linear mean is roughly correct, the estimates remain consistent even though the variance assumption is wrong. Point estimates from a Poisson fit and a negative binomial fit on the same data are usually close. **The standard errors do not survive.** They are derived from the assumed variance. When the true variance is many times larger, reported standard errors are far too small — under a simple proportional-dispersion model they are too small by roughly the square root of the dispersion factor, so a dispersion near 8 means errors understated by something on the order of 2.8-fold. Every downstream quantity inherits the error: intervals too narrow, test statistics too large, predictors that are noise appearing decisively significant. Prediction intervals are similarly overconfident. So the honest one-line summary is: **the estimates are fine, the uncertainty is fiction.** ## How to detect it - Compare the **residual deviance to its residual degrees of freedom**. Under a well-fitting count model these are roughly comparable; a deviance of 900 on 300 degrees of freedom is a loud signal. - Compute the ratio of variance to mean within groups of similar fitted values. A pooled ratio can be inflated purely by predictors the model already explains, so check it conditionally. - Plot residuals against fitted values and look for a fan. - Compare observed and model-predicted frequencies of each count value; a systematic excess in the tails or at zero tells you where the mismatch lives. The reverse case exists too. **Underdispersion** — variance clearly below the mean — points to counts that are capped, rounded, or produced by a scheduled rather than a random process. Poisson standard errors are then conservative, so the risk flips from inventing effects to missing them. ## The two standard repairs **Quasi-Poisson.** Keep the same log-linear mean model but assume the variance is `phi * mu` for an estimated dispersion `phi`. The coefficient estimates are identical to the Poisson ones; every standard error is multiplied by `sqrt(phi)`. It is simple, robust to how the variance actually scales, and requires no new distributional commitment. The cost: it is a quasi-likelihood, not a full likelihood, so likelihood-based model comparison and simulation from the fitted model are not available. **Negative binomial.** Assume a genuine count distribution whose variance is `mu + mu^2/theta`, with `theta` estimated alongside the coefficients by maximum likelihood. As `theta` grows large the extra term vanishes and the fit approaches Poisson. Because the variance grows quadratically rather than linearly with the mean, this model weights small-count observations relatively more than quasi-Poisson does, so coefficients can shift slightly. It is a proper distribution, so you can compare fits by likelihood-based criteria, simulate from it and produce honest prediction intervals. **Choosing between them** comes down to how the variance actually scales with the fitted mean: bin the data by fitted value, compute the variance in each bin, and see whether it tracks the mean linearly (favouring quasi-Poisson) or quadratically (favouring negative binomial). ## The repairs that address the cause instead Widening standard errors accepts the extra variance as noise. Often it is signal. If the excess comes from repeated sessions per user, cluster-robust standard errors or a model with a user-level random effect respects the design and can also improve the mean model. If it comes from an omitted predictor, adding it removes dispersion and improves prediction at the same time. A senior answer names the diagnosis before naming the remedy: reaching straight for negative binomial treats a symptom that may have a better cure.
- How would you choose between a quasi-Poisson and a negative binomial fit here?Look at how the variance actually scales with the fitted mean. Bin observations by fitted value and compare within-bin variance to within-bin mean: roughly linear growth favours quasi-Poisson, which just multiplies standard errors by a constant; roughly quadratic growth favours negative binomial, whose variance is mu + mu^2/theta. Negative binomial is a real distribution, so it also supports likelihood-based comparison and prediction intervals.
- Do the coefficients change much when you switch to negative binomial?Usually only slightly. Both use the same log link and the same mean structure, so the fitted rate ratios come out similar; what moves substantially are the standard errors and hence the p-values. A large shift in the coefficients themselves is a hint that the extra variability is not spread evenly — concentrated in a subgroup, say — which deserves investigating rather than absorbing into a dispersion parameter.
- The extra variance comes from many sessions per user. Is negative binomial still the right answer?Not necessarily the best one. That is clustering, a design fact, and it is better addressed by cluster-robust standard errors at the user level or by a model with a user-level random effect. Those respect the dependence structure and can improve the mean model too. A negative binomial fit would widen the intervals without acknowledging why they needed widening.
- What would variance clearly below the mean suggest instead?Underdispersion, which points to counts that are capped, rounded, or generated by a scheduled rather than a random process. Poisson standard errors are then conservative rather than optimistic, so the risk flips: you are more likely to miss a real effect than to invent one. It is rarer than overdispersion but worth naming when a deviance sits far below its degrees of freedom.
The fit is like a scale calibrated for a steady load being used on a shaking one: the average reading is still about right, but the precision printed on the dial is fantasy.
saying these in an interview costs you the question
- Says overdispersion badly biases the coefficient estimates
- Changes the link function instead of the variance model
- Ignores it because the point estimates look sensible
- Assumes overdispersion always means excess zeros
- Jumps to negative binomial without diagnosing the cause