skip to content

Why model event counts with Poisson regression instead of ordinary linear regression?

level: juniorimportance: must knowfreq 70%

answer

  1. counts are not symmetric or continuous
  2. spread grows with the level
  3. a straight line can predict negatives
  4. log link makes effects multiplicative
  5. percentage change, not extra events

basics

~20 s

Counts are non-negative integers, usually skewed, and their spread grows with their level. Poisson regression models the log of the expected count, so fitted values stay positive and each predictor acts multiplicatively rather than adding a fixed number of events.

solid answer

~50 s

Counts have properties least squares quietly ignores. They are non-negative integers, usually right-skewed with a pile-up at low values, and their variability rises with their level instead of staying constant. A straight-line fit on counts can predict negative expected values, leaves residual spread that widens with the fitted mean, and forces every effect to be additive: the same number of extra events whether the baseline is 1 or 100. Poisson regression instead models `log(E[Y|x]) = b0 + b1*x1 + ...`, so the fitted mean is `exp(...)` and is always positive, and each coefficient multiplies the expected count. That matches how count processes usually behave — a segment files about 20% more support tickets, not 3 more tickets regardless of baseline. It is fitted by maximum likelihood, and the log scale also gives a natural place to put an exposure offset when observation windows differ.

go deeper

for a junior

Be ready to say why a count outcome is awkward for a straight line: no negatives, whole numbers, skew, and spread that grows with the level. Naming the log link and multiplicative effects is enough at this stage.

for a middle

Explain that the model puts the log of the expected count on the linear scale and is fitted by maximum likelihood, so fitted means are the exponential of the linear predictor. Contrast that with regressing log(y+1), which changes the estimand and stumbles on zeros.

for a senior

Show that you check the mean model rather than defending it on principle: inspect fitted values near zero, watch residual spread against the fitted mean, and confirm effects really behave multiplicatively in the domain before shipping the interpretation.

for a principal

Own the call about when a simpler additive model earns its bias. An average events-per-user difference often communicates better to a non-technical audience, and you should be able to state precisely what that framing hides at the low and high ends.

## What a count outcome is A count outcome records how many times something happened in a fixed window of observation: support tickets filed by a customer this quarter, accidents at an intersection this year, page views in a session. Its values are whole numbers, they cannot go below zero, and there is no fixed upper bound. That shape is different from the continuous, roughly symmetric outcome that ordinary least squares was designed around. ## What ordinary least squares assumes, and where counts break it A linear model states `E[Y|x] = b0 + b1*x1 + ...` and, for its usual inference, that the scatter of observations around that line is roughly constant in size and roughly symmetric. Three things go wrong for counts. **Predictions can go negative.** A straight line is unbounded downward. If most units have small counts, the fitted line will dip below zero somewhere in the predictor range, and it will do so exactly in the region you often care about — the low-activity units. **The spread is not constant.** In count data the observations with a high expected value are also the ones that vary the most in absolute terms; units expected to see 2 events cluster tightly, units expected to see 200 swing by tens. Least squares treats all observations as equally noisy, so it under-weights the quiet, informative units and over-weights the loud ones, and the reported standard errors do not reflect the true pattern. **Additive effects are usually the wrong story.** A least-squares coefficient says: one more unit of the predictor adds `b` events. Real count processes rarely work that way. Doubling exposure, moving to a busier product surface, or serving a heavier-usage segment tends to scale events, not shift them. An additive model forced onto a multiplicative process produces a coefficient that is a blend of the effect at high and low baselines, and it fits neither. ## What Poisson regression does instead Poisson regression is a generalized linear model: the outcome is assumed to follow a count distribution, and a **link function** connects the linear predictor to the mean. With the standard log link, ``` log(E[Y|x]) = b0 + b1*x1 + b2*x2 + ... ``` which is equivalent to ``` E[Y|x] = exp(b0 + b1*x1 + b2*x2 + ...) ``` Three consequences follow immediately. 1. **The fitted mean is always positive.** The exponential of any real number is positive, so no combination of predictors ever produces a negative expected count. 2. **Effects are multiplicative.** Adding to the log of the mean multiplies the mean itself. A coefficient of `b` means the expected count is multiplied by `exp(b)` for each one-unit increase in that predictor, holding the others fixed. 3. **Predictors combine multiplicatively too.** On the log scale their contributions add, so on the count scale they multiply. Two factors each raising expected counts by 20% together give about `1.2 * 1.2 = 1.44`, a 44% increase, not 40%. The model is fitted by maximum likelihood rather than by minimising squared error. That fitting automatically gives more weight to observations whose expected count is small — precisely the ones whose variability is small and whose information about the low end is best. ## Why not just take logs of the outcome? A common alternative is to regress `log(y)` with least squares. Two problems. First, `log(0)` is undefined, and counts are full of zeros; the usual patch is `log(y + 1)`, whose results depend on a constant nobody can justify. Second, it changes the estimand: it models the mean of the log rather than the log of the mean. Exponentiating back gives something closer to a median or geometric mean than to the expected count, which is normally the quantity a stakeholder is asking about. Poisson regression models `log(E[Y])` directly, handles zeros without any shift, and returns expected counts. ## What Poisson regression does not assume It does not assume normally distributed errors. It does not require the predictors to be transformed. It does not require the outcome to be small. What it does assume is that, given the predictors, the outcome behaves like a count process whose variability is tied to its mean, and that the log-linear mean is correctly specified. When the variability turns out to be much larger than that assumption allows, the mean model can still be fine while the reported uncertainty is not — a separate issue diagnosed and repaired on its own terms. ## When least squares is still fine If every count is large and far from zero, the distribution is close to symmetric, negative predictions are impossible in the observed range, and an average additive difference is genuinely the quantity of interest, a linear fit can be a reasonable working approximation. That is a deliberate, defensible simplification — quite different from reaching for it by default.

  • When is ordinary least squares on a count outcome actually acceptable?
    When the counts are large and far from zero. In that regime the distribution is near-symmetric, negative fitted values are impossible in the observed range, and an additive average effect may be exactly what the audience wants. It stays a deliberate approximation: check fitted values near the low end and whether residual spread widens with the fitted mean before settling for it.
  • Does log-transforming the count and running a linear model give the same thing?
    No. Taking `log(y)` breaks on zeros, forcing an arbitrary `log(y + 1)` whose results depend on the constant chosen. It also models the mean of the log rather than the log of the mean, so exponentiating back gives a median-like quantity, not the expected count. Poisson regression models `log(E[Y])` directly and handles zeros untouched.
  • What does a log link imply about how two predictors combine?
    On the log scale their effects add, so on the count scale they multiply. Two predictors each raising expected counts by 20% together give roughly `1.2 * 1.2 = 1.44`, a 44% increase rather than 40%. That multiplicative baseline is why an interaction term in such a model tests departure from multiplicativity, not from additivity.

Modelling counts with a straight line is like measuring growth in absolute inches for every species: fine for one animal, misleading the moment you compare a mouse with an elephant. The log link switches you to percentage growth, which travels.

saying these in an interview costs you the question

  • Says counts are fine in least squares because the sample is large
  • Claims Poisson regression assumes normally distributed errors
  • Regresses log(y) with zeros patched by an arbitrary constant
  • Reads a log-link coefficient as extra events per unit
  • Thinks the log link is only about fixing skewness

context