skip to content

Why does a Poisson regression on counts observed over unequal exposure need an offset?

level: middleimportance: should knowfreq 58%

answer

  1. unequal observation windows inflate counts
  2. counts versus rates
  3. a term with no free coefficient
  4. log of exposure, coefficient pinned at one
  5. the model becomes events per unit exposure

basics

~20 s

Units observed longer accumulate more events for no interesting reason. Adding the log of exposure as an offset, with its coefficient fixed at 1, makes the model describe the event rate per unit of exposure instead of the raw count.

solid answer

~40 s

Say you model support tickets per customer, but customers have been subscribed for different lengths of time. Someone with 24 months on the platform has twelve times the opportunity to file a ticket as someone with 2 months, so raw counts confound the effect you care about with how long each unit was observed. The fix is an offset: add `log(months subscribed)` to the linear predictor with its coefficient forced to 1, giving `log(E[Y]) = log(months) + b0 + b1*x1 + ...`. Rearranged, that is `log(E[Y] / months) = b0 + b1*x1 + ...` — a model for the ticket rate per month. Coefficients then read as rate ratios per unit of exposure, and predictions for a new customer scale automatically with their tenure. Without it, exposure leaks into every coefficient correlated with it.

go deeper

for a junior

Be able to say what goes wrong without one: units watched for longer rack up more events regardless of anything you care about, so the model needs to compare rates rather than totals.

for a middle

Explain the mechanics: the log of exposure is added to the linear predictor with its coefficient fixed at 1, which rearranges to a model for log(count/exposure) and makes exponentiated coefficients rate ratios.

for a senior

Demonstrate that you test the proportionality the offset assumes rather than taking it on faith — fit the same term as a free predictor, compare, and be ready to explain what a coefficient away from 1 means in the domain.

for a principal

Own the definition of exposure itself. Which denominator counts as opportunity is a modelling and product decision with real consequences for what the coefficients mean, and it should be argued explicitly rather than inherited from whatever the table happened to contain.

## The problem an offset solves Counts are only comparable when the opportunity to accumulate them was the same. Customers subscribed for different lengths of time, intersections carrying different traffic volumes, servers up for different numbers of hours — in each case the raw count mixes two things: how frequently events occur for that unit, and how much opportunity that unit had. The quantity of interest is almost always the first. Suppose customer A filed 6 tickets over 24 months and customer B filed 3 tickets over 2 months. On raw counts A looks worse. On rates, A files 0.25 tickets per month and B files 1.5 — six times as many. Any model that regresses raw counts without accounting for tenure will find whatever tenure happens to correlate with, and attribute it to the wrong predictor. ## What an offset is, mechanically An **offset** is a term added to the linear predictor with its coefficient fixed at 1 rather than estimated from the data. It costs no degrees of freedom and produces no coefficient of its own. In a Poisson model with a log link and exposure `t`: ``` log(E[Y]) = log(t) + b0 + b1*x1 + ... ``` Move the offset across: ``` log(E[Y] / t) = b0 + b1*x1 + ... ``` The left-hand side is now the log of a **rate** — events per unit of exposure. Exponentiating the whole thing: ``` E[Y] = t * exp(b0 + b1*x1 + ...) ``` The expected count is the exposure multiplied by a rate that depends on the predictors. Doubling the exposure exactly doubles the expected count, which is the proportionality assumption the offset encodes. ## Why the log, and why fixed at 1 The linear predictor is on the log scale, so anything you want to multiply the mean by must be added there as its logarithm. Adding raw `t` instead would give `E[Y] = exp(t) * exp(b0 + ...)`, an exponential explosion in exposure with no rate interpretation at all. Fixing the coefficient at 1 is what imposes proportionality: one unit more log-exposure means one unit more log-count, so counts scale exactly with exposure. That is a modelling assumption, and it is testable — see below. ## Offset versus a free predictor Consider accident counts per intersection with annual vehicle volume as the exposure. Two options: - **As an offset**: `log(E[Y]) = log(volume) + b0 + ...`. You assert that accidents are proportional to traffic, and the model describes accidents per vehicle. The other coefficients answer 'which intersections are more dangerous per vehicle passing through'. - **As an ordinary predictor**: include `log(volume)` with a coefficient `c` estimated freely. Now `E[Y]` is proportional to `volume^c`. A fitted `c` near 1 supports the offset assumption; `c` clearly below 1 says accidents grow slower than traffic (congestion slows everyone down); `c` above 1 says they grow faster (interaction opportunities compound). The offset is the restricted version of the free-predictor model, so comparing the two is a direct test of whether events really scale one-for-one with exposure. When they clearly do not, using the offset anyway pushes the mismatch into the residuals and into any coefficient correlated with exposure. ## Why not just model the rate directly? A tempting alternative is to compute `y / t` and fit a model to that. It does not work well. The result is no longer a count, so the count likelihood does not apply. Worse, its variability depends on the denominator: a rate computed from one month of data is far noisier than one computed from three years, and a least-squares fit treats them as equally reliable. The offset approach keeps the integer count as the modelled outcome — preserving the likelihood and its natural weighting — while making the systematic part of the model a rate. High-exposure units automatically carry more weight, exactly as they should. ## Practical details Exposure must be strictly positive; a unit with zero exposure has no chance of an event and cannot contribute, so it belongs outside the model rather than inside it with a log of zero. Exposure must also be measured on a scale you can defend: choosing months versus days changes the intercept (`exp(b0)` is a rate per month or per day) but not the other coefficients or their interpretation as rate ratios. Finally, an offset does not fix a variance problem — it corrects the systematic part of the model, and a fit can carry a perfectly good offset while still showing far more variability than the count assumption allows. ## Reading the output With an offset in place, an exponentiated coefficient is a **rate ratio**: the factor by which events per unit of exposure change for a one-unit move in that predictor. Predictions come back on the count scale, already scaled by whatever exposure you supply for the new unit — which is what makes the model usable for forecasting a customer's tickets over the next quarter rather than only comparing groups.

  • What changes if you enter the log of exposure as an ordinary predictor instead of an offset?
    Its coefficient gets estimated rather than pinned at 1, so the model no longer assumes events scale proportionally with exposure. For accidents per intersection, a fitted coefficient on log(vehicle volume) clearly below 1 says accidents grow more slowly than traffic; above 1 says faster. It is the more flexible model, and comparing it against the offset version tests the proportionality assumption directly.
  • Why add the log of exposure rather than exposure itself?
    Because the linear predictor is on the log scale. Adding log(t) there makes the mean multiply by t once you exponentiate: E[Y] = t * exp(b0 + b1*x1 + ...). Adding raw t would instead multiply the mean by exp(t), which grows explosively with exposure and carries no rate interpretation whatsoever.
  • Could you just divide counts by exposure and model the ratio?
    Not cleanly. The ratio is no longer an integer count, so the count likelihood does not apply, and its variability depends on the denominator — a rate from one month is far noisier than one from three years, yet a least-squares fit weights them equally. The offset keeps the count as the outcome, models the rate, and weights each unit by its actual exposure.

An offset is like comparing shops by sales per hour open rather than by total takings: you divide out how long each was trading before comparing anything else about them.

saying these in an interview costs you the question

  • Compares a 2-month and a 24-month customer on raw counts
  • Says the offset coefficient is estimated like any other
  • Adds exposure directly instead of its logarithm
  • Divides counts by exposure and fits least squares
  • Believes an offset also fixes excess variability

context