How does a Gamma prior on a support-ticket rate update after observing daily counts?
answer
- the rate lives on positive numbers
- two numbers from the data: total and exposure
- shape collects events, rate collects days
- prior mean is alpha over beta
- say which Gamma convention you use
basics
~20 sEvents add to the shape, exposure to the rate parameter: in shape-and-rate form, Gamma(alpha, beta) with y tickets over n days becomes Gamma(alpha + y, beta + n). The prior reads as alpha pseudo-events over beta pseudo-days.
solid answer
~50 sA Gamma prior is conjugate to a Poisson likelihood for the arrival rate. Using the shape-and-rate form, a `Gamma(alpha, beta)` prior with `y` total tickets observed over `n` days gives a `Gamma(alpha + y, beta + n)` posterior. The prior mean is `alpha / beta` and the posterior mean is `(alpha + y) / (beta + n)`, so the two parameters read directly as pseudo-events over pseudo-days of exposure. If you believed about 5 tickets a day with the confidence of four days of watching, that is `Gamma(20, 4)`; then 30 tickets over 10 real days gives `Gamma(50, 14)`, mean about 3.6 — between the prior's 5 and the observed 3, weighted by exposure. The trap is parametrisation: with the scale form the second parameter is `1 / beta` and the update looks different, so state which one you are using before writing the rule.
go deeper
Be able to recall the pairing itself: a Gamma prior with Poisson count data yields a Gamma posterior, so the update is arithmetic on two numbers.
State the update in shape-and-rate form and explain why the parameters read as pseudo-events over pseudo-exposure.
Handle uneven exposure by summing exposures rather than counting rows, and flag overdispersion as the reason a clean posterior can still be too narrow.
Frame the prior's strength in business terms — it is worth so many days of traffic — so a stakeholder can challenge it directly rather than treating it as a hidden assumption.
## The model Counts of events in a fixed window — support tickets per day, crashes per hour, defects per batch — are commonly modelled as Poisson with an unknown rate lambda. Observing `y_1, ..., y_n` counts over `n` equal-length periods gives a likelihood proportional to `lambda^(sum of y_i) * exp(-n * lambda)` Only two numbers from the data appear: the total count and the number of periods. Those are the sufficient statistics, and their appearance is the reason a conjugate update exists. ## The conjugate prior The Gamma family on positive lambda has density proportional to `lambda^(alpha - 1) * exp(-beta * lambda)` in the **shape-and-rate** parametrisation, where `alpha` is the shape and `beta` is the rate. Its mean is `alpha / beta` and its variance is `alpha / beta^2`. Multiplying prior by likelihood: `lambda^(alpha - 1) * exp(-beta * lambda) * lambda^y * exp(-n * lambda) = lambda^(alpha + y - 1) * exp(-(beta + n) * lambda)` where `y` is the total count. That is the kernel of `Gamma(alpha + y, beta + n)`. The shape absorbs events; the rate parameter absorbs exposure. ## Pseudo-events over pseudo-days The prior mean `alpha / beta` has the form of a count divided by an exposure, and the posterior mean `(alpha + y) / (beta + n)` is literally total events over total exposure once you count the imagined ones. So a Gamma prior is fully described by the sentence "I have already seen alpha tickets in beta days". Work an example. You expect about 5 tickets a day, and you would trust that about as much as four days of observation: `alpha = 20`, `beta = 4`, giving `Gamma(20, 4)` with prior mean 5. Ten real days bring 30 tickets. The posterior is `Gamma(20 + 30, 4 + 10) = Gamma(50, 14)`, mean `50 / 14 = 3.57`. The observed rate is `30 / 10 = 3.0`. The posterior sits between 5 and 3, nearer the data because ten days of real exposure outweigh four days of imagined exposure. As exposure accumulates, `beta` becomes negligible against `n` and the posterior mean converges on the observed rate. The prior's influence is measured in days, which makes it unusually easy to defend or attack in a design review: "this prior is worth four days of traffic" is a claim a stakeholder can actually argue with. ## Why the parametrisation matters Gamma is written two ways. Shape-and-rate uses `exp(-beta * lambda)` with mean `alpha / beta`; shape-and-scale uses `exp(-lambda / scale)` with mean `alpha * scale`. The conjugate update above is stated in the rate form. In the scale form the same update reads `scale_new = 1 / (1 / scale_old + n)`, which is far easier to get wrong. Interviewers do notice when a candidate writes the rule without saying which convention they are in — announcing it first is a cheap signal of care. ## Uneven exposure Days are rarely comparable. If period `i` has exposure `e_i` — hours open, active users, machine-hours — model the count as Poisson with mean `lambda * e_i`. The likelihood becomes proportional to `lambda^(sum y_i) * exp(-lambda * sum e_i)`, so the update is `Gamma(alpha + sum of y_i, beta + sum of e_i)`. Exposure, not the number of rows, goes into the second parameter. Forgetting this is the most common real-world error with this pair: someone adds the number of days when the days were half-length. ## Assumptions worth naming The Poisson likelihood asserts that events arrive independently at a constant rate within the window, which implies the variance equals the mean. Support tickets often violate this: an incident produces a burst, so the observed variance exceeds the mean — overdispersion. The Gamma-Poisson pair still updates cleanly, but the posterior will be too narrow, because the model has been told the data is less variable than it is. The standard response is to move to a model that allows extra variance, at which point the simple conjugate update no longer applies and you have traded closed form for fidelity. The other assumption is a fixed rate over the whole observation window. If ticket volume genuinely doubled after a release, a posterior accumulated since launch will lag badly, because every old day still counts as much as every new one. ## What to say "Gamma is conjugate to Poisson; in shape-and-rate form add the total events to the shape and the total exposure to the rate; the prior mean alpha over beta reads as pseudo-events over pseudo-exposure." Then flag overdispersion, and you have covered the mechanics and the main failure mode.
- What if the observation periods have different exposure lengths?Model the count in period `i` as Poisson with mean `lambda * e_i`, where `e_i` is that period's exposure. The update becomes `Gamma(alpha + sum of counts, beta + sum of exposures)`. Total exposure, not the number of rows, goes into the second parameter — adding the number of days when the days are half-length inflates the estimate.
- Ticket arrivals are bursty rather than steady; what does that do to the posterior?The Poisson likelihood assumes variance equals mean, so bursts make the data overdispersed relative to the model. The posterior mean stays reasonable but the posterior is too narrow — you will report more certainty than you have. Fixing it means a likelihood that allows extra variance, which costs you the closed-form conjugate update.
- Why does the interpretation of a Gamma prior depend on the parametrisation?In shape-and-rate form the mean is `alpha / beta`, so the parameters read as pseudo-events over pseudo-exposure and the update adds counts to each. In shape-and-scale form the second parameter is the reciprocal, the mean is `alpha * scale`, and the same update looks like a harmonic combination. Same distribution, different arithmetic — state your convention first.
saying these in an interview costs you the question
- Adds the total count to both Gamma parameters
- Never states the shape-rate versus shape-scale convention
- Adds the number of rows when exposures differ in length
- Treats the posterior as valid despite obvious burstiness
- Says the Gamma prior mean is alpha times beta