What does a gamma distribution with shape k and rate lambda represent?
answer
- two parameters, positive support
- waiting for the k-th event
- shape one is a familiar special case
- mean is k over lambda
- rate is an inverse scale
basics
~20 sA gamma with shape k and rate lambda is the waiting time until the k-th event when events happen independently at rate lambda. Its mean is k/lambda, and shape k = 1 gives the exponential.
solid answer
~40 sThe gamma is a two-parameter family on the positive reals. With integer shape k it is the time you wait for the k-th event when single events arrive at rate `lambda`, stacking k exponential waits end to end: mean `k / lambda`, variance `k / lambda^2`. The shape controls the form of the density — `k < 1` piles mass near zero, `k = 1` is exactly the exponential, and `k > 1` gives an interior peak that grows more symmetric as k rises. The rate is an inverse scale: changing it stretches the time axis without altering the shape. Non-integer k is legal, with the gamma function supplying the normalising constant. Watch the parameterisation — many treatments use scale `theta = 1 / lambda`, writing the mean as `k * theta`.
go deeper
Be ready to say the gamma is a positive, right-skewed two-parameter family and that shape 1 is the exponential; the density formula is not expected.
Expect to describe what shape and rate each control, give the mean k/lambda and variance k/lambda^2, and explain the waiting-time-for-k-events reading.
Show you handle the rate-versus-scale convention deliberately and can connect shape to hazard behaviour when a constant-hazard model is too rigid.
Own the modelling tradeoff: the extra shape parameter buys flexibility at the cost of estimation stability on small samples, so argue when that complexity is worth carrying.
## The family The gamma distribution is a flexible two-parameter law for a positive continuous quantity. Its density is ``` f(x) proportional to x^(k - 1) * exp(-lambda * x) for x > 0 ``` with **shape** `k > 0` and **rate** `lambda > 0`. The proportionality constant is `lambda^k / Gamma(k)`, where `Gamma(k)` is the gamma function — a continuous extension of the factorial satisfying `Gamma(n) = (n - 1)!` for positive integers. That constant is what allows non-integer shapes, and it is where the family gets its name. Moments: ``` mean = k / lambda variance = k / lambda^2 SD = sqrt(k) / lambda ``` ## The waiting-time reading The interpretation that makes the family memorable applies when k is a positive integer. If single events occur independently at a constant rate `lambda`, the wait for the first one is exponential with mean `1 / lambda`. The wait for the k-th one is a gamma with shape k and the same rate: you are stacking k independent exponential waits end to end. That immediately explains the moments — k waits of mean `1/lambda` give a total mean of `k / lambda`, and k independent variances of `1/lambda^2` add to `k / lambda^2`. The integer-shape case is often called the Erlang distribution. For non-integer k there is no "count of events" story, but the density is still perfectly well defined and the same moment formulas hold. ## What each parameter actually does **Shape k changes the form of the curve.** - `k < 1`: the density is unbounded as x approaches 0 and decreases throughout — mass crowds towards zero with a long tail. - `k = 1`: the density is `lambda * exp(-lambda x)`, exactly the exponential, peaking at zero. - `k > 1`: the density starts at 0, rises to a single interior mode at `(k - 1) / lambda`, then decays. As k grows the shape becomes steadily more symmetric and bell-like, because the relative spread `SD/mean = sqrt(k)/k = 1/sqrt(k)` shrinks. **Rate lambda changes only the scale.** If X is gamma with shape k and rate lambda, then `c * X` is gamma with shape k and rate `lambda / c`. Changing the rate stretches or compresses the horizontal axis and rescales the height to keep the area at 1; the silhouette is untouched. That is the precise sense in which shape is shape and rate is scale. ## The hazard connection The hazard rate — the instantaneous chance the event happens now given it has not yet — is constant for `k = 1`, increasing for `k > 1` and decreasing for `k < 1`, approaching lambda in both directions as time grows. This is why the gamma is the natural generalisation when a constant-hazard model is too rigid: it lets you express "the longer this has run, the more likely it is to finish" (`k > 1`) or the opposite (`k < 1`) with one extra parameter. ## The parameterisation trap There are two conventions in wide use: ``` rate form: shape k, rate lambda -> mean = k / lambda, variance = k / lambda^2 scale form: shape k, scale theta -> mean = k * theta, variance = k * theta^2, theta = 1 / lambda ``` Both describe the same distribution, and mixing them up inverts your answer by a factor of `lambda^2` in the variance. Say which convention you are using before you quote a number; interviewers notice when a candidate does this unprompted. ## The beta, for contrast The beta distribution is the other standard two-parameter continuous family, and it is easy to confuse with the gamma because both have two positive shape parameters. The difference is the support. A beta lives on the unit interval `[0, 1]`, which makes it the natural way to describe a proportion, a rate between 0 and 1, or an uncertain probability. Its density is proportional to `x^(alpha - 1) * (1 - x)^(beta - 1)`, its mean is `alpha / (alpha + beta)`, and larger `alpha + beta` concentrates it more tightly around that mean. `Beta(1, 1)` is exactly the uniform distribution on `[0, 1]`. So: gamma for an unbounded positive magnitude such as a duration, beta for a bounded fraction. ## What an interviewer is checking Three things, usually. First, that you can say what the two parameters do without reciting a density. Second, that you know `k = 1` collapses to the exponential — the single most useful special case. Third, that you handle the rate-versus-scale ambiguity carefully rather than confidently quoting the wrong moment. The topic itself is a differentiator rather than a screener, so a clear, correctly qualified answer counts for more than exhaustive detail.
- Where does the beta distribution fit alongside the gamma?The beta lives on `[0, 1]`, so it describes a proportion or an uncertain probability rather than an unbounded magnitude. It has two positive shape parameters with mean `alpha / (alpha + beta)`, and a larger `alpha + beta` concentrates the density around that mean. `Beta(1, 1)` is the uniform distribution on the unit interval.
- What happens to the gamma density as the shape k grows large?It becomes progressively more symmetric and bell-shaped, with an interior mode at `(k - 1)/lambda`. The relative spread shrinks because `SD/mean = sqrt(k)/k = 1/sqrt(k)`, so a large-shape gamma is a tight, roughly symmetric hump rather than the skewed curve seen at small k.
- How do the rate and scale parameterisations of the gamma differ?Scale is the reciprocal of rate: `theta = 1/lambda`. In rate form the mean is `k / lambda`; in scale form it is `k * theta`. They describe the same family, but silently switching between them inverts your numbers — always state which convention a quoted mean or variance belongs to.
Shape is the silhouette of the curve and rate is the zoom level on the time axis: change the rate and you are looking at the same shape through a different lens.
saying these in an interview costs you the question
- Says the shape parameter stretches the time axis
- Quotes the gamma mean as k times lambda
- Thinks the shape must be a whole number
- Places the beta distribution on all positive reals
- Cannot name the exponential as the k = 1 case