Which family models a per-user conversion propensity that must lie between 0 and 1?
answer
- match the support to the range first
- hard floor at 0 and ceiling at 1
- two shape parameters, flat to U-shaped
- mean is a over a plus b
basics
~20 sThe beta distribution. Its support is exactly the interval from 0 to 1, and two shape parameters let it be flat, bell-shaped or piled at both ends. A normal would put probability on impossible values.
solid answer
~40 sUse a beta distribution. It is the standard continuous family on the open interval from 0 to 1, with two positive shape parameters `a` and `b` and mean `a / (a + b)`. Both parameters above 1 gives a single interior peak, both equal to 1 gives the uniform, and both below 1 gives a U shape with mass piled near 0 and 1. A normal is wrong because its support is the whole real line: with a propensity averaging 0.05 and a spread of 0.04 it puts a large share of its probability below zero, so sampling from it produces negative propensities. If you observe `k` conversions out of `n` visits rather than a latent rate, model the counts as binomial, or beta-binomial when users vary more than one shared rate allows.
go deeper
Be ready to say that a proportion cannot go below 0 or above 1, and that the family you choose must have support matching that range.
An interviewer expects the beta by name, its two shape parameters, its mean, and a concrete statement of what a normal produces when you sample from it near a boundary.
Show that you separate the latent propensity from the observed count, handle users with tiny visit counts without assigning them impossible certainty, and know when extra between-user spread calls for a beta-binomial.
Own the decision the model serves: whether per-user propensities are worth modelling individually at all versus segmenting, and what the cost is of a model that can emit impossible values into downstream systems.
## The constraint is the whole story A conversion propensity is a proportion: the long-run fraction of a given user's visits that end in a purchase. It cannot be negative and it cannot exceed one. Choosing a family for a bounded quantity starts with matching the **support** of the family to the range of the quantity, before anything about shape or skew is considered. A family whose support extends beyond the possible range is wrong regardless of how well it appears to fit the middle of the data. ## Why the beta The beta distribution is the workhorse continuous family on the unit interval. It has two positive shape parameters, conventionally `a` and `b`, and: - Its support is exactly the interval from 0 to 1. - Its mean is `a / (a + b)`. - Its shape is controlled entirely by the pair. With `a` and `b` both greater than 1, the density has a single interior peak. With both equal to 1, it is the uniform distribution over the interval. With both less than 1, the density is U-shaped, piling mass near both endpoints, which is exactly the right picture for a population where most users are near-never converters and a few are near-always converters. With `a` small and `b` large, it is a steeply right-skewed shape hugging zero, which is what real conversion propensities usually look like. That last point is worth dwelling on. Real propensities are typically tiny, a few percent, with a long thin upper tail. The beta accommodates that without a transform, because the boundary at zero is built into the family rather than imposed afterwards. ## What goes wrong with a normal Suppose you fit a normal with mean 0.05 and standard deviation 0.04 to per-user propensities. Three concrete failures follow: 1. **Impossible values.** More than a tenth of the probability sits below zero. Draw simulated users from this model and some of them convert a negative fraction of the time. 2. **Broken intervals.** Any symmetric interval around the mean extends below zero on one side, so a quoted range is partly nonsense and its coverage claim is void. 3. **Wrong shape.** A normal is symmetric, so it forces the distance from the mean to the upper tail to equal the distance to the lower tail, which is impossible when the mean sits close to a hard boundary. The failure is not subtle or asymptotic; it shows up in the first hundred simulated draws. ## Related choices you should distinguish **A latent rate versus an observed count.** A user's propensity is a number in the unit interval, and the beta is its family. What you actually record is `k` conversions in `n` visits, which is an integer count and belongs to the binomial family given the propensity. Confusing the two leads to fitting a continuous family to raw counts, or treating the ratio `k / n` from `n = 2` visits as if it were a precise propensity, when it can only take the values 0, 0.5 or 1. **Users who differ from one another.** When per-user propensities genuinely vary, the aggregate counts are more spread out than any single shared binomial rate can produce. Letting each user's propensity be drawn from a beta and their conversions from a binomial gives the beta-binomial, which absorbs that extra spread with one additional parameter. **Working on a transformed scale.** An alternative to modelling the proportion directly is to model its log-odds, which map the unit interval onto the whole real line, and to use an unbounded family there. This is what makes bounded outcomes tractable inside regression models. It is a legitimate family choice with a different set of conveniences, not a rival to the beta. ## Edge cases at the boundary The beta as usually stated has support on the *open* interval, so exact zeros and exact ones sit awkwardly. Real data is full of them: users with zero conversions in three visits. Options are to model the count directly rather than the ratio, to use a mixture that gives explicit probability to the endpoints, or to treat the exact zeros as the natural consequence of small `n` rather than as evidence of a zero propensity. Recognising that a user with 0 out of 3 has weak evidence rather than a propensity of exactly zero is the substance of the point. ## How to answer this in an interview Lead with the support argument, name the beta, give the mean formula and one or two shapes to prove you know the parameters do real work, and then say concretely what a normal produces when you sample from it. Finish by separating the latent propensity from the observed count, because that distinction is what the question is usually probing beneath the surface.
- You observe 0 conversions in 3 visits for a user. Is that user's propensity zero?No. Three visits carry almost no information, and 0 out of 3 is entirely ordinary for a propensity of 0.1. The observed ratio is an estimate from a tiny sample, not the parameter. Model the counts directly, or pool across users so that low-volume users are pulled toward the population level instead of being assigned an impossible certainty.
- When would you reach for a beta-binomial instead of a plain binomial?When users genuinely differ, so aggregate conversion counts are more spread out than a single shared rate can produce: too many users at zero and too many near the top. The beta-binomial lets each user draw a propensity from a beta and then draws their conversions binomially, absorbing that extra variation with one more parameter.
- What does a beta with both shape parameters below 1 describe?A U-shaped population: most of the probability piles up near 0 and near 1, with little in between. In product terms that is a base of users who essentially never convert plus a segment that almost always does, and very few in the middle. It is a shape a normal cannot represent at all.
A normal is a tape measure with no end stops; on a scale that physically stops at 0 and 1 you need a ruler that ends where the scale does.
saying these in an interview costs you the question
- Fits a normal to a quantity with hard bounds at 0 and 1
- Treats an observed ratio from three visits as the true propensity
- Thinks the beta has only one parameter
- Cannot say what shapes the two beta parameters produce
- Confuses the latent propensity with the observed conversion count