skip to content

Starting from a Beta(1,1) prior, what posterior follows from 8 clicks in 100 impressions?

level: juniorimportance: must knowfreq 58%

answer

  1. Beta(1,1) is flat on zero to one
  2. count the clicks and the non-clicks
  3. one pseudo-success, one pseudo-failure
  4. add 8 and add 92
  5. mean is a over a plus b

basics

~10 s

Beta(1 + 8, 1 + 92), that is Beta(9, 93). The prior contributes one pseudo-success and one pseudo-failure, so the posterior mean is 9/102, about 0.088, slightly above the raw rate of 0.08.

solid answer

~40 s

A `Beta(1, 1)` prior is flat on the interval from 0 to 1 and reads as one pseudo-success plus one pseudo-failure. With 8 clicks and 92 non-clicks, the posterior is `Beta(1 + 8, 1 + 92) = Beta(9, 93)`. Its mean is `9 / 102 = 0.088`, pulled a little above the observed 0.08 because the flat prior's two pseudo-observations sit at a rate of one half. Its mode is `(9 - 1) / (102 - 2) = 0.08`, which coincides exactly with the maximum-likelihood estimate — a flat prior leaves the peak of the likelihood untouched and only shifts the mean. Reporting the whole `Beta(9, 93)` rather than a single number is the point: it tells you the plausible range as well as the centre, and it becomes the prior for tomorrow's traffic.

go deeper

for a junior

Be ready to do this arithmetic instantly: add successes to the first parameter, failures to the second, and quote the mean as first over the sum of both.

for a middle

Explain why the mean is 0.088 while the mode stays at 0.080, and what the two pseudo-observations of a flat prior are actually doing.

for a senior

Point out that the posterior pair is all the state you need, that it becomes tomorrow's prior, and that 100 impressions at an 8% rate leaves a wide interval.

for a principal

Be able to argue whether a flat prior is defensible for a click-through rate at all, given that rates above one half are effectively impossible in the domain.

## What Beta(1, 1) is The `Beta(a, b)` density is proportional to `theta^(a-1) * (1-theta)^(b-1)` on the interval from 0 to 1. Setting `a = b = 1` makes both exponents zero, so the density is constant: `Beta(1, 1)` is the uniform distribution on `(0, 1)`. It says every click-through rate between 0 and 1 is equally plausible before you see data. In pseudo-count language it is the weakest prior that still carries counts: one imagined success and one imagined failure, an effective prior sample size of 2. ## The update With a binomial likelihood, the conjugate update adds successes to the first parameter and failures to the second. Eight clicks in 100 impressions means `s = 8` successes and `n - s = 92` failures, so `Beta(1 + 8, 1 + 92) = Beta(9, 93)` Two numbers now summarise everything you know about the rate. ## Reading the posterior **Mean.** The mean of `Beta(a, b)` is `a / (a + b)`, here `9 / 102 = 0.0882`. The raw click-through rate is `8 / 100 = 0.08`. The posterior mean sits slightly higher because the prior's two pseudo-observations are split evenly, dragging the estimate a hair toward one half. With 100 real observations against 2 pseudo-observations, that pull is tiny — as it should be. **Mode.** The mode of `Beta(a, b)` for `a > 1, b > 1` is `(a - 1) / (a + b - 2)`, here `8 / 100 = 0.08`, exactly the maximum-likelihood estimate. This is a useful thing to know: a flat prior does not move the peak of the posterior at all, because multiplying the likelihood by a constant cannot shift where it is maximised. It moves the *mean*, because the posterior is a skewed distribution and the flat prior fattens the upper tail. **Spread.** The posterior is far from a point estimate. `Beta(9, 93)` puts most of its mass roughly between 0.04 and 0.15; 100 impressions is simply not much data for an 8% rate. The number of *successes*, not the number of trials, drives the precision here, which is why low-rate metrics need much more traffic than people expect. ## Why the count interpretation is the whole trick Because the update is addition, you never need the dataset again — only the running totals. If tomorrow brings 3 clicks in 60 impressions, you update `Beta(9, 93)` to `Beta(12, 150)` and you are done. The posterior becomes the next prior, and the arithmetic is identical whether the data arrives all at once or in dribs. ## The zero-success case The same machinery handles the case that breaks naive estimation. Suppose a new creative gets 0 clicks in 50 impressions. The raw rate is `0 / 50 = 0`, which asserts the rate is exactly zero — a claim 50 impressions cannot support. With a `Beta(1, 1)` prior the posterior is `Beta(1, 51)`, mean `1 / 52 = 0.019`. In general the posterior mean under a flat prior is `(s + 1) / (n + 2)`, a result old enough to have its own name: Laplace's rule of succession, and the same idea appears in classification as add-one smoothing. It never returns exactly 0 or exactly 1, which is precisely what you want when a downstream step takes a logarithm, divides by the estimate, or ranks items by it. ## Common mistakes - Writing `Beta(9, 101)`: adding the number of trials to the second parameter instead of the number of failures. - Writing `Beta(8, 92)`: dropping the prior's pseudo-counts. - Calling `9 / 102` the probability that the rate is 0.088. It is the posterior *mean*; the rate is a continuous quantity, so the probability of any single exact value is zero. - Treating `Beta(1, 1)` as "no prior". It is a genuine assumption — that every rate from 0 to 1 is equally likely, which for a click-through rate is a fairly strange belief, since rates above 0.5 are basically unheard of. It is weak, not absent. ## What to report Hand back the pair `(9, 93)` alongside the summary you care about. The two parameters are the complete state, the mean or mode is the headline number, and the width of the distribution is the honest statement about how little 100 impressions tells you.

  • A new creative gets 0 clicks in 50 impressions; what does the same prior give?
    `Beta(1 + 0, 1 + 50) = Beta(1, 51)`, mean `1 / 52`, about 0.019. Under a flat prior the posterior mean is always `(s + 1) / (n + 2)` — Laplace's rule of succession, the same idea as add-one smoothing. It never returns exactly zero, which matters when something downstream takes a logarithm of the estimate or ranks by it.
  • Why is the posterior mean 0.088 while the mode is 0.080?
    The posterior is right-skewed at a low rate, so its mean sits above its peak. The flat prior does not move the peak at all — multiplying the likelihood by a constant cannot change where it is maximised, so the mode stays at the maximum-likelihood value 8/100. The prior's two pseudo-observations, split evenly, are what pull the mean up.
  • How much does the Beta(1,1) prior actually influence this answer?
    Barely. It contributes 2 pseudo-observations against 100 real ones, so it shifts the mean by under a percentage point in absolute terms. Its real effect is at the edges: it keeps the estimate strictly inside 0 and 1 and gives a usable answer when successes are zero, which is where the raw proportion breaks down.

saying these in an interview costs you the question

  • Answers Beta(9, 101) by adding trials instead of failures
  • Answers Beta(8, 92), forgetting the prior's pseudo-counts
  • Calls Beta(1,1) no prior rather than a weak one
  • Reports only 0.08 and drops the posterior spread
  • Says the posterior mean is the probability the rate equals 0.088

context