What makes a Beta prior conjugate to a binomial likelihood in Bayesian updating?
answer
- the family survives the multiplication
- look at the exponents of theta
- powers of theta and one minus theta
- successes go to a, failures to b
- no evidence integral to evaluate
basics
~20 sA Beta density and a binomial likelihood are both powers of theta and 1 minus theta, so their product is again a Beta: a Beta(a, b) prior with s successes in n trials gives Beta(a + s, b + n - s).
solid answer
~50 sConjugacy means the posterior lands in the same parametric family as the prior, so updating is arithmetic on parameters instead of an integral. A `Beta(a, b)` density is proportional to `theta^(a-1) * (1-theta)^(b-1)`, and the binomial likelihood for `s` successes in `n` trials is proportional to `theta^s * (1-theta)^(n-s)`. Multiplying gives `theta^(a+s-1) * (1-theta)^(b+n-s-1)`, which is exactly `Beta(a + s, b + n - s)` — you add successes to the first parameter and failures to the second. The normalising constant, the part that usually needs numerical work, comes free because you already recognise the family. That also gives the pseudo-count reading: `a + b` behaves like a prior sample size, and the posterior mean `(a + s) / (a + b + n)` is a weighted average of the prior mean and `s / n`. Conjugacy buys tractability, not correctness — it says nothing about whether the prior is a good description of your belief.
go deeper
Be ready to state the update rule out loud: a Beta prior plus binomial data gives a Beta posterior, with successes added to the first parameter and failures to the second.
Expect to derive it. Write the Beta kernel and the binomial likelihood as powers of theta and 1 minus theta, multiply, and read the new parameters off the exponents.
Show why it matters operationally: no evidence integral, constant-size state per item, and a posterior that serves directly as the next prior in a streaming update.
Own the tradeoff. Explain when the narrowness of the conjugate family misrepresents real belief badly enough that you should abandon closed form and pay for numerical inference.
## The general idea Bayes' rule says the posterior is proportional to the prior times the likelihood: `posterior(theta | data) = prior(theta) * L(theta; data) / evidence` The denominator, the evidence, is the integral of `prior(theta) * L(theta; data)` over all values of theta. For most models that integral has no closed form, which is why Bayesian computation is often numerical. A prior family is called **conjugate** to a likelihood when the product `prior * likelihood`, viewed as a function of theta, has the same algebraic shape as the prior. When that happens you can read the posterior parameters straight off the exponents and never touch the integral, because you already know the normaliser of that family. Conjugacy is a property of a *pair*: a family of priors together with a likelihood. It makes no sense to call a prior conjugate on its own. ## Why Beta and binomial fit together A `Beta(a, b)` distribution lives on the interval `(0, 1)`, which is exactly where a probability parameter lives. Its density is `f(theta) = theta^(a-1) * (1-theta)^(b-1) / B(a, b)` where `B(a, b)` is a constant that does not depend on theta. Its mean is `a / (a + b)`, and when `a > 1` and `b > 1` its mode is `(a - 1) / (a + b - 2)`. Observing `s` successes in `n` independent trials with success probability theta gives a likelihood proportional to `theta^s * (1-theta)^(n-s)` The binomial coefficient is dropped because it does not involve theta, and anything free of theta is absorbed into the normaliser. Multiply the two: `theta^(a-1) * (1-theta)^(b-1) * theta^s * (1-theta)^(n-s) = theta^(a+s-1) * (1-theta)^(b+n-s-1)` That is the kernel of a `Beta(a + s, b + n - s)`. The update rule is therefore: **add the successes to the first parameter and the failures to the second**. A common slip is adding `n` to the second parameter instead of `n - s`. ## The pseudo-count reading Because the update just adds counts, the prior parameters behave like counts you pretend to have already seen. A `Beta(3, 7)` prior on a conversion rate acts like 3 prior successes and 7 prior failures — a prior mean of 0.3 carrying about 10 observations of weight. `Beta(30, 70)` has the same mean but ten times the weight and moves ten times less when the data arrives. The posterior mean makes the weighting explicit: `(a + s) / (a + b + n)` which can be rewritten as `[(a + b) / (a + b + n)] * [a / (a + b)] + [n / (a + b + n)] * [s / n]` an average of the prior mean and the observed proportion, weighted by prior strength `a + b` against sample size `n`. As `n` grows, the second term dominates and the posterior mean converges on the observed proportion. That is the precise sense in which data eventually outweighs a fixed prior. ## Other standard pairs The same pattern recurs across the exponential family: - **Gamma prior, Poisson likelihood** for a rate: the prior contributes pseudo-events over pseudo-time-units. - **Normal prior, Normal likelihood with known variance** for a mean: precisions add and the posterior mean is a precision-weighted average. - **Dirichlet prior, multinomial likelihood** for a probability vector: the multivariate generalisation of Beta-binomial, updated by adding category counts. - **Gamma prior, exponential likelihood** for a rate: add the number of observations to the shape and the sum of the waiting times to the rate parameter. Every likelihood in the exponential family has a conjugate prior family; the recipe is always "prior parameters accumulate the sufficient statistics of the data". ## What conjugacy does and does not buy you It buys: closed-form posteriors, exact answers with no sampler and no convergence diagnostics, constant-size state (you carry two numbers, not a dataset), and a posterior that is itself a valid prior for the next batch. That combination is why conjugate updates are the natural fit for online systems that maintain a rate estimate per item. It does not buy correctness. Conjugate families are narrow and unimodal in the usual parametrisations, so if your genuine belief is "this rate is either near zero or near one half", no single Beta expresses it and forcing one distorts the analysis. Conjugacy also evaporates the moment the model gains structure the family cannot absorb — covariates, unknown dispersion, measurement-error layers. The honest framing in an interview is: conjugacy is a computational convenience that happens to be available for a handful of very common models, and choosing a prior because it is convenient rather than because it describes your belief is a decision you should be able to defend.
- What does the sum a + b tell you about a Beta prior?It is the prior's strength, an effective prior sample size. `Beta(3, 7)` and `Beta(30, 70)` both centre on 0.3, but the second carries about 100 pseudo-observations, so 100 real trials move it roughly half way while they nearly overwhelm the first. Quoting `a + b` is the quickest way to say how much evidence your prior is worth.
- Name another conjugate pair and its update rule.A Gamma prior on a Poisson rate. With the shape-and-rate parametrisation, `Gamma(alpha, beta)` observing `y` total events over `n` time units becomes `Gamma(alpha + y, beta + n)`. The shape accumulates events and the rate parameter accumulates exposure, so the prior reads as `alpha` pseudo-events observed over `beta` pseudo-units of time.
- Does conjugacy mean the prior is a good choice?No. Conjugacy is a statement about algebra, not about belief. It guarantees the posterior stays in a known family; it says nothing about whether that family can represent what you actually believe. If your belief is bimodal or has hard bounds a Beta cannot express, the convenient prior is the wrong prior, and you should say so rather than bend the belief to fit the math.
It is like adding two numbers in the same units: counts plus counts stay counts. The prior arrives already denominated in successes and failures, so the data just increments the tally.
saying these in an interview costs you the question
- Says conjugacy makes the prior correct rather than convenient
- Updates Beta(a, b) to Beta(a + s, b + n) instead of b + n - s
- Claims any prior is conjugate to any likelihood
- Thinks the evidence integral still has to be computed numerically
- Treats a and b as arbitrary knobs with no count interpretation