Why is the Jeffreys prior for a binomial proportion Beta(1/2, 1/2) rather than uniform?
answer
- a rule, not a taste
- square root of Fisher information
- binomial information is n over p(1-p)
- invariance under reparameterisation is the point
- U-shaped, still proper
basics
~10 sJeffreys' rule sets the prior proportional to the square root of the Fisher information. For a binomial proportion that gives the Beta(1/2, 1/2) density. The motivation is invariance under reparameterisation, not neutrality.
solid answer
~50 sJeffreys' rule defines the prior as proportional to the square root of the Fisher information for the parameter. For a binomial proportion `p` the Fisher information from `n` trials is `n/(p(1-p))`, so the prior is proportional to `p^(-1/2) * (1-p)^(-1/2)`, which is the Beta(1/2, 1/2) density — a proper, U-shaped curve with more density near 0 and 1 than a uniform prior has. The motivation is **invariance**: the square root of the Fisher information transforms under a change of variables exactly the way a density does, so applying Jeffreys' rule in one parameterisation and transforming gives the same prior as applying it directly in the other. You get the same inferences whether you model the probability or its log-odds. That fixes the scale-dependence complaint against uniform priors, but it does not make the prior informationless, and Jeffreys' rule is not automatically proper in other models.
go deeper
You are not expected to derive it. Knowing that Jeffreys' rule exists, that it is built from the Fisher information, and that it is about invariance is already a strong showing.
Be ready to state the rule as proportional to the square root of the Fisher information and to land on Beta(1/2,1/2) for a proportion, describing the U shape in words.
Show where it bites and where it does not: small samples and near-boundary proportions versus large data where any reasonable prior gives the same posterior. Volunteer the impropriety caveat for other models.
Take a position on defaults. Decide whether the team's reference analyses use an invariance-motivated rule or an explicit weakly informative prior, and be able to justify that choice to a reviewer who will ask about parameterisation.
## The rule Jeffreys' rule says: take the Fisher information `I(theta)` for the parameter of your model and set the prior proportional to `sqrt(I(theta))`. Fisher information measures how sharply the likelihood distinguishes nearby parameter values — how much a single observation is expected to tell you about `theta` when `theta` is the truth. So the rule places more prior density exactly where the data are most informative per unit of the parameter. ## Working it out for a proportion For `n` independent Bernoulli trials with success probability `p`, the Fisher information is `I(p) = n / (p(1-p))` Taking the square root and dropping the constant `sqrt(n)`: `prior(p) proportional to p^(-1/2) * (1-p)^(-1/2)` That is the kernel of a Beta distribution with both shape parameters equal to 1/2. It is **proper** — it integrates to a finite value despite the spikes, because `x^(-1/2)` is integrable at zero. Its shape is a U: density is highest near `p = 0` and `p = 1`, lowest at `p = 0.5`. Compared with a uniform prior on [0,1], it therefore leans slightly toward the extremes rather than toward even odds. ## Why anyone would want it: invariance The reason to prefer this over uniform is a specific, checkable property. Suppose you reparameterise from `p` to some smooth `psi = g(p)`. A density transforms by multiplying by `|dp/dpsi|`. Fisher information transforms by the *square* of that same derivative, so `sqrt(I)` transforms by `|dp/dpsi|` — exactly the density rule. The consequence: > Applying Jeffreys' rule to the log-odds and transforming back to the probability gives the same prior as applying Jeffreys' rule to the probability directly. So the answer no longer depends on which algebraically equivalent way you wrote your model. That is a real fix for a real complaint: a uniform prior on a probability is a peaked prior on the log-odds, and there is no scale-free notion of "uniform". Jeffreys' rule gives a construction that is stable under relabelling. ## What it is not Three honest caveats separate a good answer from a memorised one. **It is not uninformative.** A U-shaped density is a genuine statement: it says extreme probabilities are a priori more plausible per unit of `p` than middling ones. On a handful of trials this is visible in the posterior. Invariance is the property being bought; ignorance is not for sale. **It is not always proper.** Jeffreys' rule applied to the mean of a normal with known variance yields a flat density over the whole real line, which is improper; applied to the scale parameter of a normal it yields something proportional to `1/sigma`, also improper. Propriety of the binomial case is a happy accident of the algebra, not a guarantee of the rule. **It behaves awkwardly with several parameters at once.** The multivariate version, built from the determinant of the Fisher information matrix, can give priors that practitioners find unattractive in models with a location and a scale together, which is why variants that treat parameters separately, and the broader family of reference priors, exist at all. ## When it actually matters Almost never on large data. With thousands of trials, Beta(1/2,1/2) and a uniform prior give posteriors you cannot tell apart, because the likelihood dominates any prior that puts nonzero density near the truth. The choice shows up in the small-sample and near-boundary cases: a handful of trials, a proportion close to zero or one, an interval you are quoting for a rare event. There it changes the estimate and the interval enough for someone to notice. ## What interviewers are checking This question is a differentiator, not a screener. A strong answer names the `sqrt(Fisher information)` construction, produces Beta(1/2,1/2) with at least a gesture at the derivation, states invariance as the motivation, and then volunteers a caveat — usually that invariance is not the same as ignorance, or that the rule is improper elsewhere. Reciting "Jeffreys is the uninformative prior" is the answer that gets probed and falls over.
- Is the Jeffreys prior always a proper distribution?No. For a binomial proportion it happens to be proper, but for the mean of a normal with known variance Jeffreys' rule gives a flat density over the whole real line, and for the scale parameter it gives something proportional to `1/sigma`. Both are improper, so you have to check that the posterior is a genuine distribution before relying on the result.
- Does the Jeffreys prior count as uninformative?Not in the sense people usually mean. Beta(1/2,1/2) is U-shaped, so it places more density near 0 and 1 than a uniform prior does, and on a handful of trials that visibly shifts the posterior. What it buys is invariance — the same inferences whichever parameterisation you write the model in — not an absence of information.
- When would you actually reach for it instead of a weakly informative prior?When invariance is the property you need to defend — a reference analysis where someone will ask why the answer changes if the model is rewritten in another parameterisation, or a small-sample summary where you want a default nobody chose. If you have genuine outside knowledge, a weakly informative prior stated on the modelling scale is usually the more useful choice.
saying these in an interview costs you the question
- Calls Jeffreys the uninformative prior with no caveat
- Thinks Jeffreys is always uniform
- Assumes the Jeffreys prior is always proper
- Cannot say what invariance means here
- Believes Jeffreys matters at large sample sizes