skip to content

Why is a flat prior on a probability not flat on the log-odds scale?

level: middleimportance: should knowfreq 46%

answer

  1. densities carry units
  2. change of variables adds a factor
  3. derivative of p with respect to log-odds
  4. p(1-p) makes it peak at 0.5
  5. reverse direction is improper

basics

~20 s

A density picks up a Jacobian factor under a change of variables. A uniform prior on p in [0,1] becomes a bell-shaped density on log-odds peaked at zero, favouring p near 0.5. No prior is uninformative on every scale.

solid answer

~50 s

A density is mass per unit of the parameter, so when you reparameterise you must multiply by the derivative of the old parameter with respect to the new one. Take `p` uniform on [0,1] and move to the log-odds `psi = log(p/(1-p))`. Then `dp/dpsi = p(1-p)`, so the implied density on `psi` is `p(1-p)` — the standard logistic density, peaked at `psi = 0` and thinning out toward the extremes. In other words, a "flat" prior on the probability is a mildly informative prior on the log-odds that favours values near even odds. Run it the other way and a flat prior on the log-odds implies a density proportional to `1/(p(1-p))` on the probability, which is improper and piles mass at 0 and 1. The practical lesson: decide which scale you actually model on, and state the prior there deliberately.

go deeper

for a junior

Remember the headline: a uniform prior on a probability is not uniform once you rewrite the model in log-odds. Being able to state that, without the algebra, already beats most answers.

for a middle

Do the transformation out loud. Show the Jacobian, get dp/dpsi = p(1-p), and say what the resulting curve looks like and where it peaks.

for a senior

Connect it to a real fit: name where you have seen a default prior on the wrong scale distort a small-sample estimate, and describe how you now check what a prior implies on the scale you have intuition for.

for a principal

Frame the standard for the team. Argue for specifying priors on the modelling scale with documented implications, rather than letting each analyst pick whatever looks neutral in their own notation.

## Why flatness is not a property of ignorance A prior is a probability density, and a density is defined relative to a unit of the parameter — mass per unit of `p`, or mass per unit of log-odds. Those units are not the same thing, so the same beliefs look like different curves on different scales. Formally, if `psi = g(p)` is a smooth, monotone transformation, the density transforms as `f_psi(psi) = f_p(p) * |dp/dpsi|` The extra `|dp/dpsi|` is the Jacobian factor. It is what makes "flat" a statement about a parameterisation rather than a statement about knowing nothing. ## The concrete case Let `p` be a success probability with a uniform prior on [0,1], so `f_p(p) = 1`. Work on the log-odds scale, `psi = log(p/(1-p))`, whose inverse is the logistic function `p = 1/(1 + exp(-psi))`. Differentiating gives `dp/dpsi = p(1-p)`. Therefore `f_psi(psi) = 1 * p(1-p) = exp(psi) / (1 + exp(psi))^2` That is the standard logistic density: symmetric, peaked at `psi = 0` (which is `p = 0.5`), with most of its mass roughly between `-4` and `+4` in log-odds, i.e. between about 0.02 and 0.98 in probability. So a uniform prior on the probability is *not* agnostic about the log-odds at all — it prefers moderate odds and actively discounts very large or very small ones. Run the mapping the other way. Put a flat density on `psi` across the whole real line. Transforming back, `f_p(p)` is proportional to `1/(p(1-p))`, which integrates to infinity: an improper prior whose mass piles up at both 0 and 1. Combined with a binomial likelihood it produces a proper posterior only when you have seen at least one success and at least one failure; with zero successes there is no posterior at all. "Flat" on the log-odds scale is therefore a strong statement on the probability scale — the belief that the truth is nearly certain in one direction or the other. ## Why this bites in practice It bites whenever a model is written on a transformed scale, which is most of them. Rates and probabilities are commonly modelled through log-odds; durations, counts and variances through logs. If you declare a prior on the natural scale because it "feels neutral" and the model works on the transformed scale, you have specified something you did not intend, and on small samples that shows up directly in the posterior. The effect fades with data. As the sample grows the likelihood becomes sharply peaked and swamps the Jacobian-sized difference between two competing vague priors, so the same analysis on tens of thousands of observations barely notices. It is small samples, rare events and boundary cases — where you were reaching for a "safe" prior in the first place — where the choice moves the answer. ## What to do instead Three defensible responses: 1. **Specify on the modelling scale.** Decide which parameter the model is actually written in and put a proper, weakly informative prior there. For log-odds, a normal centred at the baseline log-odds with a standard deviation around 1.5 keeps most of the mass within roughly plus or minus 3 in log-odds — about 5% to 95% in probability around the centre — while ruling out impossible extremes. 2. **Check the implication on the scale you have intuition for.** Whatever scale you specify on, look at what the prior implies about the quantity you can reason about, and adjust until that looks sane. 3. **Use an invariance rule if you genuinely want one.** There are construction rules whose output transforms consistently under reparameterisation, so the same inferences come out whichever scale you write the model in. They answer the invariance complaint, but they still are not "no information". ## What interviewers listen for The strong answer contains the Jacobian, one worked direction of the mapping, and the honest conclusion: there is no universally uninformative prior, so the professional move is to choose deliberately and then show the conclusion does not hinge on the choice. The weak answer insists that uniform means neutral because every value has the same density — which is exactly the claim the change-of-variables formula refutes.

  • What does a flat prior on the log-odds imply back on the probability scale?
    A density proportional to `1/(p(1-p))`, which is improper — it integrates to infinity and piles mass toward 0 and 1. With a binomial likelihood it yields a proper posterior only when the data contain at least one success and at least one failure; an all-successes sample leaves you with no posterior distribution at all.
  • Which prior would you actually put on a rate you model in log-odds?
    A normal centred on the baseline log-odds you expect, with a standard deviation of roughly 1 to 2. That keeps most of the mass within about plus or minus 3 in log-odds — very roughly 5% to 95% in probability around the centre — which excludes nothing anyone would defend while stopping the estimate running to infinity when a group has no observed failures.
  • Does this scale-dependence matter with a large sample?
    Rarely. Once the likelihood is sharply peaked it dominates any vague prior, and the difference between two reasonable parameterisations of vagueness shrinks to nothing. The issue is confined to thin data, rare events and parameters near a boundary — which is precisely where people reach for a neutral-sounding prior.

Spreading paint evenly on a flat map does not spread it evenly on the globe: the projection stretches some regions and squashes others, and the density changes with it.

saying these in an interview costs you the question

  • Says a uniform prior is neutral on every scale
  • Forgets the Jacobian when transforming a density
  • Assumes flat on log-odds is proper on the probability
  • Confuses transforming a density with transforming a likelihood
  • Treats scale-dependence as a purely theoretical curiosity

context