skip to content

What distinguishes an informative prior from a weakly informative or flat prior?

level: juniorimportance: must knowfreq 72%

answer

  1. three rungs, not two
  2. how much work the prior does
  3. weakly informative rules out the absurd
  4. flat is not the same as neutral
  5. unbounded flat means improper

basics

~20 s

An informative prior encodes real outside knowledge and moves the posterior by itself. A weakly informative prior only rules out implausible values, letting the data dominate. A flat prior spreads density evenly and is not genuinely neutral.

solid answer

~40 s

An **informative prior** carries real external evidence — a previous study, a physical constraint, a known base rate — and is deliberately narrow, so it visibly shifts the posterior on its own. A **weakly informative prior** is deliberately vague but still bounded: it puts most of its mass on scientifically plausible values and rules out the absurd, so the likelihood drives the answer while the fit stays stable and identified. A **flat prior** puts equal density across the parameter range; it feels neutral but is not, because flatness depends on which scale you happen to parameterise in, and on an unbounded parameter it is improper. Most applied work defaults to weakly informative priors, reserving informative ones for cases where the outside evidence is real, documented, and defensible to a sceptic.

go deeper

for a junior

Be ready to name the three kinds and give one concrete example of each in a sentence. The trap is saying a flat prior means you added no information.

for a middle

Explain the mechanics: why a uniform density over an unbounded parameter is improper, what propriety of the posterior means, and how a weakly informative prior stabilises an estimate that the data barely identify.

for a senior

Show judgment about which rung a real analysis needs. Talk about thin segments, boundary estimates and separation, and about verifying that the conclusion survives widening the prior before you ship it.

for a principal

Own the standard. Decide what the team's default prior is for routine readouts, what evidence justifies escalating to an informative prior, and how prior choices get documented so no one can tune them after seeing results.

## The choice you are actually making Bayesian inference combines a prior distribution over the unknown parameter with the likelihood of the observed data to produce a posterior. The likelihood comes from the model and the data; the prior is the part you choose. So "which prior?" is a design decision you must be able to justify, and interviewers use it as a proxy for whether you understand what the machinery is doing. The useful taxonomy has three rungs, and they differ in *how much of the posterior they are responsible for*. ## Informative priors An informative prior is narrow enough that it materially moves the answer. It is the right choice when there is genuine, citable outside information: a previous trial on the same drug, a physical bound (a conversion rate cannot exceed 1, a reaction time cannot be negative), an established base rate from years of the same measurement. The test is not "do I have an opinion?" but "could I write down where this came from and would a sceptical reviewer accept it?" Informative priors are most valuable exactly where frequentist estimation struggles: small samples, rare events, and hierarchies with few groups. Their cost is that the conclusion is now partly borrowed, which is why any analysis leaning on one should be shown alongside the answer you would get without it. ## Weakly informative priors A weakly informative prior is the modern default. It is centred on a neutral value and scaled so that the range everyone would call plausible sits comfortably inside it, while genuinely absurd values sit in the tails. If a treatment effect is measured in percentage points and nobody has ever seen one bigger than 15, a distribution centred at zero whose mass mostly lies within roughly plus or minus 10 or 15 points is weakly informative: it asserts almost nothing about the sign or size of the real effect, but it refuses to entertain an effect of 4,000 points. This does real work beyond tidiness. It regularises: when the data barely identify a parameter — a category with three observations, a logistic model where one group has no failures at all and the unconstrained estimate runs off to infinity — a weakly informative prior keeps the estimate finite and the computation stable. And because it is proper (its density integrates to one), the posterior is guaranteed to be a genuine distribution. ## Flat priors, and why "flat" is not "none" A flat prior assigns equal density to every value in the range. On a bounded parameter such as a probability, a uniform density is proper. On an unbounded one — a mean, a regression coefficient — a uniform density over the whole real line has infinite total mass and is **improper**: it is not a probability distribution at all. Improper priors are still usable when the resulting posterior can be shown to be proper, but that is something you check, not assume, and it fails often enough in hierarchical variance parameters to be a real hazard. Improper priors also make marginal-likelihood model comparison meaningless, because the normalising constant is arbitrary. The deeper problem is that flatness is a property of a scale, not of a state of ignorance. Uniform density on a probability is not uniform density on that probability's log-odds, and the two give different answers on small samples. So "I used a flat prior to stay objective" is a claim that does not survive a change of variables — a point interviewers love to press on. ## Choosing in practice A workable rule: pick the scale you actually model on, put a proper weakly informative distribution on it, and check what the prior implies about quantities you have intuition for. Escalate to an informative prior only when you can name the evidence. Then show that the conclusion does not hinge on the choice — or say plainly that it does. As the sample grows, all three rungs converge, provided the prior places nonzero density around the true value. With tens of thousands of informative observations, a weakly informative and a flat prior give practically the same posterior. The choice matters most where data are thin, effects are near a boundary, or a parameter is weakly identified — which is precisely where most real analyses live.

  • What does it mean for a prior to be improper, and is that always a problem?
    An improper prior has infinite total mass — a uniform density over the whole real line, say — so it is not a probability distribution. It can still be used if the resulting posterior is proper, but that has to be verified rather than assumed, and it commonly fails for hierarchical variance parameters. Improper priors also make marginal-likelihood model comparison meaningless, so a proper weakly informative prior is the safer default.
  • How would you defend a weakly informative prior to a sceptical stakeholder?
    Show what it excludes rather than what it asserts. State the range it treats as plausible, point out that every value anyone in the room would defend sits inside it, and demonstrate that the posterior barely moves when you widen it. Framing it as a guardrail against nonsense values, not an opinion about the answer, usually lands better than the mathematics.
  • Where does a weakly informative prior change the answer most?
    Where the data are thin or the parameter is weakly identified: a segment with a handful of observations, a rare event, a group with no observed failures. In those cases an unconstrained estimate can run to a boundary or to infinity, and the prior is what keeps it finite. With a large, informative sample the same prior is nearly invisible.

An informative prior is a witness statement, a weakly informative prior is a fence around the property, and a flat prior is a claim to have no opinion that quietly depends on which map you drew the fence on.

saying these in an interview costs you the question

  • Calls a flat prior no prior at all
  • Says any prior biases the analysis and must be avoided
  • Assumes wider priors are always safer
  • Uses an improper prior without checking posterior propriety
  • Treats an informative prior as a personal hunch rather than cited evidence

context