skip to content

Priors and Likelihoods

What a prior encodes, what the likelihood contributes, and how multiplying them gives a posterior you can update again. Interviewers start here because a vague prior sinks everything downstream.

on this pageshow

explore

questions

21

What is the difference between probability as long-run frequency and probability as degree of belief?

level: juniorimportance: must knowfreq 78%

answer

  1. two meanings behind the same number
  2. repeatable trials versus one-off events
  3. what odds would you take on it?
  4. same axioms, different referent
  5. coherence stops beliefs being arbitrary

basics

~20 s

The frequency reading defines probability as the proportion of times an outcome occurs in repeatable trials. The belief reading defines it as a numeric degree of confidence, so it can also score one-off events such as a single rocket launch.

solid answer

~50 s

Under the frequency interpretation, a probability is a property of a repeatable experiment: it is the limiting share of trials in which the outcome occurs, so it only means something where you can imagine running the same setup many times. Under the degree-of-belief interpretation, a probability is a number describing how confident a particular person is given what they know, so it applies equally to things that will happen exactly once — one specific rocket launch, one specific election — and even to settled facts you have not observed, like whether it rained here yesterday or whether a sealed envelope holds the ace of spades. Both readings use the same probability axioms and the same arithmetic; they disagree about what the number refers to. Where a genuine reference class exists, a well-calibrated believer's number should line up with the observed frequency.

go deeper

for a junior

Be ready to state both definitions in one sentence each and give one example each: a coin toss for frequency, a single specific launch for belief. Knowing that both use the same axioms is enough at this level.

for a middle

Explain why one-off and already-settled events break the frequency reading, and name coherence or the Dutch-book argument as what disciplines personal probabilities. Show you can move between the two readings rather than defending one.

for a senior

Show judgment about which reading fits the question in front of you. Most product questions are one-offs, so be able to say when you would report a degree of belief and how you would keep it honest and calibrated against outcomes.

for a principal

Own the framing for the organisation: which decisions deserve belief-style probabilities, how those numbers are sourced and reviewed, and how forecasts get scored for calibration so that stated confidence is accountable rather than rhetorical.

## Two readings of the same number Everybody agrees on the mathematics of probability: numbers between 0 and 1, the probabilities of mutually exclusive outcomes add, everything sums to 1, and conditional probability is defined by `P(A|B) = P(A and B) / P(B)`. The interview question is about what the number *refers to*, and there are two standard answers. **Frequency.** A probability is a property of a repeatable experiment. Saying a coin lands heads with probability 0.5 means: set up the same experiment over and over, and the share of heads settles down near 0.5. The probability lives in the world, attached to a physical setup, not to any observer. Two analysts who disagree about it disagree about a fact. **Degree of belief.** A probability is a numeric summary of how confident someone is, given what they know. Saying "0.5" means you would be equally happy taking either side of an even-money bet. The probability lives in a mind, indexed by an information state. Two analysts with different information can hold different numbers and both be perfectly reasonable. ## Where the difference bites: one-off and settled events The frequency reading needs a reference class — a family of repetitions the event belongs to. Many questions people genuinely care about have none. - *Will this particular rocket launch succeed?* This vehicle, this crew, this weather. It happens once. There is no infinite sequence of identical launches to take a share of. - *Did it rain at this location yesterday?* Not even a future event. It either rained or it did not; the fact is settled. Nothing about it is random any more. - *Does the sealed envelope in front of you hold the ace of spades?* The card was placed there already. Its identity is fixed. A strict frequentist has to say that these are not probability statements about the event itself; at most they can talk about the mechanism that produced it, if such a mechanism can be described as repeatable. The degree-of-belief reading answers all three without strain: the *event* is settled, your *knowledge* is not, and probability quantifies the knowledge. If a card was drawn uniformly from a shuffled deck and sealed away unseen, your degree of belief that it is the ace of spades is 1/52, even though the card in the envelope is not fluctuating. This is exactly why data-science interviews open here. Most product and business questions are one-offs: will *this* launch move retention, is *this* model good enough to ship. A framework that can attach numbers to those questions is doing something a strict frequency reading cannot. ## Belief is not "anything goes" The usual objection is that if probability is personal, one can pick any numbers at all. The standard reply is the **Dutch-book** argument, developed most sharply by de Finetti. Interpret your stated probability for a claim as the price at which you are willing to buy or sell a bet paying 1 if the claim is true. If your prices violate the probability axioms — you price a claim and its negation at 0.6 each, say — an opponent can construct a set of bets you accept individually that together lose you money no matter what happens. Avoiding such a guaranteed loss (being *coherent*) forces your numbers to satisfy exactly the probability axioms. So beliefs are personal in their content but disciplined in their structure. de Finetti's other contribution is **exchangeability**: if your beliefs about a sequence of trials do not depend on the order in which they arrive, then you behave *as if* there were an unknown fixed rate with a distribution over it, and your beliefs about long runs reproduce ordinary frequency reasoning. Frequencies are not discarded by the belief reading; they are recovered from it wherever repetition is meaningful. ## Where the two must agree A belief that ignores available frequency data is a bad belief, not a valid alternative. If a machine has produced defects at 2% across a million units under stable conditions, a coherent analyst's degree of belief that the next unit is defective is about 2%. **Calibration** is the practical bridge: over many statements you make at "80% confident", roughly 80% should turn out true. That test is itself a frequency check applied to beliefs, and forecasters are routinely scored on it. ## How to answer this in an interview Say the two definitions crisply, give one example that only the belief reading handles (a specific one-off event, or a settled fact you have not observed), say that both obey the same axioms, and add that where a reference class exists a good belief matches the frequency. That last sentence is what separates a candidate who has thought about it from one reciting a slogan. Avoid framing it as a war between camps: in practice most working statisticians pick the reading that suits the question in front of them.

  • If probability is personal belief, what stops two people picking whatever numbers they like?
    Coherence. Read your stated probabilities as betting prices you would take either side of. If they violate the probability axioms, an opponent can assemble bets you accept that lose you money in every outcome — a Dutch book. Avoiding a guaranteed loss forces exactly the standard axioms. Content is personal; structure is not optional. Beliefs also get scored against reality through calibration.
  • Does a strict frequentist deny that a specific one-off event has a probability?
    Roughly, yes. With no repeatable reference class there is no limiting frequency, and for an event already settled there is nothing left to be random. A frequentist can still model the mechanism that generated it as repeatable and reason about that, but declines to attach a probability to the particular settled outcome. Most practitioners are pragmatic about this in conversation.
  • Where do the two interpretations have to agree?
    Wherever a genuine reference class exists and the setup is stable. If a process has produced an outcome in 2% of a million comparable trials, a coherent degree of belief about the next trial is close to 2%. Ignoring solid frequency evidence is a bad belief, not a legitimate alternative reading, and calibration scoring makes that failure visible.

A frequency probability is a factory's defect rate, read off from many nearly identical items. A belief probability is the price you would accept on either side of a bet about one sealed envelope, where no repetition is available at all.

saying these in an interview costs you the question

  • Says Bayesian probability is just guessing, with no rules
  • Claims the two interpretations use different probability axioms
  • Insists every probability requires a repeatable experiment, including for settled facts
  • Treats a degree of belief as fixed and immune to data
  • Thinks frequency probabilities are objective and beliefs are therefore worthless

context

open as a page

Starting from a Beta(1,1) prior, what posterior follows from 8 clicks in 100 impressions?

level: juniorimportance: must knowfreq 58%

basics

~10 s

Beta(1 + 8, 1 + 92), that is Beta(9, 93). The prior contributes one pseudo-success and one pseudo-failure, so the posterior mean is 9/102, about 0.088, slightly above the raw rate of 0.08.

open as a page

What is the difference between probability and likelihood after seeing 7 heads in 10 coin flips?

level: juniorimportance: must knowfreq 75%

basics

~20 s

Probability fixes the coin's bias and varies the data; likelihood fixes the observed data and varies the bias. After 7 heads in 10 flips the likelihood is a curve over p that is not a probability distribution.

open as a page

What distinguishes an informative prior from a weakly informative or flat prior?

level: juniorimportance: must knowfreq 72%

basics

~20 s

An informative prior encodes real outside knowledge and moves the posterior by itself. A weakly informative prior only rules out implausible values, letting the data dominate. A flat prior spreads density evenly and is not genuinely neutral.

open as a page

In Bayesian versus frequentist inference, is the unknown parameter treated as random or as fixed?

level: middleimportance: must knowfreq 64%

basics

~20 s

Frequentists treat the unknown parameter as a fixed constant and put all the randomness in the data. Bayesians give the parameter a probability distribution that describes how uncertain they are about its value, and condition on the data actually observed.

open as a page

What makes a Beta prior conjugate to a binomial likelihood in Bayesian updating?

level: middleimportance: must knowfreq 70%

basics

~20 s

A Beta density and a binomial likelihood are both powers of theta and 1 minus theta, so their product is again a Beta: a Beta(a, b) prior with s successes in n trials gives Beta(a + s, b + n - s).

open as a page

In Bayes' rule for a parameter, what does 'posterior is proportional to prior times likelihood' mean?

level: middleimportance: must knowfreq 70%

basics

~20 s

It means the posterior density at each parameter value is prior times likelihood divided by a single constant, the evidence. That constant does not depend on the parameter, so the product alone determines the posterior's shape.

open as a page

In Normal-Normal conjugate updating, how does the posterior mean combine prior and data?

level: middleimportance: should knowfreq 44%

basics

~20 s

As a precision-weighted average, where precision is one over variance. Precisions add, so the posterior is always more precise than the prior, and its mean lies between the prior mean and the sample mean, closer to whichever is more precise.

open as a page

When can you ignore the marginal likelihood in Bayes' rule, and when must you actually compute it?

level: middleimportance: should knowfreq 45%

basics

~20 s

Inside a single model the marginal likelihood is a constant that only rescales the posterior, so you can ignore it. You must compute it to compare models, since a Bayes factor is a ratio of two marginal likelihoods.

open as a page

Why is a flat prior on a probability not flat on the log-odds scale?

level: middleimportance: should knowfreq 46%

basics

~20 s

A density picks up a Jacobian factor under a change of variables. A uniform prior on p in [0,1] becomes a bell-shaped density on log-odds peaked at zero, favouring p near 0.5. No prior is uninformative on every scale.

open as a page

A new feature shows 3 conversions in 40 sessions: why would a prior earn its keep here?

level: seniorimportance: should knowfreq 48%

basics

~20 s

At 40 sessions the raw 7.5% rate is mostly noise: one more conversion would read 10%. A prior built from comparable past features supplies the information the data lacks and pulls the estimate toward plausible values.

open as a page

Why does updating a Beta posterior event-by-event match one batch update of the same data?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Because the posterior becomes the next prior and the update only adds sufficient statistics. For exchangeable data the likelihood factorises and multiplication commutes, so any batching or ordering accumulates the same success and failure counts and lands on identical parameters.

open as a page

Your posterior for a parameter is bimodal — what goes wrong if you report only the posterior mean?

level: seniorimportance: should knowfreq 48%

basics

~10 s

With two separated modes the posterior mean lands in the trough between them, a value the posterior itself calls unlikely. The single number also hides that two competing explanations are in play.

open as a page

How does a hierarchical prior partially pool eight noisy per-group effect estimates?

level: seniorimportance: should knowfreq 52%

basics

~20 s

It treats the eight group effects as draws from a common distribution whose mean and spread are learned from the data. Each estimate is then pulled toward the overall mean, the noisiest groups moving furthest.

open as a page

How would you decide whether a team reports Bayesian or frequentist results for its recurring decisions?

level: principalimportance: should knowfreq 37%

basics

~20 s

Decide by the shape of the decisions, not by taste. Bayesian reporting pays off when data per decision is thin, credible prior information exists, and stakeholders need a probability about the claim itself. Frequentist reporting suits high-volume standardised readouts.

open as a page

When is the convenience of a conjugate prior not worth the constraint it puts on your model?

level: principalimportance: should knowfreq 34%

basics

~20 s

When the family cannot express the belief or the structure the problem has. Closed form buys exact, constant-memory updates worth keeping at high throughput, but bending a bimodal belief or a covariate-driven model into a convenient family is a modelling error.

open as a page

How do you handle a Bayesian efficacy readout whose conclusion flips between a sceptical and an enthusiastic prior?

level: principalimportance: should knowfreq 38%

basics

~20 s

Report the flip as the finding: if the decision changes across priors reasonable people hold, the data are not decisive. Pre-specify the prior set, quantify the tipping point, and decide on the cost of being wrong.

open as a page

Why do two analysts with different reasonable priors reach nearly the same conclusion as data grows?

level: middleimportance: nice to knowfreq 31%

basics

~20 s

Each observation multiplies more likelihood into the posterior, while the prior enters only once. With enough data the likelihood swamps it, so both posteriors concentrate on the same value and take the same approximately normal shape — the Bernstein-von Mises result.

open as a page

How does a Gamma prior on a support-ticket rate update after observing daily counts?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Events add to the shape, exposure to the rate parameter: in shape-and-rate form, Gamma(alpha, beta) with y tickets over n days becomes Gamma(alpha + y, beta + n). The prior reads as alpha pseudo-events over beta pseudo-days.

open as a page

Why is the Jeffreys prior for a binomial proportion Beta(1/2, 1/2) rather than uniform?

level: middleimportance: nice to knowfreq 26%

basics

~10 s

Jeffreys' rule sets the prior proportional to the square root of the Fisher information. For a binomial proportion that gives the Beta(1/2, 1/2) density. The motivation is invariance under reparameterisation, not neutrality.

open as a page

Two coin experiments with different stopping rules give proportional likelihoods — why is the posterior identical?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Two likelihood functions that differ only by a factor free of the parameter give the same posterior, because that factor is absorbed by the normalising constant. With the same prior the two experiments' posteriors coincide exactly.

open as a page