What is the difference between probability and likelihood after seeing 7 heads in 10 coin flips?
answer
- one formula, two readings
- which symbol is held fixed
- data frozen, parameter free
- a curve over p, not a density
- integrates to 1/11, not 1
basics
~20 sProbability fixes the coin's bias and varies the data; likelihood fixes the observed data and varies the bias. After 7 heads in 10 flips the likelihood is a curve over p that is not a probability distribution.
solid answer
~40 sBoth use the same expression, `P(7 heads in 10 flips | p) = C(10,7) * p^7 * (1-p)^3`, but they treat different symbols as free. Read as a probability, p is fixed at some value such as 0.5 and the data vary: across the 11 possible head counts those probabilities sum to 1. Read as a likelihood, the data are frozen at 7 heads and p is free, giving a curve `L(p)` on the interval from 0 to 1 that peaks at 0.7. That curve is not a density over p — it integrates to 1/11, not to 1 — and multiplying it by any positive constant changes nothing, because a likelihood enters inference only through ratios and shape. Turning it into a distribution over p requires a prior.
go deeper
Be ready to state the swap in one breath: probability holds the parameter fixed and lets data vary, likelihood holds the data fixed and lets the parameter vary. Know that the likelihood curve does not integrate to 1.
Expect to write the binomial expression and explain why constant factors are irrelevant, so that only the shape of the curve matters. Be able to say why a density over the parameter needs a prior.
Show that you read a likelihood curve for both peak and width, and connect width to how much the data actually discriminate. Interviewers listen for whether you resist collapsing the curve to a single number too early.
Own the framing question: when is a full likelihood or posterior worth carrying through a decision pipeline versus summarising, and what does an organisation lose when every analysis reports only a peak with no sense of how sharp it is.
## One formula, two readings Take the standard model for coin flips: the coin lands heads with probability `p`, flips are independent, and you record how many heads appear in `n` flips. The model says `P(k heads in n flips | p) = C(n, k) * p^k * (1-p)^(n-k)` where `C(n, k)` counts the orderings of k heads among n flips. With `n = 10` and `k = 7` this is `120 * p^7 * (1-p)^3`. The expression on its own does not tell you which symbol is the variable and which is held fixed. That choice is the entire distinction between a probability and a likelihood. ### Probability: parameter fixed, data varying Pick a value of the parameter, say `p = 0.5`, and let the data range. You now have a function of `k` over the values 0 through 10. It is a genuine probability distribution: every value is non-negative and the eleven values add to 1. This is the reading you use when you ask forward questions — what outcomes should I expect if the coin behaves like this? ### Likelihood: data fixed, parameter varying Now do the opposite. The data are in hand and no longer uncertain: you saw 7 heads. Let `p` range over the interval from 0 to 1 and write `L(p) = 120 * p^7 * (1-p)^3`. This is the likelihood function. It answers a backward question — for each candidate value of the bias, how well does that value account for what actually happened? It is highest at `p = 0.7` and falls away on either side, quickly toward 0 and 1 where the observed mix of heads and tails becomes implausible. ### Why the likelihood is not a distribution over p Two reasons, and interviewers probe both. First, arithmetic. Integrating `120 * p^7 * (1-p)^3` over p from 0 to 1 gives `1/11`, not 1. (For a binomial likelihood with n flips the integral is always `1/(n+1)`.) So the curve is not normalised over the parameter, and there is no reason it should be — normalisation was imposed on the data axis, not on the parameter axis. Second, meaning. In the classical set-up `p` is a fixed unknown constant, not a random quantity, so there is nothing for a density over `p` to describe. A distribution over `p` only comes into existence when you supply a prior and multiply: the posterior is proportional to prior times likelihood, and the prior is the piece that makes the product integrable to 1 in a meaningful way. ### Only shape and ratios matter Because a likelihood is used through ratios, the constant `C(10,7) = 120` is decorative. Comparing two candidate biases, the factor cancels: `L(0.7) / L(0.5) = (0.7^7 * 0.3^3) / (0.5^10)` and the same cancellation happens inside Bayes rule, where any constant factor is swallowed by the normaliser. This is why textbooks say a likelihood is defined only up to a positive multiplicative constant, and why two experiments whose likelihoods are proportional carry the same information about the parameter. ### Reading the curve A likelihood curve has two features worth naming out loud in an interview. Its peak says which parameter value best accounts for the data. Its width says how sharply the data discriminate between values. With 7 heads in 10 flips the curve is broad: `p = 0.5` still has appreciable support, so the data are far from settling whether the coin is fair. With 70 heads in 100 flips the peak sits at the same 0.7 but the curve is far narrower, and 0.5 is effectively excluded. Same peak, very different evidence — a point a candidate who only quotes the peak will miss. ### Traps to avoid - Saying `L(0.7)` is the probability that the bias equals 0.7. It is not a probability of anything about `p`; it is the probability of the observed data under that value of `p`. - Swapping the conditioning: `P(data | p)` and `P(p | data)` are different objects, and only the second requires a prior. - Comparing likelihood values across different datasets. Within one dataset ratios are meaningful; across datasets the arbitrary constants differ and the comparison is empty.
- Does multiplying a likelihood function by a positive constant change any Bayesian conclusion?No. The posterior is proportional to prior times likelihood, and the normaliser rescales whatever product you hand it, so a constant factor cancels. Likelihood ratios between two parameter values are also unchanged. That is why the binomial coefficient `C(10,7)` can be dropped: it carries no information about p.
- If the likelihood is not a distribution over p, what makes the posterior one?The prior. A prior density over p multiplied by the likelihood gives a function of p with finite integral; dividing by that integral, the evidence, produces a proper density that integrates to 1. Without a prior there is no density over the parameter at all, only a likelihood curve.
- What does a nearly flat likelihood curve tell you?That the data barely discriminate between parameter values: every candidate explains the observations about equally well. In Bayesian terms the posterior comes out close to the prior, because the likelihood contributes almost no reshaping. The practical reading is that you need more data or a more informative design, not a tighter conclusion.
One table of numbers, two ways to read it. Read a row, with the cause fixed, and you get a distribution over outcomes that must add to 1. Read a column, with the outcome fixed, and you get a score for each candidate cause — a column that is under no obligation to add to anything.
saying these in an interview costs you the question
- Says L(0.7) is the probability that the true bias is 0.7
- Insists a likelihood curve must integrate to 1 over the parameter
- Treats P(data given p) and P(p given data) as the same quantity
- Quotes only the peak of the curve and ignores its width
- Compares raw likelihood values computed on different datasets