Two coin experiments with different stopping rules give proportional likelihoods — why is the posterior identical?
answer
- what changes between the two designs
- compare the two expressions in p
- 220 versus 55, both free of p
- constants are absorbed by the normaliser
- tail areas need a sample space, likelihoods do not
basics
~20 sTwo likelihood functions that differ only by a factor free of the parameter give the same posterior, because that factor is absorbed by the normalising constant. With the same prior the two experiments' posteriors coincide exactly.
solid answer
~40 sSuppose one experimenter fixes 12 flips and records 9 heads and 3 tails, giving `C(12,3) * p^9 * (1-p)^3`, while another flips until the third tail appears and it arrives on flip 12, giving `C(11,2) * p^9 * (1-p)^3`. The constants differ, 220 against 55, but neither depends on `p`. Since the posterior is proportional to prior times likelihood and constant factors are swallowed by the evidence, the same prior produces literally the same posterior. This is the likelihood principle: once data are observed, everything they say about the parameter is carried by the likelihood up to a constant. Tail-area p-values need not agree, because they sum probability over the outcomes each design could have produced. The ignorability holds when the stopping rule depends only on observed data.
go deeper
Recall that a constant factor multiplying a likelihood cancels in Bayes rule, so it cannot change a posterior. That single fact is the engine behind the whole example.
Be able to write both probabilities, the fixed-count one and the stop-at-third-tail one, and point out that they share the same function of the parameter with different leading constants.
State the likelihood principle, show the cancellation, and explain why tail-area calculations depend on the design while likelihoods do not. Then add the caveat that ignorability needs a rule depending only on observed data.
Own the practical consequence: what your organisation's analysis policy says about interim looks and early stopping, and how to keep the honest Bayesian answer from being read as permission to peek at frequentist readouts without adjustment.
## The two designs Let `p` be the probability of heads on a single flip, with flips independent. **Design A, fixed sample size.** Flip exactly 12 times and count. The experimenter observes 9 heads and 3 tails. The probability of that outcome is `C(12,3) * p^9 * (1-p)^3 = 220 * p^9 * (1-p)^3` **Design B, stop at the third tail.** Keep flipping until the third tail appears, then stop. It happens to appear on flip 12, so again 9 heads and 3 tails. The probability of that outcome requires exactly 2 tails among the first 11 flips and a tail on flip 12: `C(11,2) * p^9 * (1-p)^2 * (1-p) = 55 * p^9 * (1-p)^3` The two functions of `p` are identical apart from the factor 220 versus 55. ## Why the posterior cannot notice Bayes rule says the posterior is proportional to prior times likelihood, with the evidence supplying whatever scale is needed to make the result integrate to 1. Multiply a likelihood by 4 and the evidence quadruples too; the ratio is unchanged at every parameter value. Formally, if `L_B(p) = k * L_A(p)` for a constant `k` that does not depend on `p`, then `prior(p) * L_B(p) / integral of prior * L_B = prior(p) * L_A(p) / integral of prior * L_A` because `k` factors out of both the numerator and the integral. Same prior, same posterior — not approximately, exactly. ## The likelihood principle The general statement is the likelihood principle: if two experiments yield likelihood functions for the same parameter that are proportional, they carry the same evidence about that parameter, and inference should be the same. Bayesian updating satisfies it automatically, since it only ever uses the likelihood through the product with the prior. That is a genuine structural property, not a coincidence of this coin example. The intuition is worth stating plainly. The stopping rule describes what the experimenter would have done in situations that did not occur. Once the data are in hand, those counterfactual branches contributed nothing to what was seen. Whether you would have kept flipping had the third tail come later does not change how well `p = 0.7` explains 9 heads and 3 tails. ## Where the two frameworks diverge Tail-area procedures do notice the design, because they are defined over the sample space of the experiment. A p-value adds up the probability of outcomes at least as extreme as the observed one, and *at least as extreme* has to be enumerated over the outcomes the design could have produced. Design A can produce any head count from 0 to 12; design B can produce any number of flips from 3 upward. These are different sets, so the two computations sum over different things and can return different numbers from the same observed sequence. This is the classic illustration that frequentist tail-area inference violates the likelihood principle, and it is the reason optional stopping is a genuine hazard for naive fixed-sample p-values in a way that has no direct Bayesian counterpart. ## The honest caveats A candidate who says stopping rules never matter has overclaimed. The ignorability argument requires that the rule depend only on data the model already conditions on. - If the decision to stop depends on something the model omits — the experimenter peeked at an unrecorded covariate, or halted because of an external event correlated with the outcome — the stopping mechanism is informative and the simple factorisation into a constant times the likelihood fails. - If stopping depends on the parameter itself in a way not represented by the model, the same failure occurs. - Design still matters *before* the data exist. Choosing how long to run determines how much information you will collect and therefore how precise the posterior will be. The principle says the rule adds nothing once the data are observed; it does not say the rule is irrelevant to planning. - Model checking is not covered either. Assessing whether a model is adequate can legitimately consider what the design could have produced, which is outside the parameter-inference statement. ## What to say in an interview Three beats. First, show the two likelihoods and point out that they differ by a constant in `p`. Second, show that the constant cancels in the posterior, so Bayesian updating is untouched. Third, add the caveat: this holds because the stopping rule depends only on observed data, and it does not make experimental design irrelevant — it makes the design irrelevant *to the parameter's likelihood once the data are in*.
- Does this mean experimental design and sample-size planning are irrelevant to a Bayesian?No. Design determines what data you will get and therefore how precise the posterior will be, so planning matters a great deal beforehand. The principle is narrower: once the data are observed, the stopping rule contributes nothing about the parameter beyond the likelihood function it produced.
- When does a stopping rule genuinely affect Bayesian inference?When it depends on something outside the model — an unrecorded variable the experimenter peeked at, an external event correlated with outcomes, or the parameter itself in a way the model does not represent. Then the stopping mechanism is informative and no longer factors out as a constant, so it must be modelled explicitly.
- Why can two p-values differ for these same 12 flips?Because a p-value sums the probability of outcomes at least as extreme as the one observed, and that sum is taken over the outcomes the design could have produced. Fixed-12 flips and stop-at-third-tail have different sets of possible outcomes, so the two tail areas are computed over different sample spaces.
saying these in an interview costs you the question
- Claims the two likelihoods are identical rather than proportional
- Says optional stopping distorts a posterior the way it inflates a p-value
- Concludes that stopping rules and design never matter at all
- Thinks the 220 versus 55 factor shifts the posterior slightly
- Cannot say why tail-area calculations depend on the design