Does checking P(B > A) every morning inflate error the way repeated p-value checks do?
answer
- the posterior is a snapshot, not a procedure
- no correction to apply to the number itself
- the stopping rule is the thing with behaviour
- early stops select lucky data
- magnitude-aware rules are hard to trip on noise
basics
~20 sThe posterior itself needs no correction for how often you look: it summarises the data in hand at any moment. But a rule that stops the first time the probability crosses a bar still selects lucky data and overstates the lift.
solid answer
~50 sSeparate the quantity from the rule. The posterior probability that B beats A conditions only on the data observed so far, so recomputing it daily neither corrupts it nor calls for any correction. The rule built on top of it is a different object. "Stop the first morning the probability exceeds 0.95" is a procedure, and procedures have operating characteristics: the chance of crossing the bar at some point in a long test exceeds the chance of being above it at any single moment, and the effect estimate at the stop is biased away from zero, because you stopped on a favourable fluctuation. The defence is the shape of the rule, not a correction to the posterior — magnitude-aware criteria such as an expected-loss tolerance shrink toward zero as data accumulates and are hard to trip on noise, plus a sceptical prior and a minimum run covering the weekly cycle.
go deeper
Know that recomputing a posterior as data arrives is legitimate and needs no adjustment, but that stopping the instant a number looks good is a separate decision with its own risks.
Explain why the posterior needs no correction — it conditions only on the data observed — and why a threshold-crossing stopping rule still has behaviour that depends on how often you evaluate it.
Name the concrete exposures, especially the inflated effect estimate at the moment of stopping, and describe the practical defences: magnitude-aware criteria, a sceptical prior, a minimum run covering the weekly cycle.
Own the consequence at portfolio scale: a program that stops aggressively books wins that never show up in the top-line metric. Decide the monitoring policy, the priors, and which launches get a holdback.
## Two different objects This question separates candidates who have memorised "Bayesian methods are immune to peeking" from those who understand why the claim is half true. **The posterior is a snapshot.** It is defined as the prior combined with the likelihood of the data you have. It carries no assumption about how much data you intended to collect, and it does not change depending on how many times you have looked at it. If you recompute the probability that B beats A every morning, each morning's number is a correct summary of the evidence available that morning. There is nothing to correct, no multiplicity adjustment to apply, and applying one would be a category error. **A stopping rule is a procedure.** "Ship the first morning the probability of superiority exceeds 0.95" is not a summary; it is an algorithm that consumes a data stream and emits decisions. Algorithms have long-run behaviour, and that behaviour depends on how often you evaluate them. Nothing about being Bayesian exempts a procedure from having operating characteristics. ## What the daily check actually costs you Two concrete exposures, and a good answer names both. **Crossing on noise.** The probability that a threshold is crossed *at some point* during a long run is larger than the probability of finding it crossed at one arbitrary moment, simply because more opportunities exist. If two arms are genuinely identical, the probability of superiority wanders — it is a random walk driven by accumulating noise — and given enough looks it will visit high values. The more often you check and the earlier you allow yourself to stop, the more of that wandering you expose yourself to. **Magnitude inflation at the stop.** This is the cost people underrate. If you stop precisely when the evidence looks most favourable, the evidence in front of you at that moment is favourable partly by luck. The posterior mean lift at the moment of stopping therefore tends to overstate the true lift, and the earlier you stop the worse it is. The decision to ship may still be right; the number you put in the launch summary is optimistic, and it is that number the business will use to forecast, to prioritise the next quarter's roadmap, and to judge whether the program is working. Programs that stop aggressively accumulate a portfolio of launches whose claimed wins do not add up to the observed movement in the top-line metric. ## What actually protects you None of the fixes involve correcting the posterior. - **Decide on magnitude, not on direction.** An expected-loss tolerance in metric units, or a practical-equivalence band, both require the posterior to have become genuinely concentrated before they fire. Direction can wander cheaply on thin data; "the expected cost of being wrong is under two hundredths of a point of conversion" cannot be satisfied by a lucky morning early in the test. This is the single most effective protection, and it costs nothing to adopt. - **Use a prior that is actually sceptical.** A flat prior lets a handful of early conversions swing the posterior hard. A prior centred on no effect with a width reflecting the historical distribution of true effects — most product changes do very little — damps early excursions and materially reduces stopping on noise. The cost is slower detection of genuine effects, which is a tradeoff to state, not to hide. - **Set a minimum run length for non-statistical reasons.** Weekday and weekend populations differ, and so do the days after a marketing push. A test stopped on Tuesday morning has measured Tuesday-ish traffic. A full weekly cycle is a business requirement independent of any statistics, and it happens to remove the most damaging early looks as a side effect. - **Hold back and verify the ones that matter.** For a launch whose claimed effect drives planning, keep a holdback and re-measure it later. That checks the magnitude rather than the direction, which is precisely what early stopping distorts. ## The honest summary Say it in three beats: the posterior needs no correction for looking, because it conditions only on the data in hand; the stopping rule is a separate object whose behaviour does depend on how often you look, and stopping early biases the reported effect upward; and the defence is a magnitude-aware criterion plus a sceptical prior and a minimum run length, not an adjustment to the posterior. The confident "Bayesian methods have no peeking problem, full stop" is the answer that fails.
- So what does the daily check still expose you to?Two things. A higher chance of crossing the threshold at some point during a long run purely on noise, because each look is another opportunity. And an inflated effect estimate at the moment of stopping, since you stop when the data looks most favourable. The direction of the call may survive; the magnitude you report will tend to be optimistic.
- Does a stronger prior help, and what does it cost?Yes. A prior centred on no effect with a realistic width damps early excursions, so the probability of superiority is far less likely to spike on a handful of lucky conversions. The cost is slower detection of genuine effects and a more conservative program overall. That is a legitimate tradeoff, and the prior should be stated in the readout so nobody mistakes it for neutrality.
- Should any correction be applied to the posterior for the number of looks?No, and proposing one signals a misunderstanding. The posterior conditions on the data observed, not on the observer's schedule, so there is no quantity being spent by looking. If you want to control how often a rule fires on noise, change the rule — raise the bar, require a magnitude, impose a minimum run length — rather than adjusting the posterior.
Reading a thermometer hourly does not damage the thermometer. Deciding to leave the house the first hour it reads warm is a rule, and that rule will send you out on a freak sunny hour.
saying these in an interview costs you the question
- Claims Bayesian analysis is immune to peeking, full stop
- Wants to correct the posterior for the number of looks
- Reports the lift at the moment of stopping as unbiased
- Stops the first morning the bar is crossed regardless of magnitude
- Uses a flat prior and then stops very early on thin data