What guarantee does an always-valid confidence sequence give that a fixed-horizon confidence interval does not?
answer
- coverage at one time versus all times
- the quantifier sits inside the probability
- any stopping rule is covered
- test martingales and Ville's inequality
- the price is extra width
basics
~20 sA confidence sequence covers the true parameter at every sample size simultaneously with probability at least 1 minus alpha, so it can be read or stopped at any moment. A fixed-horizon interval guarantees coverage only at one pre-specified sample size.
solid answer
~50 sWrite the parameter as `theta` and the interval after `t` observations as `CI_t`. A 95% fixed-horizon interval at a planned `n` guarantees `P(theta not in CI_n) <= 0.05` for that single `n`. A confidence sequence guarantees the far stronger uniform statement `P(there exists any t with theta not in CI_t) <= 0.05` — one promise covering the whole trajectory. So you may read it whenever you like, stop for any reason including a data-driven one, and the interval at that moment is still valid. The construction is a test martingale: a mixture sequential probability ratio test averages the likelihood ratio over a distribution of effect sizes, giving a nonnegative process with expectation one under the null, and Ville's inequality bounds the chance it ever exceeds `1/alpha`. The price is width — at the same sample size the sequence is strictly wider, since it must survive every future look.
go deeper
Recall that a fixed-horizon interval is only guaranteed at the sample size you planned, while a confidence sequence is designed to be read at any moment during the run.
State the two coverage statements precisely and explain why the uniform one implies validity at any stopping time, including a stopping rule driven by the data itself.
Show you have shipped one: choosing it for a live dashboard, quantifying the width penalty against tighter fixed-horizon intervals, and correcting the estimate reported after an early stop.
Decide where anytime-valid inference belongs in the platform and what it costs the company in data volume, and set the reporting standard for effect magnitudes from experiments that stopped early.
## Two different promises The distinction is about **when** the coverage promise applies. A fixed-horizon 95% confidence interval is a procedure attached to one sample size `n`, chosen before data collection. Over repeated runs of the whole experiment, the interval computed at that `n` contains the true parameter `theta` in at least 95% of runs. It says nothing about the intervals you could have computed at other sample sizes along the way, and a rule that stops when the interval happens to exclude a value is not covered by the guarantee. A **confidence sequence** is an entire sequence of intervals `CI_1, CI_2, ...`, one per observation or per batch, satisfying the uniform statement `P(exists t : theta not in CI_t) <= alpha`. The quantifier is inside the probability. In at least 95% of runs, **every** interval in the whole sequence contains `theta`. ## Why the uniform version buys anytime validity Because the promise covers all times at once, it also covers any time you might choose to stop, including a time chosen by looking at the data. Formally, for any stopping rule `tau`, `P(theta not in CI_tau) <= alpha`, since the failure event at `tau` is contained in the event that some interval failed. This is what people mean by anytime-valid or always-valid inference: validity is no longer conditional on having pre-committed to a sample size. ## How they are built The machinery is supermartingale-based rather than central-limit-based. - An **e-value** is a nonnegative statistic whose expectation under the null is at most one. By Markov's inequality, `P(E >= 1/alpha) <= alpha`, so a large e-value is evidence against the null at level alpha, and `1/E` behaves like a conservative p-value. - An **e-process** or test martingale is a sequence of such quantities that keeps this property as data accrue. **Ville's inequality** — the martingale analogue of Markov's — bounds the probability that the process ever reaches `1/alpha` by alpha, for the whole infinite trajectory. - A **mixture sequential probability ratio test** builds one concretely. A plain likelihood ratio needs a specific alternative; the mixture version integrates the likelihood ratio over a distribution across candidate effect sizes, giving a single process valid without committing to one alternative. Inverting it — collecting all parameter values the process never rejects — produces the confidence sequence. A second family is built from **time-uniform boundaries** derived from the law of the iterated logarithm, which is why widths often carry a `sqrt(log log t / t)` term. ## The cost Anytime validity is not free, and quantifying the cost is what separates a real answer from a slogan. - **Width.** At any given sample size the confidence sequence is wider than the fixed-horizon interval for the same data. It has to be: it is protecting against every future look, not one. - **Asymptotic rate.** A fixed-horizon half-width shrinks like `1 / sqrt(n)`. A time-uniform one shrinks like `sqrt(log log n / n)`, an extra slowly-growing factor. The gap narrows in relative terms as data accumulate but never closes. - **Tuning.** With a mixture construction, the mixing distribution is a design knob. Concentrating it near the effect sizes you actually care about tightens the sequence there and loosens it elsewhere; a badly chosen spread wastes power on effects nobody would ship. ## Practical consequences - **Dashboards.** A confidence sequence is the right object to render on a live experiment dashboard, because a reader who checks it daily is doing exactly what it was built to permit. - **Stopping for external reasons.** A launch date moves, a bug forces a shutdown, a holiday truncates traffic. Under a confidence sequence the interval at that moment is valid; under a fixed-horizon design an early truncation forces a judgement call. - **The point estimate is still biased.** Uniform coverage of the interval does not de-bias the centre. If you stopped because the sequence excluded zero, the estimate at that moment is on the favourable side, and shipping decisions based on its magnitude should use a bias-adjusted version. - **It does not license changing the metric.** Anytime validity concerns repeated looks at one pre-specified quantity, not swapping in whichever metric currently looks good. ## Answering it well State the two probability statements side by side, note that the quantifier over time sits inside the probability in one and outside in the other, name the martingale mechanism and Ville's inequality, and be honest about the width cost and the residual estimate bias. Claiming anytime validity is free is the fastest way to lose the point.
- What defines an e-value, and why is it useful under repeated analysis?An e-value is a nonnegative statistic whose expectation under the null is at most one. Markov's inequality then gives `P(E >= 1/alpha) <= alpha`, so a threshold at `1/alpha` is a valid test. Because e-values multiply across independent batches and their martingale versions obey Ville's inequality over the whole trajectory, they stay valid however many times you look.
- If a confidence sequence is always wider, why not run fixed-horizon and just be disciplined?Discipline is the right choice when it is achievable: a fixed-horizon design at the same data volume gives tighter intervals. The case for a confidence sequence is organisational — self-serve dashboards, stakeholders who read results daily, and shutdowns for reasons unrelated to the data. You buy validity under real behaviour, and pay for it in width.
- You stop the moment the sequence excludes zero. Is the point estimate at that moment unbiased?No. The interval is valid, but stopping at the first exclusion selects moments where noise favoured the effect, so the naive estimate is biased away from zero, most severely on early stops. Report a bias-adjusted estimate, or treat the boundary of the interval nearest zero as the defensible magnitude when sizing the business impact.
- How does the mixing distribution in a mixture SPRT affect the result?It sets where the procedure is most sensitive. Concentrating the mixture on the effect sizes you would actually act on tightens the sequence in that range and loosens it far away, while a very diffuse mixture spreads sensitivity thinly and delays detection everywhere. It must be fixed in advance; tuning it after seeing data destroys the guarantee.
A fixed-horizon interval is a ticket valid only for one specific train. A confidence sequence is a pass valid on every train that day, and costs more for exactly that reason.
saying these in an interview costs you the question
- Says each interval covers with 95% probability one at a time
- Claims anytime validity comes at no cost in width
- Thinks the guarantee also de-biases the point estimate
- Confuses a confidence sequence with a prediction interval
- Believes it licenses swapping the metric mid-experiment