Would you make always-valid sequential testing the default on an experimentation platform?
answer
- match the method to how results are read
- validity under real behaviour, not policy
- efficiency is the currency you pay
- tier the default by traffic volume
- early stops overstate the effect
basics
~20 sUsually yes on a self-serve platform, because people read results whenever they like and sequential methods stay honest under that behaviour. The tradeoff is efficiency: the same conclusion needs more data than a fixed-horizon design.
solid answer
~50 sThe decision follows from how results are actually consumed. If dozens of teams read a live dashboard daily, a method whose validity depends on nobody looking is a fiction, and anytime-valid inference makes the platform's stated guarantee true. The cost should be quoted: confidence sequences are wider at every sample size, and a group-sequential boundary inflates the maximum sample needed for a given power — a few percent for an O'Brien-Fleming shape, roughly 15% to 25% for a Pocock shape with several looks. So I would tier it. High-traffic surfaces get anytime-valid readouts by default, where the efficiency loss costs days rather than viability. Low-traffic surfaces keep a fixed design with a locked read-out and a few pre-planned looks. Either way, pair it with bias-adjusted estimates for early stops and copy that distinguishes not yet conclusive from no effect.
go deeper
Know that sequential methods let results be read during a run while fixed-horizon designs assume one planned read-out, and that the flexibility is paid for with data.
Be able to state the efficiency cost concretely: wider intervals at every sample size, and a larger maximum sample for the same power, with the size of the penalty depending on the boundary shape.
Show operational judgement: choosing per surface based on traffic and how results are consumed, enforcing adjusted estimates on early stops, and refusing a mid-flight method switch.
Own the platform policy end to end — default boundary shape, tiering by traffic, futility rules, adjusted reporting and stakeholder language — and defend the aggregate efficiency cost to the business as the price of guarantees that hold under real behaviour.
## Frame the decision correctly This is not a question about which method is statistically superior; both control error when used as specified. It is a question about which specification survives contact with the organisation. A fixed-horizon design is more efficient per unit of data, but its guarantee is conditional on a discipline: nobody acts on the result before the planned read-out. A sequential or anytime-valid design gives away some efficiency to buy a guarantee that holds under continuous reading. The right default is whichever assumption is true about your company. ## Arguments for making it the default - **The consumption pattern is the real design constraint.** On a self-serve platform, results are read by product managers, executives and engineers on whatever day they care. A method that stays valid under that behaviour is describing reality; one that assumes a single read-out is describing a policy nobody follows. - **Faster kills and faster wins.** A boundary crossed early ends a clear winner or a clear disaster sooner, freeing the surface. With a futility rule, hopeless experiments end early too, which compounds across a portfolio. - **One artefact, no exceptions process.** Everything on the dashboard means the same thing all the time, and a truncated run — a launch date moved, an incident, a holiday — still yields a valid readout instead of a case-by-case judgement. - **It removes an argument.** Nobody has to litigate whether it was acceptable to look, which is a governance win as much as a statistical one. ## Arguments against, and where they bite - **Efficiency.** Wider intervals and larger maximum samples mean some effects that a fixed design would have resolved never resolve. On a low-traffic surface where an experiment already runs for weeks, a 20% inflation can turn a feasible experiment into an infeasible one. - **Not every decision is monitored.** A quarterly pricing test read once, by one team, on one date, does not need anytime validity, and paying for it is waste. - **Early stopping distorts effect magnitudes.** Estimates at a crossing are biased away from the null, and organisations that stop early accumulate a portfolio of overstated wins unless adjusted estimates are enforced. - **Comprehension cost.** Sequential boundaries and confidence sequences are harder to explain than a single number at the end, and a method stakeholders misread is not obviously an improvement. ## A defensible policy 1. **Tier by traffic.** High-traffic surfaces: anytime-valid by default, since the efficiency loss costs days. Low-traffic surfaces: a fixed design with a locked read-out plus a small number of pre-planned interim looks under a stringent early boundary, which recovers most of the safety at a fraction of the cost. 2. **Pick a house boundary shape.** An O'Brien-Fleming-style spending function as the default, because its inflation is a few percent and its early looks exist mainly as an escape hatch for extreme results. Reserve flatter, Pocock-style shapes for experiments where finishing early is worth a lot. 3. **Always include futility, non-binding.** Non-binding so that a team can argue to continue without breaking the false-positive guarantee, which is what would happen in practice anyway. 4. **Report adjusted estimates on early stops.** Show a bias-adjusted effect, or lead with the conservative end of the interval when sizing business impact. 5. **Fix the language.** An experiment that has not crossed a boundary is *not yet conclusive*. Copy that says *no effect* converts a bounded-power result into a false claim about the world. 6. **Constrain look scheduling.** For group-sequential designs, looks are scheduled on operational grounds and recorded before the analysis, never chosen because the estimate looked close. ## How to argue it in an interview The weak answer picks a side on aesthetics. The strong answer names the consumption pattern as the deciding variable, quotes the efficiency cost with an actual magnitude, tiers the policy rather than imposing one rule, and closes the loop on the second-order problems — biased effect sizes at early stops and stakeholder language — that adopting sequential methods creates and that nobody else in the room will mention.
- When is a fixed-horizon design still clearly the right call?When looking early is genuinely impossible or pointless: a single decision read once on a fixed date, a low-traffic surface where every point of sensitivity matters, or an offline evaluation with no live dashboard. Fixed-horizon intervals are tighter at the same data volume, so if the discipline is real rather than aspirational, taking the efficiency is the correct trade.
- How would you stop a portfolio of early-stopped experiments from overstating cumulative impact?Require bias-adjusted estimates for any experiment that stopped at a boundary, since crossing selects favourable noise. Size business impact from the conservative end of the interval rather than the point estimate, and reconcile a sample of shipped wins against a holdout so the aggregate claim is checked against reality rather than against the sum of individual readouts.
- A team asks to switch mid-experiment from a fixed design to a sequential one after seeing early data. What do you say?No. Both methods control error only under a specification fixed before the data were seen, and choosing the method because of what the early estimate showed is exactly the selection those methods are built to prevent. The honest options are to run the fixed design to its planned read-out, or to abandon it and restart under a pre-registered sequential design.
saying these in an interview costs you the question
- Presents sequential methods as free with no efficiency cost
- Applies one rule to every surface regardless of traffic
- Ignores that early stops bias reported effect sizes
- Lets stakeholders read not yet conclusive as no effect
- Allows switching methods after seeing interim data