skip to content

What does a futility boundary add to a group-sequential experiment, and what does it cost?

level: seniorimportance: nice to knowfreq 26%

answer

  1. a second reason to stop early
  2. not about winning, about giving up
  3. conditional power or beta spending
  4. the currency spent is power
  5. binding versus non-binding overrides

basics

~20 s

A futility boundary stops an experiment early when the accumulated data make crossing the efficacy boundary implausible, freeing traffic and time. The cost is power: some experiments that would eventually have reached significance are killed before they get there.

solid answer

~50 s

A futility boundary is a second, inner boundary on the same standardised scale: cross the outer one and you stop for success, fall below the inner one and you stop for futility. It is usually derived from **conditional power** — the probability of a significant final result given the data so far under an assumed effect — with a threshold such as 20%, or from a beta-spending function that budgets Type II error across looks. The distinction that matters is binding versus non-binding. A **non-binding** boundary leaves the efficacy cutoffs computed as if the experiment might continue regardless, so overriding it does not inflate the false-positive rate. A **binding** boundary assumes you really will stop, which permits slightly less stringent efficacy cutoffs but breaks the guarantee if you override it. Non-binding is the safer default in product work, because someone eventually wants to keep a promising experiment alive.

go deeper

for a junior

Know that a sequential design can stop early for two opposite reasons, a clear win and a clearly hopeless result, and that the second one is called futility.

for a middle

Explain how the boundary is derived from conditional power or a beta-spending budget, and that the error it spends is the false-negative one rather than the false-positive budget.

for a senior

Demonstrate judgement in the room: refusing a futility call at a tiny information fraction, naming the assumed effect behind a conditional-power number, and writing the result up as inconclusive rather than null.

for a principal

Own the policy. Decide how aggressive the house futility rule should be given what held traffic costs, whether rules are binding, and how many promising experiments the organisation is willing to lose for faster turnover.

## The second reason to stop An efficacy boundary answers one question at each interim analysis: is the effect large enough to stop and declare a win, or a harm? A futility boundary answers the complementary one: is the effect so unpromising that continuing is a waste of traffic? Without a futility rule, a group-sequential design will run to its maximum sample even when the estimate has been flat at zero from the first look. That is expensive: traffic held on an inert change, a decision deferred, and a slot on the surface that another idea could have used. ## How the boundary is derived Two standard constructions. **Conditional power.** At an interim analysis, conditional power is the probability that the final analysis will cross the efficacy boundary, given the data observed so far and an assumption about the true effect going forward. The assumption matters enormously and has to be stated. Common choices: - the effect the experiment was designed to detect (optimistic, gives a lenient futility rule); - the currently observed effect (much more pessimistic once the estimate is near zero); - a value in between, or a range reported as a sensitivity band. A rule such as *stop for futility if conditional power under the designed effect falls below 20%* is typical. **Beta spending.** The mirror of alpha spending. A beta-spending function allocates the total Type II error budget across the looks, and the futility boundary at each look is the value below which the cumulative probability of a false negative reaches the allowed amount. This is the more formal route and slots naturally into a spending-function implementation. ## Binding versus non-binding This is the distinction interviews probe, and it is easy to state backwards. - **Binding futility.** The design assumes the experiment stops whenever the futility boundary is crossed. That assumption removes sample paths that could otherwise have come back and crossed the efficacy boundary, so the efficacy cutoffs can be relaxed slightly while still hitting the nominal error rate. The obligation is real: if you cross the futility boundary and continue anyway, the false-positive rate exceeds the nominal level, because the efficacy boundary was priced on the assumption you would have quit. - **Non-binding futility.** The efficacy boundary is computed ignoring the futility rule entirely, as if the experiment always continues. The futility boundary is then advisory. Overriding it and running to the end costs nothing in false-positive terms; you simply did not take the savings the rule offered. In regulated trials binding boundaries appear where the stopping commitment can genuinely be enforced. In product experimentation, where a team can and will argue to keep an experiment running, non-binding is the sane default: it preserves the guarantee under human behaviour rather than assuming it away. ## The costs, stated honestly - **Power.** Every futility rule kills some experiments that would have crossed the efficacy boundary at full data. A lenient rule costs a little power; an aggressive one costs a lot. If you want to hold the design's power constant after adding futility, the maximum sample has to grow to compensate. - **Assumption sensitivity.** A conditional-power rule computed under the observed effect is far more trigger-happy at an early look than the same rule computed under the designed effect, because early estimates are noisy and often near zero. Reporting which assumption was used is part of reporting the decision. - **Interpretation.** Stopping for futility is not evidence that the effect is zero. It is a statement that this experiment, with this remaining budget, is unlikely to demonstrate the effect it was designed to find. A small real effect may well exist below the resolution of the design, and the write-up should say so rather than record a proven null. - **Timing.** Futility rules at very early looks are the ones most likely to be regretted, since conditional power at a small information fraction is dominated by noise. Many designs deliberately start futility monitoring only after a meaningful share of the information has arrived. ## What a strong answer includes Define the boundary, name conditional power or beta spending as the derivation, get the binding versus non-binding direction right — non-binding may be overridden without inflating false positives, binding may not — and be explicit that the currency spent is power, not the false-positive budget.

  • Conditional power at an interim look is 12%. What assumption do you need before acting on that?
    Which effect the projection assumes going forward. Under the effect the experiment was designed to detect, 12% is genuinely discouraging. Under the currently observed effect, a low value is nearly automatic once the estimate sits near zero, especially at a small information fraction. State the assumption, and ideally report conditional power under both before recommending a stop.
  • Does adding a futility boundary change the efficacy boundary?
    Only if it is binding. A binding rule lets the design assume that unpromising paths are removed, which permits slightly less stringent efficacy cutoffs but obliges you to stop. A non-binding rule leaves the efficacy boundary exactly as it was, so you may override it without inflating the false-positive rate, at the price of forgoing the efficiency gain.
  • How should a futility stop be written up for stakeholders?
    As an inconclusive result, not a proven null. Say that the experiment was unlikely to demonstrate the effect it was powered for within its remaining budget, report the estimate with its interval so the range of effects still compatible with the data is visible, and note that a smaller real effect could exist below the design's resolution.

saying these in an interview costs you the question

  • Treats a futility stop as proof the effect is zero
  • Says futility stopping inflates the false-positive rate
  • Gets binding and non-binding the wrong way round
  • Quotes conditional power without naming its assumed effect
  • Adds an aggressive futility rule without accounting for lost power

context