skip to content

In a group-sequential A/B test, what does an alpha-spending function do?

level: middleimportance: must knowfreq 55%

answer

  1. the error rate is a budget
  2. spread across the planned looks
  3. cumulative function of information fraction
  4. boundaries solved from the increments
  5. Lan-DeMets frees the look timing

basics

~20 s

An alpha-spending function states in advance how much of the total 0.05 false-positive budget each interim analysis may consume, as a function of how much of the planned data has arrived. The pieces sum to 0.05.

solid answer

~50 s

It is a cumulative budget curve. Write `t` for the information fraction, the share of the planned maximum sample already observed, and define `alpha(t)` as the cumulative probability of having rejected the null by that point when the null is true. It must satisfy `alpha(0) = 0`, `alpha(1) = 0.05` and be non-decreasing. At each interim analysis you spend the increment `alpha(t_k) - alpha(t_{k-1})`, and the critical value for that look is solved numerically so that the chance of first crossing there equals exactly that increment. The shape of the curve is the design choice: a curve that spends almost nothing early gives stringent early boundaries and protects the final analysis, while a curve that spends evenly stops earlier but penalises the end. The Lan-DeMets formulation is what makes this practical, because only the function has to be fixed up front, not the number or the timing of the looks.

go deeper

for a junior

Be able to say that a sequential design analyses the data several times and that the total false-positive budget is divided among those analyses in advance rather than used in full at each one.

for a middle

Explain the three defining properties of the curve, that spending is indexed by information fraction rather than calendar time, and that each look's critical value is derived from the increment the curve allows.

for a senior

Show that you have operated one: information fractions on ramping traffic, adding or moving a look for operational reasons, and refusing to let an informal glance at the estimate decide when the formal analysis happens.

for a principal

Own the choice of curve as a policy decision. Argue how much of the budget the organisation should be able to spend early, given how costly a wrong early launch is versus a slow one, and make it a platform default rather than a per-experiment negotiation.

## What is being budgeted A group-sequential experiment is analysed several times while it is still collecting data: at each of a set of interim analyses you may either stop and declare an effect, or continue. Every one of those analyses is an extra chance for a null result to be declared significant, so the design has to control the probability that **any** of them rejects a true null. That total probability is the error budget, usually 0.05 two-sided. An **alpha-spending function** is the object that allocates that budget across the looks. ## Information fraction Spending is indexed by the **information fraction** `t`, not by calendar time. Information is the reciprocal of the variance of the effect estimate, so for a simple two-arm comparison of means with a fixed allocation it is essentially `n_current / n_max`, the share of the planned maximum sample already collected. At the final analysis `t = 1`. ## The function itself A spending function is a map `alpha(t)` on `[0, 1]` with three properties: - `alpha(0) = 0` — nothing is spent before data arrive; - `alpha(t)` is non-decreasing — you cannot get budget back; - `alpha(1) = alpha` — the whole budget is used by the final analysis. `alpha(t)` is read as the cumulative probability, under the null hypothesis, that the test has stopped and rejected by information fraction `t`. Two classical shapes, both due to the Lan-DeMets construction: - **O'Brien-Fleming-like**: `alpha(t) = 2 * (1 - Phi(z_(alpha/2) / sqrt(t)))`, where `Phi` is the standard normal cumulative distribution function. It is almost flat near zero, so nearly nothing is spent at an early look. - **Pocock-like**: `alpha(t) = alpha * ln(1 + (e - 1) * t)`. It rises quickly and then flattens, spreading the budget much more evenly. Both hit `alpha` exactly at `t = 1`, which is the defining requirement. ## Turning spending into boundaries The boundaries are derived, not chosen. Let `Z_1, ..., Z_K` be the standardised test statistics at the successive looks. Under the null they are jointly normal with a known correlation structure — `corr(Z_j, Z_k) = sqrt(t_j / t_k)` for `j < k` — because later statistics reuse the earlier data. The first critical value solves `P(|Z_1| >= c_1) = alpha(t_1)`. The second solves `P(|Z_1| < c_1, |Z_2| >= c_2) = alpha(t_2) - alpha(t_1)`, and so on. These are multivariate normal probabilities evaluated by recursive numerical integration. The output is a table of nominal cutoffs, each stricter than the fixed-sample 1.96 at some looks and, at the final look, only a little stricter. ## Why the Lan-DeMets form matters in practice The original group-sequential designs required the number of looks and their exact information times to be locked before the first subject enrolled. That is unworkable in product experimentation, where an analysis happens when someone is available and traffic is uneven. The spending-function formulation decouples the two: you commit to the **curve**, and whenever an analysis actually happens you compute the information fraction reached, read the cumulative spending off the curve, and derive the boundary for that look conditional on the boundaries already used. The number of looks can even change mid-flight. The one thing you may not do is let the **observed effect** decide when to look. If an analyst peeks at the estimate informally and then schedules a formal analysis because it looks close, the information times are data-dependent and the error guarantee no longer holds. ## What alpha spending does not give you - It does not make the effect estimate at the stopping moment unbiased. Stopping when the statistic is large selects on noise, so the naive estimate at a boundary crossing is biased away from the null and needs a bias-adjusted or median-unbiased version. - It says nothing about the second error. Stopping early because the effect looks hopeless is a separate budget, spent through a futility boundary or a beta-spending function. - It cannot be applied retroactively. A curve chosen after seeing where the statistic went is not a pre-specification, and the arithmetic is meaningless. ## What a good answer sounds like Name the three defining properties of the curve, say that boundaries are solved from the increments under the joint normal distribution of the sequential statistics, and note the Lan-DeMets flexibility together with its one condition: look timing must not be driven by the observed results.

  • What is the information fraction, and how would you compute it for a two-arm test of means?
    It is the share of the planned maximum statistical information already collected, where information is one over the variance of the effect estimate. For a two-arm comparison of means with a fixed allocation and stable variance it reduces to the observed sample size divided by the planned maximum sample size. It is not elapsed calendar time, which matters when traffic is seasonal or ramping.
  • Can you add a fifth interim analysis to a design that planned four?
    Yes, under a spending-function design. You recompute the boundary at whatever information fraction the new look lands on, spending the increment the curve allows there. What you cannot do is choose to add that look because the current estimate looks promising — data-dependent look timing invalidates the guarantee. Schedule additions on operational grounds, and record the decision before looking.
  • Two spending functions reach 0.05 at t = 1. What differs between them?
    How the budget is distributed over the run, and therefore the whole stopping behaviour. A curve that is nearly flat early yields very stringent early cutoffs, so the experiment almost never stops at the first look but the final analysis is barely penalised. A curve that rises quickly buys realistic early stopping at the cost of a harsher final cutoff and lower power at full data.

It is a spending plan for a fixed budget: you decide in advance what share of the 0.05 each check-in may draw down, and the final analysis lives on whatever is left.

saying these in an interview costs you the question

  • Says the total alpha is applied at every look independently
  • Treats alpha spending as a correction applied after the fact
  • Confuses the information fraction with elapsed calendar time
  • Thinks the spending function controls power rather than false positives
  • Believes the estimate at a boundary crossing needs no adjustment

context