skip to content

How do O'Brien-Fleming and Pocock group-sequential boundaries differ?

level: middleimportance: should knowfreq 45%

answer

  1. the shape of the cutoff over time
  2. one boundary is flat, one decreasing
  3. stringent early protects the final analysis
  4. a constant cutoff buys early stopping
  5. the price shows up in end-of-test power

basics

~20 s

O'Brien-Fleming boundaries are extremely stringent at early looks and relax to nearly the fixed-sample cutoff at the end. Pocock uses one constant, looser cutoff at every look, so it stops earlier but gives up power at the final analysis.

solid answer

~50 s

Both control the same overall false-positive rate; they differ in the shape of the critical values over time. An O'Brien-Fleming boundary scales roughly as `sqrt(K / k)` at look `k` of `K`, so with four planned looks the first cutoff sits above 4 on the standardised scale — essentially uncrossable — and the final cutoff is about 2.02, barely above the fixed-sample 1.96. A Pocock boundary is a single constant applied at every look, around 2.36 for four looks at two-sided 0.05, a nominal per-look level near 0.018. The tradeoff is real: Pocock can genuinely stop at look one or two for a moderate effect, but if the signal only emerges at the end you must clear 2.36 rather than 1.96, so power at full data is noticeably lower. O'Brien-Fleming buys almost no early stopping and loses almost nothing at the end, which is why it is the more common default.

go deeper

for a junior

Recall that both are tables of critical values for repeated analyses, and that one keeps the same cutoff throughout while the other starts far stricter and eases off.

for a middle

Explain the shapes and their consequences in numbers: a constant cutoff near 2.36 for four looks versus a decreasing sequence starting above 4 and ending near 2.02, and what each does to end-of-test power.

for a senior

Argue the choice from the experiment's economics, handle a run that finishes just under a Pocock cutoff without renegotiating the rule, and recompute boundaries when realised information fractions drift from the plan.

for a principal

Set the house shape and defend it. Decide what early stopping is worth to the business against the sensitivity lost at full data, and whether teams may deviate per experiment or must justify it.

## The same guarantee, two shapes Both families are group-sequential boundaries: a table of critical values, one per planned analysis, constructed so that the probability of crossing at **any** of them under the null equals the nominal level. What separates them is how the error budget is distributed across time, which shows up as the shape of the boundary on the standardised scale. ## Pocock: a flat boundary Pocock's boundary uses the **same** critical value `c` at every look, with `c` solved so that the overall crossing probability equals the nominal level. It grows with the number of looks: roughly 2.18 for two looks, 2.29 for three, 2.36 for four, and 2.41 for five, at two-sided 0.05. Read as a per-look nominal level, 2.36 corresponds to about 0.018 rather than 0.05 — noticeably stricter than a one-shot test, but reachable. Consequences: - Early stopping is a genuine possibility. A moderate true effect has real probability of pushing the statistic past 2.36 at the first or second look. - The final analysis is penalised for the whole run. If the effect is real but small, so that the statistic only reaches, say, 2.1 at full data, the experiment is declared inconclusive even though a fixed-horizon test at 1.96 would have rejected. - Holding power fixed, the design needs a meaningfully larger maximum sample than a fixed-horizon design — on the order of 15% to 25% more with several looks. ## O'Brien-Fleming: stringent early, near-normal at the end The O'Brien-Fleming boundary makes the critical value a decreasing function of the information fraction: with `K` equally spaced looks the cutoff at look `k` scales roughly as `c * sqrt(K / k)`. With four looks at two-sided 0.05 the sequence runs approximately 4.05, 2.86, 2.34, 2.02. Consequences: - The first look is effectively unreachable. Only a spectacular effect — or a catastrophic regression, in the harm direction — will cross it, which is exactly the intent: early data are noisy, and an early stop on noise is expensive. - Almost none of the budget is consumed early, so the final cutoff, about 2.02, is barely stricter than the fixed-sample 1.96. Power at full data is close to that of a fixed design, and the maximum sample inflation is a few percent. - Early stopping does happen, but essentially only for large effects, which is the case where stopping early is most obviously right. ## Choosing between them The question to ask is what an early stop is worth relative to a missed small effect. - Choose an O'Brien-Fleming shape when the endpoint decision matters most and you mainly want an escape hatch for enormous wins or losses. This is the usual default in product experimentation, where small effects are common and abandoning them because the boundary was flat is costly. - Choose a Pocock shape when the cost of running is high, effects are expected to be large if they exist, and finishing weeks earlier is worth losing sensitivity at the end. - Anything in between is available: a power-family spending function with a tunable exponent interpolates between the two shapes, so you are not restricted to the two classical tables. ## Details that separate a good answer - Both boundaries have a harm-direction counterpart. A two-sided boundary can be crossed downward, which is how a sequential design stops a damaging change early. - The boundaries are computed for the **standardised** statistic under a known correlation structure between successive looks. If the actual information fractions drift from the plan — traffic arrives faster or slower than expected — a spending-function implementation recomputes the boundary at the realised fractions rather than reusing the original table. - Under either boundary, the point estimate at the moment of stopping is biased away from the null, because crossing selects for a favourable draw. A design that stops early should report a bias-adjusted estimate alongside the decision, and the earlier the stop, the larger the correction. - Neither shape is more or less honest than the other. They control exactly the same overall error rate; they trade when that error may be spent. ## What interviewers listen for The strongest answers give the shape in one sentence each, quote the direction of the tradeoff correctly — Pocock stops earlier and pays at the end, O'Brien-Fleming almost never stops early and pays almost nothing at the end — and then tie the choice to the economics of the experiment rather than declaring one boundary universally superior.

  • Why is the O'Brien-Fleming final critical value only slightly above the fixed-sample 1.96?
    Because almost none of the error budget was spent at the earlier looks. The early cutoffs are so high that the probability of crossing them under the null is tiny, leaving nearly the whole 0.05 available for the final analysis. That is the whole design intent: an escape hatch for extreme results with a final test that behaves almost like the fixed-horizon one.
  • Your Pocock-boundary experiment ends with a standardised statistic of 2.1. What do you report?
    Inconclusive by the pre-registered rule, since 2.1 is below the roughly 2.36 constant cutoff. You report the estimate with its sequential interval and note that the design deliberately traded end-of-test sensitivity for early-stopping ability. What you do not do is switch to the 1.96 cutoff after the fact, which discards the design's guarantee entirely.
  • Can you get a shape between the two classical boundaries?
    Yes. A power-family spending function with a tunable exponent interpolates continuously: a small exponent gives an O'Brien-Fleming-like curve that spends almost nothing early, a larger one approaches even spending and Pocock-like behaviour. You pick the exponent from how much early-stopping ability the decision is worth, rather than accepting one of two fixed tables.

saying these in an interview costs you the question

  • Says one boundary controls the error rate better than the other
  • Claims O'Brien-Fleming stops earlier because it is more sensitive
  • Assumes the final cutoff is always 1.96 under any boundary
  • Ignores that Pocock costs power at the final analysis
  • Reports the raw estimate after an early stop without adjustment

context