An A/B variant read +12% lift at a 1% ramp and +2% at 50% — what explains the drop?
answer
- huge standard error at tiny traffic
- you selected on the number being big
- unbiased overall, inflated once selected
- regression to the mean, not decay
- Type-M exaggeration grows as power falls
basics
~20 sMost likely the winner's curse. The 1% estimate was extremely noisy, and the feature was promoted because that noisy number looked big. Selecting on a large observed effect inflates it, so the precise later read of +2% is the better estimate, not a decay.
solid answer
~50 sAt 1% traffic the standard error of the lift is enormous — plausibly comparable to the lift itself — so individual reads scatter widely around the truth. Promotion happened *because* the read was big, which means the reported number was conditioned on being large. Conditional on that selection, an estimator that is unbiased in general is biased upward in magnitude; this is the winner's curse, or Type-M (magnitude) error, and it grows as power falls. The +2% at 50% comes with a far smaller standard error and is much closer to the truth; the shrinkage is regression to the mean, not the feature getting worse. Before concluding that, rule out the real alternatives: the exposed population may have changed between stages, the treatment may have been patched, or the effect may genuinely differ at scale. The operational fix is to decide in advance which stage produces the number you report.
code
python · 15 linesimport random, statistics
BASE, TRUE_LIFT, N = 0.10, 0.02, 1200 # per-arm users at a tiny ramp stage
def one_read():
ctrl = sum(random.random() < BASE for _ in range(N)) / N
trt = sum(random.random() < BASE * (1 + TRUE_LIFT) for _ in range(N)) / N
return (trt - ctrl) / ctrl
reads = [one_read() for _ in range(2000)]
promoted = [r for r in reads if r > 0.08] # only big early reads get promoted
print(len(promoted), "of", len(reads), "reads looked big")
print("all reads:", round(statistics.mean(reads), 3)) # near the true 0.02
print("promoted only:", round(statistics.mean(promoted), 3)) # far above 0.02go deeper
Remember that small samples give wildly variable estimates, and that a number from a tiny slice of traffic is not a reliable measurement of a small effect.
Explain the mechanism: the estimator is unbiased in general, but conditioning on a large observed value makes the reported figure an overstatement, and the exaggeration grows as precision falls.
Show the diagnosis, not just the label — rule out population shift, a patched treatment and genuine scale effects before attributing the drop to selection, and say which stage's number you would report.
Own the organisational consequence: if launches are approved on early reads, claimed lifts will not sum to observed company-metric movement, and the fix is a reporting standard, not a per-team argument.
## The phenomenon An estimate from a small ramp stage that looks spectacular and then shrinks as traffic grows is one of the most reliably reproduced patterns in experimentation. It has a name — the **winner's curse** — and a precise mechanism. ## Noise at 1% Start with the sampling variability. For a binary metric with baseline rate p and n users per arm, the standard error of a single arm's rate is sqrt(p(1-p)/n), and the standard error of the difference is roughly sqrt(2) times that. Expressed as a *relative* lift, divide by p. With a 10% baseline and a small early stage, the standard error of the relative lift is easily ten percentage points or more. That means a true effect of +2% produces observed reads spread across a range from clearly negative to well above +15%, purely from chance. Nothing is broken; that is what the estimator does at that sample size. ## Selection is what turns noise into bias The estimator of the lift is (approximately) unbiased: average it over all possible experiments and you recover the truth. But nobody averages over all possible experiments. The feature got promoted, celebrated and written up **because** its early read was large. You are looking at a number drawn from the right-hand tail of that wide distribution, selected for being there. Once you condition on "the observed effect exceeded some bar", the conditional expectation of the estimate exceeds the true effect. Formally, if the estimate is roughly normal with mean equal to the true effect and standard error SE, then E[estimate | estimate > c] is larger than the true effect, and the gap grows as SE grows relative to the true effect. Gelman and Carlin call the ratio between the expected selected magnitude and the truth the **Type-M (magnitude) error**, or exaggeration ratio, and pair it with **Type-S (sign) error**: at very low power, a selected result can even have the wrong sign. The headline is that low power does not merely make a result unlikely to be found — it makes any result you *do* find an overstatement. ## Why the 50% read is the better number At 50% traffic the standard error is much smaller, and — crucially — that read was not selected for being large; you took it because the ramp schedule said to. An unselected, precise estimate is what you want. So the honest summary is: the effect was probably always about +2%, the early stage happened to draw a high number, and the later stage regressed toward the truth. It is a property of the measurement process, not a claim that anything degraded. ## Ruling out the alternatives A senior answer does not stop at "winner's curse", because three other explanations produce the same shape and demand different actions. - **Population shift.** Early ramp stages are often restricted — one country, internal users, a low-risk platform, or simply the most active users because they trigger the feature soonest. If the later stage reaches a broader population with a different baseline, the two stages estimate effects on different groups. Diagnose by re-estimating the later stage restricted to the earlier stage's segment; if the effect there is still large, the drop is composition, not selection. - **The treatment changed.** Bug fixes and tweaks during a ramp are common, and they mean the two stages measured different interventions. Check the deployment history against the stage boundaries before believing any statistical story. - **Genuine scale effects.** Some effects really are smaller at scale — a recommendation change that works when few users compete for the same inventory, an infrastructure win that disappears when the cache is shared. These are rarer, but they exist and are worth naming. ## What to do about it The durable fixes are procedural. Decide before the ramp which stage produces the number that goes in the launch document, and report that one with its interval rather than the best number seen during the ramp. Tell stakeholders in advance that early reads are gates, not estimates, and that they should expect them to shrink — an expectation set before the shrink lands is a calibration exercise, whereas the same message afterwards sounds like an excuse. When decisions must be made on small-stage data, prefer explicitly shrunk estimates: a hierarchical or Bayesian model with a prior centred near zero pulls extreme reads back toward the bulk of historical effects, and historical effect sizes at most companies are small, which is exactly why the prior helps. Finally, keep the honest bookkeeping: if you launch on the basis of early reads across many features, the sum of the claimed lifts will exceed the measured movement of the company metric, and that gap is the winner's curse showing up in the aggregate.
- How would you check whether the drop is a population shift rather than the winner's curse?Re-estimate the later stage restricted to the same segment the early stage covered — same country, platform, or eligibility rule — and compare like with like. If the effect within that segment is still large, the drop came from reaching a different population, and the right report is a segmented one. If it shrinks there too, selection on a noisy early read is the better explanation.
- Which number belongs in the launch write-up?The estimate from the largest, latest stage whose population matches the one you are launching to, reported with its confidence interval. Never the maximum seen across stages, and never an average of the stage readings — both are selected quantities. Say explicitly which stage the number came from, so a reader can judge its precision instead of inheriting your selection.
- Does the winner's curse mean the early result was wrong or fabricated?No. The early estimate was a legitimate draw from a very wide sampling distribution; nothing was miscomputed. What was wrong was treating a single high draw as the effect size. That is why the fix is procedural — pre-committing to which stage reports the number — rather than an investigation into the early data.
- How do you set expectations before a ramp so the shrink is not read as failure?State in the plan that early stages are safety gates and that their point estimates are expected to be exaggerated, with a rough sense of how wide they are. Name the stage that will produce the reported effect. Teams that hear this in advance treat the shrink as calibration; teams that hear it afterwards hear an excuse, and the analyst loses credibility they will need later.
It is the tallest person in a room of five: pick the maximum of a few noisy draws and you have picked partly for luck, so the next, larger measurement almost always looks smaller.
saying these in an interview costs you the question
- Claims the early result was a bug or fabricated
- Averages the 1% and 50% estimates together
- Concludes the feature genuinely got worse over time
- Treats an estimate as unbiased after selecting on it
- Reports the largest lift seen during the ramp