Should data from earlier ramp stages be pooled into the final effect estimate of an experiment?
answer
- pooling is about comparability, not size
- same treatment, same population, same metric
- arm share changes stage to stage
- stratify by stage, then weight
- disagreement between stages is the finding
basics
~10 sOnly when every stage estimated the same thing: unchanged treatment, unchanged eligible population, unchanged metric. Even then, pool by combining per-stage effect estimates, never by summing raw counts across stages with different traffic splits.
solid answer
~50 sPooling is a question about comparability, not about sample size. If the treatment was patched mid-ramp, or early stages were limited to one country or platform, or the metric definition changed, the stages estimate different quantities and adding them produces a number that answers no question. When the stages are genuinely comparable, the safe method is a stratified estimate: compute the effect within each stage and combine those estimates with weights, treating stage as a stratum. Naive count-pooling is the trap, because the treatment share and the baseline level both change from stage to stage, so period acts as a confounder — the pooled difference can be much smaller than every stage's difference, or even carry the opposite sign. If the stages disagree substantially, that heterogeneity is itself the finding: report the stage whose population matches the launch population rather than smoothing it away.
go deeper
Know that data from earlier ramp stages is not automatically part of the final number, and that a bigger pile of data does not by itself make an estimate better.
Be ready to state the comparability conditions — unchanged treatment, population and metric — and to explain that pooling should combine per-stage estimates rather than raw counts.
Demonstrate the diagnosis: check deploy history against stage boundaries, compare per-stage estimates for heterogeneity, and choose between precision weighting and population weighting with a stated reason.
Own the reporting standard — which stages count toward a headline effect, how combination is weighted, and the requirement to publish per-stage numbers so pooled claims stay auditable across teams.
## Why anyone wants to pool A ramp accumulates data at every stage, and it is tempting to treat all of it as one experiment: more users, tighter interval, faster decision. Sometimes that is right. Often it is not, and the failure is not obvious from the output — a pooled number always looks more authoritative than a stage number. ## The three comparability conditions Before pooling anything, check that the stages were measuring the same quantity. **1. Same treatment.** Ramps are exactly when fixes land: a rendering bug at 1%, a copy tweak at 5%, a performance patch at 25%. If the intervention changed, the earlier data measures a different intervention. Compare the deploy history against the stage boundaries before you trust any pooled figure. **2. Same eligible population.** Early stages are frequently restricted — internal users first, then one locale, then one platform, then everyone. Even without an explicit restriction, the earliest exposed users are the ones who visit most often, so they skew active. If the composition changed, each stage estimates the effect on its own population, and their sum estimates the effect on a population that does not exist. **3. Same metric definition and pipeline.** Instrumentation added or corrected mid-ramp changes what the number means. A metric that only started logging correctly at the 5% stage cannot contribute a clean 1% contribution. ## Why summing raw counts is the specific trap Suppose you simply add up conversions and users per arm across all stages and take the difference of the two overall rates. Two things vary by stage: the **share of traffic in treatment** (1%, then 5%, then 25%) and the **baseline conversion level** (which drifts with time, season, and who is being exposed). When the arm share and the outcome level both vary across periods, the period becomes a confounder of the pooled comparison: the treatment arm's overall rate is weighted mostly toward late stages, while the control arm's overall rate is weighted mostly toward early stages. The pooled difference then mixes the treatment effect with the difference between periods. The effect can be understated, overstated, or reversed relative to the effect present in every single stage. Nothing in the arithmetic warns you. ## The correct way to combine Use stage as a stratum. Compute the effect estimate and its standard error within each stage, where the arm shares are fixed and the period is common to both arms, then combine the per-stage estimates with weights. Two weighting choices, with different meanings: - **Inverse-variance (fixed-effect) weights** — each stage contributes in proportion to its precision. This gives the most precise combined estimate *under the assumption that all stages share one true effect*. - **Population weights** — weight stages by how much they resemble the population you are about to launch to. Less precise, more relevant to the decision. Either way, look at the per-stage estimates before combining. If they disagree by more than their intervals comfortably allow, the single-effect assumption is failing, and the honest report is the disagreement, not the average that hides it. ## Practical stance For most ramps, the pragmatic answer is: use the largest, most recent stage as the primary estimate, and treat the earlier stages as supporting evidence about direction and safety rather than as contributions to the headline number. You lose some precision and gain a number that means exactly one thing — the effect on the population you are launching to, under the version of the treatment you are shipping. When traffic is scarce enough that discarding earlier stages is genuinely costly, do the stratified combination properly and state the weighting you used. ## What to say in the write-up Name which stages entered the estimate and which were excluded and why; state the weighting; show the per-stage estimates alongside the combined one. A reader who can see the per-stage numbers can judge for themselves whether the pooling was reasonable, which is the difference between an analysis and an assertion.
- If you do combine stages, how should the per-stage estimates be weighted?Inverse-variance weights give the most precise combination when you believe all stages share one true effect — each stage contributes in proportion to its precision. Population weights, matching stages to the population you are launching to, give a less precise but more decision-relevant number. Choose deliberately, say which you used, and show the per-stage estimates so a reader can see whether the single-effect assumption was plausible.
- Why can summing raw counts across ramp stages distort the pooled difference?Because the traffic split and the baseline level both change from stage to stage. The treatment total is dominated by late stages and the control total by early ones, so the pooled difference blends the treatment effect with period differences. The result can be smaller than every stage's effect, or of the opposite sign, with nothing in the arithmetic signalling a problem.
- The per-stage estimates disagree well beyond their intervals. What do you report?Report the disagreement. Substantial heterogeneity means the stages are not measuring one quantity — the treatment, the population or the pipeline changed — so a combined number would hide the actual finding. Investigate the cause, then present the stage whose population and treatment version match what you are shipping, with the others shown as context.
It is like averaging two surveys that used different questionnaires on different cities: the combined number is precise and answers nothing anyone asked.
saying these in an interview costs you the question
- Adds up conversions across stages and takes a ratio
- Treats ramp data as free extra statistical power
- Pools despite the treatment being patched mid-ramp
- Ignores that early stages covered a narrower population
- Averages stage estimates without weighting or checking agreement