When would you use pre-experiment stratification on platform or country instead of a CUPED covariate?
answer
- both remove variance something known explains
- one acts at assignment, one at analysis
- only the between-group share is removable
- thin cells get noisy fast
- a user predicts themselves best
basics
~20 sStratify when no continuous pre-period metric exists and the metric differs sharply between a few large discrete groups. Its ceiling is the share of variance sitting between strata, which is usually well below what a user's own pre-period value explains.
solid answer
~50 sBoth remove variance that something known in advance explains; they differ in what that is and when it acts. Stratification balances assignment inside discrete cells - platform, country - at randomisation time, so the arms match on those cells by construction rather than being corrected afterwards. Its ceiling is the fraction of variance lying *between* strata: if revenue varies far more within a country than across countries, the gain is small however many cells you add. CUPED uses a continuous pre-period covariate and typically wins outright when one exists, because a user's own past behaviour predicts their future behaviour far better than their country does. Prefer stratification when no usable pre-period metric exists, when the metric is dominated by a few large groups, or when you want balance guaranteed rather than adjusted. They also compose: stratify at assignment and adjust at analysis.
go deeper
Know that both approaches use information available before the test to cut noise, and that stratification does it by balancing groups at assignment while a covariate adjustment does it in the analysis.
Be able to split total variance into a between-cell and a within-cell part and say that stratification can only remove the first. That decomposition is the whole answer to how much it can possibly buy.
Show you have made the call in practice: which attributes you stratified on, why you stopped adding cells, whether the analysis actually respected the design, and how much reduction you measured rather than hoped for.
Own the design standard. Decide which attributes the platform stratifies on by default, resist per-team slicing that produces thin cells, and be explicit that the two techniques overlap so nobody budgets for the sum of their gains.
## Two routes to the same goal Variance reduction in an experiment always works the same way: identify variation in the metric that something already known explains, and stop paying for it in the error bars. The two mechanisms differ in what that known thing is and at which stage it acts. **Stratification** partitions users into discrete cells before assignment - platform, country, maybe a coarse tenure bucket - and randomises within each cell. The arms then hold the same mix of cells by construction. **CUPED** leaves assignment alone and subtracts, at analysis time, a scaled pre-period covariate from each user's outcome. ## The ceiling on stratification Decompose total metric variance into a between-cell part and a within-cell part. Stratification removes only the between-cell part - the piece attributable to the fact that, say, one platform's users are systematically heavier spenders than another's. Whatever spread remains among users inside the same cell is untouched. For most per-user product metrics that decomposition is unkind. Revenue per user, sessions per user and time spent are heavy-tailed and dominated by individual propensity. Two users on the same platform in the same country can differ by an order of magnitude. So a stratification scheme whose cells explain, say, six percent of the variance removes six percent, and the remaining ninety-four percent is exactly where the noise lives. This also explains the shape of the tradeoff around cell count. Adding cells can only increase the variance explained by the partition, so it is tempting to slice finely. But thin cells create their own problems: within-cell estimates become noisy, some cells end up with too few users in one arm, and the bookkeeping grows. The practical rule is few, large, meaningfully-different cells. If the between-cell spread is small, no amount of slicing rescues it. ## When stratification is the right tool anyway - **No pre-period metric exists.** A brand-new surface, a newly defined metric, or a population that did not exist before launch leaves nothing to correlate against. Discrete attributes known at assignment are all you have. - **The metric really is group-dominated.** Some metrics - conversion in markets with very different baselines, latency across device classes - genuinely carry a large share of their variance between groups. - **Small experiments where balance matters.** With few users, chance imbalance on an important attribute is likely and awkward to explain. Guaranteeing balance at assignment removes the conversation entirely. - **Legibility.** A stratified design is easy to describe to stakeholders, and per-cell results fall out naturally. ## Requirements and traps The stratifying attribute must be **pre-treatment**, exactly as a CUPED covariate must be. An attribute recomputed from recent behaviour - an activity tier refreshed daily, a churn-risk segment - looks static and is not; if the treatment moves the underlying behaviour, cells stop being fixed and the design's guarantee evaporates. The analysis has to match the design. If you stratified at assignment, the estimate should be formed within cells and combined across them using each cell's share of the population, rather than pooling everything and hoping. Pooling is not catastrophic when cells are balanced, but it leaves part of the intended gain on the table. ## Composing them These are not exclusive. A common arrangement is to stratify assignment on the handful of attributes with genuinely different baselines and then apply a pre-period covariate adjustment at analysis. The gains overlap: to the extent the strata predict the covariate, they are removing the same variance twice, so the combined reduction is less than the sum of the two. Measure the joint effect on historical data rather than adding the numbers. ## The one-line comparison Where a pre-period version of the metric exists, it usually explains far more of the outcome than any coarse grouping, because a user is the best predictor of themselves. Stratification is the tool for the case where that self-prediction is unavailable, or where you want the balance guaranteed at randomisation rather than repaired at analysis.
- Can you use stratified assignment and a CUPED covariate in the same experiment?Yes, and it is a reasonable default: stratify on a few attributes with genuinely different baselines, then adjust with the pre-period covariate at analysis. The two gains overlap wherever the strata predict the covariate, so the combined reduction is smaller than the sum - measure it on historical data rather than adding the numbers.
- What goes wrong if you stratify on two hundred fine-grained cells?Cells become thin. Within-cell estimates get noisy, some cells end up with very few users in one arm, and the bookkeeping cost rises with no matching benefit. Keep cells few, large, and genuinely different in their metric baselines, or coarsen until they are.
- Which attributes are actually worth stratifying on?Ones where the metric baseline differs a lot between cells and each cell holds plenty of users, and where the attribute is fixed before exposure. If the between-cell spread is small, the reduction is small however you slice it, and an attribute recomputed from recent behaviour is not fixed at all.
saying these in an interview costs you the question
- Believes more strata always means more variance removed
- Stratifies on a segment recomputed from recent behaviour
- Assumes stratification and CUPED gains simply add up
- Expects country cells to beat a user's own pre-period metric
- Stratifies at assignment then ignores cells in the analysis