As experimentation lead, would you make CUPED the default analysis for every test on the platform?
answer
- the risk is selection, not the estimator
- lock the covariate per metric in advance
- validate with A/A before trusting it
- the gain varies by metric
- show both numbers during the transition
basics
~20 sDefault it on only where the platform computes a locked, pre-registered covariate per metric and A/A validation confirms nominal error rates. Blanket promises fail because the gain is metric-dependent, and analyst-chosen covariates turn a variance tool into a p-hacking surface.
solid answer
~50 sA qualified yes. Default-on is right for a curated set of metrics where the platform computes the pre-period covariate itself, the covariate is fixed per metric before any test starts, and the pre-period window is enforced to close before exposure. That uniformity is the main safety argument: if analysts choose covariates after seeing results, the technique becomes a way to shop for significance. Before switching it on I want large-scale A/A validation showing adjusted effects centred at zero, false-positive rates at nominal alpha, and realised variance ratios near the theoretical one. I refuse to sell it as a fixed multiplier on power - the gain is the squared correlation, strong for sticky per-user metrics and near zero for new-user ones. For a transition period both numbers get shown, because an unexplained gap between the adjusted delta and the dashboard destroys trust.
go deeper
You are unlikely to be asked to own this decision, but know the headline: applying the adjustment uniformly and deciding the covariate in advance is safer than letting each analysis choose its own.
Be ready to explain why a covariate chosen after seeing the result is a problem even though each individual adjusted estimate is unbiased - the harm comes from selecting among estimates, which inflates the false-positive rate.
Talk concretely about validation: A/A runs at scale, the arm-balance check on the covariate itself, realised variance ratios per metric, and replaying historical experiments to see whether decisions move.
Own the whole rollout as a governance problem. Decide which metrics get it, who can change a covariate and under what review, how the gain is communicated without overpromising, and how you keep it from becoming a lever teams pull after a disappointing readout.
## The question behind the question This is not really about statistics; the estimator is settled. It is about whether a technique with a metric-dependent payoff and one sharp failure mode should be applied uniformly by a platform, or left to teams. The answer that holds up in an interview is: uniform application is safer than discretionary application, but only if the platform takes on the obligations that discretion was implicitly covering. ## The case for default-on **Free precision where it applies.** For sticky per-user metrics with a good pre-period analogue, the reduction is real and large, and it costs nothing at decision time. **Removing a p-hacking surface.** This is the strongest argument, and it is counterintuitive. A per-team, opt-in adjustment invites the worst possible workflow: run the test, see a null result, try a covariate, report whichever number is significant. The covariate is pre-treatment either way so each individual estimate is unbiased, but *choosing among estimates after seeing them* inflates the false-positive rate. A fixed, platform-computed covariate per metric, locked before the test starts, means there is exactly one number and nothing to select from. **Consistency of the historical record.** If half the experiments are adjusted and half are not, comparisons across time and teams are apples to oranges. ## The obligations that come with it 1. **Pre-registration by construction.** The covariate for each metric is defined in the metric registry, versioned, and cannot be changed mid-experiment. Changing it is a metric change with the same review as any other. 2. **Enforced windows.** The pre-period must close strictly before each user's exposure, not before the analysis date. This is where contamination actually enters, and it should be a property of the pipeline rather than of an analyst's care. 3. **Automated validity checks.** Every experiment reports the arm difference on the covariate itself. A persistent significant gap means either the covariate is contaminated or the assignment is broken, and either should block the readout rather than appear in a footnote. 4. **A/A validation before rollout, and continuously after.** At scale, adjusted effects must centre on zero, the false-positive rate must sit at nominal alpha, and the realised variance ratio should be close to one minus the squared correlation. Replaying historical experiments through the new pipeline and checking that decisions do not flip arbitrarily is a useful complement. 5. **Per-metric honesty about the gain.** Publish the realised correlation and variance ratio per metric rather than a headline number. A metric whose reduction is three percent should say so. ## Where it should not be default Populations with no pre-period history get nothing from it, and quietly shipping a no-op is fine as long as the reporting says so rather than implying a gain. Brand-new metrics have no track record for the correlation, so they start unadjusted until there is historical evidence. Metrics that are already heavily transformed or capped may show much weaker correlation than intuition suggests, which is again a matter for measurement rather than argument. ## The organisational problem nobody mentions The adjusted point estimate is a different number from the naive difference of means that stakeholders see in dashboards. It shifts, sometimes across a decision boundary. If the first time a product manager encounters this is in a launch review, the credibility damage outlasts the precision gain. The rollout plan therefore includes showing both numbers side by side with their intervals for a transition period, a short written explanation of what the adjustment removes, and worked examples from the team's own past experiments. The failure mode to name explicitly is the one where CUPED becomes a knob teams turn after a disappointing readout. Any design that leaves that door open - opt-in adjustment, analyst-selected covariates, re-running with a different pre-period - undoes the statistical benefit and worse, teaches the organisation that the analysis is negotiable. ## The answer in one breath Default-on for a curated metric set with platform-computed, locked covariates and enforced pre-exposure windows; validated by A/A before and monitored after; reported with realised per-metric gains and, initially, alongside the unadjusted number. Never sold as a blanket multiplier on power, and never available as a post-hoc option.
- How would you validate a CUPED rollout before trusting it?Run it at scale on A/A data: adjusted effects should centre on zero, the false-positive rate should sit at nominal alpha, and the realised variance ratio should be close to one minus the squared correlation. Then replay past experiments through the pipeline and check that decisions do not flip in ways nobody can explain.
- What specifically stops CUPED from becoming a p-hacking tool?The covariate is defined per metric in the registry and locked before the test starts, and the platform computes it. Analysts cannot select it after seeing results, and only one adjusted number exists per metric. Individual estimates were never the problem; choosing among several after the fact was.
- How do you explain to a product manager that the adjusted number differs from the dashboard?Same quantity, less noise. The adjustment removes the part of the arm gap that pre-existing differences between the users already explain, so the estimate moves slightly and the interval narrows. Show both numbers with their intervals for a transition period so the narrower one is visibly the same story told more precisely.
saying these in an interview costs you the question
- Lets analysts pick the covariate after seeing the results
- Promises one fixed variance reduction across all metrics
- Ships the rollout without any A/A validation
- Hides the unadjusted estimate from stakeholders entirely
- Treats a re-run with a new pre-period as a harmless retry