skip to content

What is CUPED, and how does a pre-experiment covariate reduce the variance of an A/B test?

level: middleimportance: must knowfreq 62%

answer

  1. reuse data from before the test started
  2. subtract a scaled pre-period deviation
  3. pick the multiplier that minimises variance
  4. covariance over variance, a slope
  5. the gain goes as the squared correlation

basics

~20 s

CUPED replaces outcome Y with Y - theta*(X - mean X), where X is a pre-experiment covariate. With theta = cov(Y, X) / var(X), variance falls to (1 - rho^2) of the original, rho being the correlation.

solid answer

~50 s

CUPED stands for Controlled-experiment Using Pre-Experiment Data. You pick a covariate `X` measured before the experiment started - most naturally the same metric over a pre-period, such as each user's revenue in the two weeks before launch - and analyse the adjusted outcome `Y_adj = Y - theta * (X - mean(X))` instead of the raw `Y`. Minimising `var(Y_adj) = var(Y) - 2*theta*cov(Y, X) + theta^2*var(X)` over theta gives `theta = cov(Y, X) / var(X)`, the least-squares slope of `Y` on `X`, and at that value the variance becomes `var(Y) * (1 - rho^2)` where `rho = corr(Y, X)`. Nothing about the users, the assignment or the true effect changes - you have only subtracted the part of each user's outcome that was already predictable from their own past, so the arm difference is measured against a quieter background.

code

python · 19 lines
python
import random, statistics

def arm_mean(vals, arms, a):
    return statistics.fmean([v for v, g in zip(vals, arms) if g == a])

def trial(n=2000, rho=0.7):            # true treatment effect is exactly 0
    arms = [i % 2 for i in range(n)]
    xs = [random.gauss(0, 1) for _ in range(n)]                  # pre-period covariate
    ys = [rho * x + random.gauss(0, (1 - rho * rho) ** 0.5) for x in xs]
    mx, my = statistics.fmean(xs), statistics.fmean(ys)
    cov = sum((x - mx) * (y - my) for x, y in zip(xs, ys)) / (n - 1)
    theta = cov / statistics.variance(xs)
    adj = [y - theta * x for x, y in zip(xs, ys)]
    return (arm_mean(ys, arms, 1) - arm_mean(ys, arms, 0),
            arm_mean(adj, arms, 1) - arm_mean(adj, arms, 0))

raw, cuped = zip(*[trial() for _ in range(500)])
print(round(statistics.fmean(raw), 4), round(statistics.fmean(cuped), 4))   # both ~ 0
print(round(statistics.variance(cuped) / statistics.variance(raw), 3))      # ~ 1 - 0.7**2

go deeper

for a junior

Be ready to state what CUPED does in one line: subtract a scaled version of each user's pre-experiment value from their outcome so the remaining spread is smaller. Knowing that the covariate must come from before the test already puts you ahead.

for a middle

You are expected to derive theta. Write the variance of the adjusted outcome, minimise it, land on covariance over variance, and say why the resulting variance is (1 - rho squared) times the original. Explaining that the gain squares the correlation is the point of the question.

for a senior

Show you have measured this rather than read it. Talk about which pre-period window you chose and why, what correlation you actually observed for the metrics you ran it on, and how you handled users whose adjusted number disagreed with the raw dashboard number.

for a principal

Own the decision of when the technique earns its keep. Weigh the pipeline and backfill cost against the realised correlations across your metric portfolio, and be clear that a technique with a metric-dependent payoff should never be sold internally as a fixed multiplier on power.

## The problem CUPED solves In a randomised online experiment the treatment effect is estimated as a difference in means: the average metric in treatment minus the average metric in control. The precision of that estimate is governed by the variance of the metric across users. For most product metrics - revenue per user, sessions, minutes watched - that variance is enormous and is dominated not by anything the experiment did but by who the users are. A handful of heavy users can move an arm's mean more than the feature under test does. The key observation is that much of that user-to-user spread is **predictable in advance**. A user who spent a lot last month tends to spend a lot this month. If you already know that, you do not have to pay for it in your error bars. ## The estimator CUPED (Controlled-experiment Using Pre-Experiment Data) takes a covariate `X` measured strictly **before** the experiment began and analyses ``` Y_adj = Y - theta * (X - mean(X)) ``` instead of `Y`. Here `mean(X)` is the pooled mean of the covariate across all users in the experiment, and `theta` is a single constant estimated from pooled data across both arms. Because the mean of `X` is subtracted, the adjustment term has mean zero across the experiment, so `mean(Y_adj) = mean(Y)`. For the treatment effect the estimator becomes ``` delta_adj = (Ybar_t - Ybar_c) - theta * (Xbar_t - Xbar_c) ``` Randomisation makes `E[Xbar_t - Xbar_c] = 0` for a pre-treatment covariate, so `delta_adj` targets the same true effect as the raw difference. What it removes is the *chance imbalance*: if by luck the treatment arm happened to contain slightly heavier spenders before the test, the second term subtracts exactly the part of the observed gap that imbalance explains. ## Choosing theta Expand the variance of the adjusted outcome: ``` var(Y_adj) = var(Y) - 2*theta*cov(Y, X) + theta^2 * var(X) ``` This is a parabola in `theta`, minimised where its derivative vanishes: ``` -2*cov(Y, X) + 2*theta*var(X) = 0 => theta = cov(Y, X) / var(X) ``` That is exactly the least-squares slope of `Y` regressed on `X`. Substituting it back: ``` var(Y_adj) = var(Y) * (1 - rho^2), rho = corr(Y, X) ``` So the whole benefit is summarised by one number: the squared correlation between the covariate and the outcome. This is also why CUPED is algebraically the same thing as fitting the outcome on a treatment indicator plus the pre-period covariate - the covariate soaks up the variance it can explain and the treatment coefficient is left with less noise around it. ## How large is the gain, really Because the reduction is `rho^2`, not `rho`, weak correlations buy almost nothing: - `rho = 0.2` removes 4 percent of the variance - not worth a pipeline. - `rho = 0.5` removes 25 percent. - `rho = 0.7` removes about 51 percent. - `rho = 0.9` removes 81 percent, which is rare outside very sticky per-user metrics. Two practical consequences follow. First, the covariate almost always wants to be *the same metric over a pre-period*, because a user's own past value of a metric is usually its single best predictor - a different metric, or a coarse attribute, correlates much less. Second, the length of the pre-period matters: too short and `X` is itself mostly noise, which depresses `rho`; very long and old behaviour has drifted away from current behaviour, and more users get excluded for lacking full history. A window of a couple of weeks is a common compromise. ## What CUPED does not do It does not change the true effect, does not change the estimand, and does not fix a badly chosen metric, a broken assignment, or a mismatched traffic split. It reduces noise, and only the noise a pre-period measurement can predict. It also requires the covariate to be genuinely pre-treatment: a quantity measured after users were exposed can itself have been moved by the treatment, and subtracting it removes part of the real effect. ## Reporting The adjusted point estimate is a different number from the raw difference in means on any single experiment - it is a different estimator with the same expectation, not a relabelling. Expect the confidence interval to be visibly narrower and the point estimate to shift slightly, occasionally enough to move a borderline result across the significance line. Teams that roll CUPED out usually show both numbers for a while so nobody thinks the metric definition changed.

  • How strong does the correlation have to be before CUPED is worth building a pipeline for?
    Because the variance reduction is `rho^2`, a correlation of 0.3 removes only about 9 percent of the variance, which rarely justifies the engineering. Around 0.5 you remove a quarter, and at 0.7 roughly half - that is where teams usually feel it. Measure the realised correlation on historical data for each metric before promising anything.
  • How does CUPED relate to fitting the outcome on the covariate?
    They are the same estimator written two ways. `theta = cov(Y, X) / var(X)` is exactly the least-squares slope of the outcome on the covariate, and subtracting `theta * (X - mean(X))` is what including the covariate alongside a treatment indicator does to the treatment coefficient. CUPED is the closed-form version for a single covariate.
  • Does CUPED change the point estimate you report for a given experiment?
    Yes. The adjusted difference in means is a different number on any particular dataset - it removes the part of the arm gap explained by pre-existing differences. What is unchanged is the expectation: both estimators target the same true effect. On borderline experiments the shift plus the narrower interval can change the decision, which is why the covariate must be fixed in advance.

It is like weighing sealed boxes when you already know each empty box's weight. Subtract the known tare and the spread that remains tells you far more about what is actually inside.

saying these in an interview costs you the question

  • Claims CUPED increases the true effect rather than cutting noise
  • States the variance reduction is rho, not rho squared
  • Picks theta by hand instead of cov(Y, X) / var(X)
  • Uses a covariate measured during the experiment window
  • Expects big gains from a covariate correlated 0.2 with the metric

context