skip to content

Why does a post-exposure covariate bias a CUPED-adjusted treatment effect estimate?

level: middleimportance: should knowfreq 45%

answer

  1. the correction term must average to zero
  2. randomisation only balances what already happened
  3. treatment can move the covariate too
  4. subtracting a moved covariate deletes real effect
  5. test the covariate itself across arms

basics

~20 s

CUPED subtracts theta times the arms' difference in the covariate, which averages to zero only for a covariate fixed before exposure. A covariate the treatment can move carries real effect, so subtracting it deletes part of that effect.

solid answer

~50 s

The CUPED estimator is `(Ybar_t - Ybar_c) - theta * (Xbar_t - Xbar_c)`. Its unbiasedness rests on one fact: under randomisation, users in the two arms are drawn from the same population, so for a covariate whose value was already determined before exposure, `E[Xbar_t - Xbar_c] = 0` and the correction term has expectation zero. If instead `X` is measured after users see the treatment - sessions during the test window, in-experiment engagement, a downstream funnel step - the treatment can shift it, so `E[Xbar_t - Xbar_c] = delta_X` is nonzero and the estimator is biased by `-theta * delta_X`. Concretely, if the feature raises both visits and revenue, adjusting revenue by in-test visits subtracts a slice of the very effect you are trying to measure, usually shrinking it toward zero. The discipline is that the covariate window must close strictly before exposure begins.

go deeper

for a junior

Remember the one hard rule: the CUPED covariate must be measured before users are exposed. If you can say why data from during the test is off limits, you have the core of the answer.

for a middle

Write the estimator out and show that its correction term has zero expectation only when the covariate is fixed before assignment. Then state the bias explicitly as theta times the treatment's effect on the covariate.

for a senior

Talk about how contamination actually sneaks in - windows anchored to analysis time rather than exposure time, rolling segment tables, late-triggering users - and about the arm-balance check on the covariate that you would automate before trusting any adjusted result.

for a principal

The organisational angle matters here: a biased adjusted estimate looks perfectly healthy in a dashboard. Argue for platform-computed covariates with enforced windows and automated validity checks rather than per-team covariate choice.

## Where the unbiasedness comes from Write the CUPED estimator for the treatment effect explicitly: ``` delta_adj = (Ybar_t - Ybar_c) - theta * (Xbar_t - Xbar_c) ``` The raw difference `Ybar_t - Ybar_c` is unbiased for the average treatment effect because randomisation makes both arms representative samples of the same population. The adjusted estimator inherits that property **if and only if** the correction term has expectation zero. For a covariate whose value was already fixed before any user was exposed, that is guaranteed. Assignment is independent of anything already determined, so treatment and control users have the same distribution of `X`, hence `E[Xbar_t] = E[Xbar_c] = mu_X` and `E[Xbar_t - Xbar_c] = 0`. Therefore `E[delta_adj] = E[Ybar_t - Ybar_c]` - the true effect. The adjustment is a mean-zero correction whose only job is to cancel the chance imbalance that randomisation leaves behind in any single run. ## What breaks when the covariate is post-exposure Suppose `X` is measured during the experiment - a user's session count in the test window, their clicks on the new surface, or whether they reached a later funnel step. Treatment can change it. Let the true treatment effect on the covariate be `delta_X`. Then ``` E[delta_adj] = ATE - theta * delta_X ``` The estimator is biased, and the bias is neither small nor conveniently signed. Its magnitude scales with how strongly the covariate predicts the outcome, because that is what sets `theta`. The tighter the covariate is coupled to the metric - which is exactly what makes it attractive as a variance reducer - the more of the real effect it silently deletes. In the common case where the feature moves the covariate in the same direction it moves the metric, the adjusted effect is pulled toward zero, and the team concludes the feature did nothing. This is the failure mode with the worst ergonomics in the whole technique: it produces confident, narrow confidence intervals around a wrong number. Nothing in the output looks broken. ## How covariates go wrong in practice - **Windowing slips.** The pre-period is defined as the two weeks before the *analysis* rather than before *exposure*, so for late-entering users part of the window sits inside the experiment. - **Late-triggering users.** Users are exposed when they hit a particular surface, and the covariate is computed from an interval that includes the moments after their trigger. - **Backfilled tables.** The covariate is read from a table that is recomputed on a schedule and quietly reflects post-launch behaviour. - **Derived attributes.** A user segment such as an activity tier that looks like a static attribute but is recomputed daily from recent behaviour. The pattern is the same in each case: the pipeline treats the covariate as a fixed user property when it is really a rolling measurement. ## Diagnosing it The direct check is to treat the covariate as if it were a metric and test it across arms. Compute the arm means of `X` and their difference with the same machinery you would use for any metric. Under a correct setup, that difference should be noise: with a properly pre-treatment covariate and a healthy assignment, arms should differ on `X` no more often than chance allows. A persistent, significant difference means one of two things - the covariate is being contaminated by the treatment, or the randomisation itself is compromised (a sample-ratio or assignment-leakage problem). Both are worth stopping the analysis over. A complementary check is to run the whole pipeline on historical A/A data, where the true effect is zero. Adjusted estimates should centre on zero. If they do not, something in the covariate definition is picking up the arm. ## Chance imbalance is not a bug One nuance frequently confuses candidates. If the covariate is genuinely pre-treatment and the arms nevertheless differ on it in a particular run, CUPED is still unbiased - that imbalance is exactly what the correction is for. Unbiasedness is a statement about the average over repeated randomisations, not a promise that any single run is balanced. Removing the observed imbalance is precisely how the estimator gets its precision. ## Estimating theta A second-order point: `theta` is not known, it is estimated from the same data, which introduces a small finite-sample bias that is negligible at online-experiment sample sizes and vanishes asymptotically. The standard discipline is a single pooled `theta` computed across both arms. Fitting a separate `theta` in each arm makes the adjustment depend on arm-specific noise and on the effect itself, which is a needless way to reintroduce the very problem the pooled estimate avoids.

  • How would you check that a candidate covariate really is pre-treatment?
    Treat it as a metric and compare its arm means with the same test you would use on any outcome. Under a correct setup the difference should look like noise. A persistent, significant gap means either the covariate is contaminated by the treatment or the assignment itself is broken - and both stop the analysis.
  • If the arms happen to differ on a valid pre-period covariate, is CUPED still unbiased?
    Yes. Chance imbalance in a pre-treatment covariate is exactly what the adjustment removes. Unbiasedness holds over repeated randomisations, not within one run, and correcting the observed imbalance is where the precision gain comes from.
  • Does estimating theta from the experiment data introduce a problem?
    Only a small finite-sample bias that is negligible at typical online sample sizes. The discipline is one pooled theta across both arms. Fitting a separate theta per arm lets the adjustment depend on arm-specific noise and on the effect itself, which is an avoidable way to reintroduce bias.

saying these in an interview costs you the question

  • Uses in-experiment engagement as the CUPED covariate
  • Says CUPED removes confounding rather than chance imbalance
  • Assumes any covariate correlated with the metric is safe
  • Fits a separate theta in each arm without thinking
  • Treats an arm difference on the covariate as harmless

context