skip to content

How does the delta method turn per-user means into a standard error for a ratio metric?

level: seniorimportance: should knowfreq 52%

answer

  1. linearise a nonlinear function first
  2. Taylor expansion around the true means
  3. variance of a linear combination of means
  4. three terms, covariance carries a minus
  5. n is the user count, not events

basics

~20 s

The delta method takes a first-order Taylor expansion of the ratio of two per-user means around their expectations, then reads off the variance. The result combines the numerator variance, the denominator variance and, subtracted, their covariance.

solid answer

~50 s

Aggregate each user into a numerator total `x_i` and a denominator total `y_i`, so the metric is `R = mean(x) / mean(y)`. Linearise `g(a, b) = a/b` around the true means with a first-order Taylor expansion, which makes R approximately a linear combination of the two sample means, and take the variance of that linear combination. Writing per-user quantities, `Var(R) ~= (R^2 / n) * [ Var(x)/mean(x)^2 + Var(y)/mean(y)^2 - 2*Cov(x, y)/(mean(x)*mean(y)) ]`, where n is the number of users. The covariance term enters with a minus sign, and since heavy users usually have both a large numerator and a large denominator, that covariance is strongly positive and pulls the variance down — dropping it is a real error, not a conservative simplification. It needs a denominator mean safely away from zero and enough users for the linearisation to hold.

go deeper

for a junior

Recall that a ratio of two random quantities needs its own variance formula, and that the delta method is the standard one. Knowing that it works from per-user totals rather than raw events is enough at this level.

for a middle

Be able to walk through the mechanics: linearise the ratio with a first-order Taylor expansion, take the variance of the resulting linear combination, and identify the three terms. State that n is the number of randomization units.

for a senior

Get the covariance sign and its practical size right, and be ready to combine arms into a difference. Say when you would distrust the approximation — small user counts, near-zero denominators, extreme skew — and what you would use instead.

for a principal

Argue for a single variance convention across the metric catalogue: closed-form delta-method variance as the default because it scales to thousands of metrics, with resampling reserved for the metrics whose assumptions genuinely fail.

## What the delta method is for You can compute the variance of a sample mean directly. You cannot do the same for a ratio of two sample means, because the ratio is a nonlinear function of them and expectations do not pass through nonlinear functions. The delta method is the standard workaround: approximate the nonlinear function by its first-order Taylor expansion, which *is* linear, and take the variance of the linear approximation. ## Setting up the units Start by collapsing the data to the randomization unit. For every user i, form: - `x_i` — that user's numerator total (clicks, revenue, successful requests) - `y_i` — that user's denominator total (impressions, orders, requests) The metric is the ratio of totals, which is identical to the ratio of means: ``` R = sum(x_i) / sum(y_i) = mean(x) / mean(y) ``` The pairs `(x_i, y_i)` across users are independent, because the randomization was per user. That independence is what licenses everything below. ## The expansion Let `g(a, b) = a / b`, and let the true means be `mx = E[x]` and `my = E[y]`. The partial derivatives are `dg/da = 1/my` and `dg/db = -mx/my^2`. First-order Taylor expansion about `(mx, my)`: ``` R = g(mean(x), mean(y)) ~= mx/my + (1/my)*(mean(x) - mx) - (mx/my^2)*(mean(y) - my) ``` The right-hand side is a linear combination of the two sample means, and the variance of a linear combination is a formula everyone knows: ``` Var(R) ~= (1/my^2)*Var(mean(x)) + (mx^2/my^4)*Var(mean(y)) - 2*(mx/my^3)*Cov(mean(x), mean(y)) ``` Substituting `Var(mean(x)) = Var(x)/n` for n users, and factoring out `R^2 = mx^2/my^2`, gives the form usually implemented: ``` Var(R) ~= (R^2 / n) * [ Var(x)/mx^2 + Var(y)/my^2 - 2*Cov(x, y)/(mx*my) ] ``` In practice you plug in the sample means, sample variances and sample covariance of the per-user pairs. The standard error is the square root. Note where n lives: it is the number of **users**, which is the entire point — the event count never appears. ## Why the covariance term is not optional The covariance enters with a minus sign. Users with many impressions typically also have many clicks, so `Cov(x, y)` is large and positive, and the bracket shrinks substantially. Drop the term and you overstate the variance, sometimes badly. That direction is worth internalising, because it is the opposite of the clustering mistake. Ignoring clustering makes intervals too narrow; ignoring covariance makes them too wide. Neither is safe, and a candidate who can state both directions correctly stands out. There is a useful limiting intuition: if x were a fixed multiple of y for every user, the ratio would be a constant with zero variance, and it is exactly the covariance term cancelling the other two that produces that zero. ## Comparing two arms The delta method gives the variance of R within one arm. Because treatment and control contain disjoint, independently randomized users, the variances add: ``` Var(R_treatment - R_control) = Var(R_treatment) + Var(R_control) ``` Divide the observed difference by the square root of that sum to get a z statistic, or build the interval directly. For a relative lift `R_treatment / R_control - 1`, apply the delta method a second time to the ratio of the two independent ratios; that second application has no covariance term, since the arms are independent. ## Assumptions and where it breaks - **`my` bounded away from zero.** The expansion divides by `my`, and near zero the linearisation degrades and the variance blows up. A metric whose denominator is often zero or tiny is a bad candidate. - **Enough units.** The approximation is asymptotic. It relies on the pair of sample means being close to their expectations and on a central limit theorem applying to them, which needs a decent number of users, more when the per-user totals are skewed. - **First order only.** The expansion discards the curvature term, so both the variance and the small bias of the ratio are approximated. Higher-order terms shrink like 1/n and are usually ignorable at experiment scale. - **Independence across units.** Guaranteed by the per-user randomization; if the same person appears as two units, that assumption is gone. When these hold, the delta method and a user-level bootstrap agree closely. The delta method wins on cost — it is a closed-form expression over per-user sufficient statistics, so a platform can evaluate it for thousands of metrics in a single pass.

  • What happens to the estimate if you drop the covariance term?
    You leave the point estimate untouched and inflate the variance. Numerator and denominator per-user totals are usually strongly positively correlated, so the subtracted covariance term is large; dropping it produces intervals that are noticeably too wide and an underpowered test. It is a conservative error rather than an anti-conservative one, but it still costs real experiments their conclusions.
  • How do you get the variance of the difference between the treatment and control ratios?
    Compute the delta-method variance separately within each arm and add them. The arms contain disjoint sets of independently randomized users, so there is no cross-arm covariance term. The standard error of the difference is the square root of that sum, and the usual z or t interval follows.
  • When would you distrust the delta-method standard error and reach for something else?
    When the denominator mean sits near zero or many users have a zero denominator, when the number of users is small, or when per-user totals are so skewed that the sample means are far from normal. In those cases the first-order expansion is unreliable and a user-level resampling estimate is the safer standard error.

saying these in an interview costs you the question

  • Adds the covariance term instead of subtracting it
  • Uses the event count as n in the formula
  • Treats numerator and denominator as independent by default
  • Applies it when the denominator mean is near zero
  • Claims the delta method changes the point estimate

context