skip to content

Delta Method for Ratios

Metrics like clicks per session have a random denominator, so the naive standard error is wrong; the delta method or a bootstrap over randomization units gives the honest variance.

on this pageshow

questions

4

Why does a per-impression standard error for click-through rate understate uncertainty when users are randomized?

level: middleimportance: must knowfreq 72%

answer

  1. which unit was actually randomized?
  2. events inside a user are correlated
  3. a few power users carry the traffic
  4. effective n is users, not impressions
  5. variance too small, point estimate fine

basics

~20 s

A per-impression standard error treats every impression as an independent draw, but impressions from one user are correlated and a few heavy users supply most of them. The effective sample size is users, not impressions.

solid answer

~40 s

Click-through rate is usually reported as total clicks over total impressions, so the analysis unit is the impression. But randomization happened at the user, and a user's impressions share that user's own click propensity, so they are positively correlated rather than independent. The binomial formula `sqrt(p*(1-p)/impressions)` charges you the full information of every impression, so it comes out far too small — often several times too small when a minority of power users contribute most of the traffic. The consequence is intervals that are too narrow and a false-positive rate well above the nominal 5%. The fix is to compute variance at the user level: either the delta method on per-user numerator and denominator totals, or a bootstrap that resamples users and carries all of their events along.

code

python · 19 lines
python
import random, statistics
random.seed(0)
users = []
for _ in range(2000):                      # a few users generate most impressions
    n = random.choice([1, 1, 2, 5, 200])   # impressions for this user
    p = random.uniform(0.02, 0.50)         # this user's own click propensity
    users.append((sum(random.random() < p for _ in range(n)), n))

clicks = sum(c for c, _ in users)
imps = sum(n for _, n in users)
ctr = clicks / imps
naive = (ctr * (1 - ctr) / imps) ** 0.5    # pretends every impression is independent

boot = []
for _ in range(500):                       # resample USERS, carry their events along
    s = [random.choice(users) for _ in range(len(users))]
    boot.append(sum(c for c, _ in s) / sum(n for _, n in s))

print(round(ctr, 4), round(naive, 5), round(statistics.stdev(boot), 5))

go deeper

for a junior

Be ready to say which unit was randomized and which unit the metric is summed over, and to notice when those differ. Knowing that impressions from one user are not independent draws is the whole first step.

for a middle

Explain the mechanics: the binomial formula assumes independent draws, within-user correlation and skewed event counts break that, and the effective sample size is the user count. Be able to say the standard error scales with the square root of the variance inflation.

for a senior

Show you would catch this in production — A/A checks whose significance rate exceeds 5%, results that fail to replicate, user-level standard errors several times the per-event ones. Then name the concrete fix and confirm the point estimate is unchanged.

for a principal

Own the framing that this is a platform default, not a per-analysis choice. Argue for computing ratio-metric variance at the randomization unit everywhere, and for the cost of the alternative: a steady stream of confidently wrong launch decisions.

## The setup A ratio metric is one whose value is a quotient of two sums, for example click-through rate = total clicks / total impressions, or average order value = total revenue / total orders. The **analysis unit** — the row the quotient is built from — is the impression or the order. The **randomization unit** in a typical online experiment is the user: the coin was flipped once per user, and every event that user generates inherits the same arm. When those two units differ, the naive variance formula silently assumes the wrong thing. ## What the naive standard error computes Treating CTR as a proportion over N impressions gives ``` SE_naive = sqrt( p * (1 - p) / N_impressions ) ``` That formula is exactly right for N independent Bernoulli draws with common success probability p. Both parts of that sentence fail here. 1. **Independence fails.** Impressions from one user are not independent draws. A user who habitually clicks has a high personal click rate; a user who habitually ignores ads has a low one. Conditional on the user, the impressions are still correlated with each other through that shared propensity, which is exactly positive intra-user correlation. 2. **Equal weight fails.** Event counts per user are heavily skewed. If 5% of users generate half the impressions, the ratio of totals is effectively an average dominated by a handful of people. Their idiosyncrasies move the metric, and there are far fewer of them than the impression count suggests. ## How wrong it gets The usual way to express the damage is the **design effect**: the ratio of the true variance under clustered sampling to the variance the independent-sampling formula assumes. A rough guide for equal-sized clusters is ``` design effect ~= 1 + (m - 1) * rho ``` where `m` is the average number of events per user and `rho` is the within-user correlation of the event-level outcome. With 20 impressions per user and a modest rho = 0.1, the variance is inflated about threefold and the honest standard error is about `sqrt(3) ~= 1.7` times the naive one. With heavier tails in the event counts — a realistic pattern where a small set of users supplies most impressions — the gap regularly reaches three to five times. Remember the square-root relationship: a ninefold variance inflation is a threefold standard-error error, not ninefold. Because the standard error sits in the denominator of the test statistic, a standard error that is three times too small turns a z of 0.7 into a z of 2.1. The nominal 5% false-positive rate becomes something in the tens of percent. Teams that analyse ratio metrics per event ship a stream of "significant" results that do not replicate, and the pattern is easy to catch after the fact: A/A tests on the same metric come back significant far more often than 5% of the time. ## What to do instead Collapse to the randomization unit first, then compute variance across those units. - **Per-user totals.** For each user i record `x_i` (clicks) and `y_i` (impressions). The metric is `R = sum(x_i) / sum(y_i)`, unchanged — you are not redefining the estimand, only computing its uncertainty correctly. - **Delta method.** Apply a first-order Taylor expansion of the ratio of the two per-user means around their expectations. This yields a closed-form variance built from the per-user variances of x and y and their covariance, divided by the number of users. It is cheap enough to run over thousands of metrics in a scheduled job. - **Bootstrap over users.** Resample users with replacement, carry each sampled user's whole event history along, recompute the ratio of totals, and take the spread of the recomputed values as the standard error. Resampling *events* instead of users reproduces exactly the bug you are trying to fix. Both approaches converge to the same answer at reasonable sample sizes; the delta method is a formula, the bootstrap is a simulation. ## Sanity checks worth naming in an interview - The user-level standard error should be **larger** than the per-event one; if it comes back smaller, something in the aggregation is wrong. - The ratio of the two is an estimate of the square root of the design effect, and it should be stable across reruns of similar experiments. - The point estimate does not change. Clustering is a variance problem, not a bias problem — anyone who says the ratio itself is biased by this has misdiagnosed it. - Metrics whose analysis unit already *is* the user — such as fraction of users who clicked at least once — do not have this problem, which is one honest reason to prefer user-level metrics where they answer the question.

  • Does this clustering problem bias the CTR point estimate itself?
    No. The ratio of totals is the same number however you compute the variance around it. Correlated events within a user affect how much information the data carries, not where the estimate lands. What is biased is the standard error, and therefore the p-value, the confidence interval and the false-positive rate. If someone claims clustering shifts the estimate, they have confused variance with bias.
  • How would you demonstrate the problem to a sceptical stakeholder without any theory?
    Run A/A analyses: split the control population in half at the user level, compute the ratio metric and its per-event p-value, and repeat a few hundred times over historical logs. If the standard errors were right, about 5% of those splits would come back significant at the 5% level. Seeing 20-40% is a concrete, theory-free demonstration that the intervals are too narrow.
  • Which metrics on a dashboard are immune to this problem?
    Any metric whose analysis unit already equals the randomization unit: fraction of users who clicked at least once, revenue per user, sessions per user. Each user contributes exactly one number, so the usual independent-sample formulas apply directly. The problem is specific to quotients of event-level sums where one user supplies many events.
  • Does having millions of impressions rescue the naive standard error?
    No — it makes it worse in practice. The naive formula shrinks with the impression count, so more events from the same user base produce ever narrower intervals while the real information, which grows only with the number of users, stays flat. Large event counts are what makes the gap between the two standard errors most dramatic.

Polling one household of five and recording five opinions is not the same as polling five strangers. You have one household's worth of information, and reporting it as five inflates your confidence.

saying these in an interview costs you the question

  • Says the point estimate is biased rather than the standard error
  • Uses the binomial formula with the impression count as n
  • Claims a huge impression count makes the problem disappear
  • Bootstraps by resampling individual impressions instead of users
  • Assumes clustering makes intervals too wide rather than too narrow

context

open as a page

How does the delta method turn per-user means into a standard error for a ratio metric?

level: seniorimportance: should knowfreq 52%

basics

~20 s

The delta method takes a first-order Taylor expansion of the ratio of two per-user means around their expectations, then reads off the variance. The result combines the numerator variance, the denominator variance and, subtracted, their covariance.

open as a page

Average order value is measured per order but users were randomized: how do you get a valid standard error?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Aggregate each user into revenue and order-count totals, then compute uncertainty across users: either the delta method on those two per-user means, or a bootstrap that resamples users and carries all their orders. Never treat orders as independent rows.

open as a page

Your experiment platform reports per-event standard errors for every ratio metric: how do you prioritise fixing it?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Quantify the damage first by re-analysing past experiments with user-level variance and counting how many decisions flip. Then make correct variance the platform default for ratio metrics, and prepare teams for wider intervals and larger sample-size requirements.

open as a page