skip to content

What does capping revenue per user at the 99th percentile do to the effect you are estimating in an A/B test?

level: middleimportance: must knowfreq 64%

answer

  1. not cleaning, a different estimand
  2. the average of min(X, c)
  3. bias traded for variance
  4. one percent of users, a third of revenue
  5. tail effects get flattened away

basics

~20 s

Capping replaces values above the 99th percentile with that value, so you estimate mean capped spend, not mean revenue. Variance falls sharply and power rises, but a real effect in the tail is shrunk or erased.

solid answer

~50 s

Capping is a deliberate change of estimand, not data cleaning. If `c` is the 99th percentile of the pooled pre-period spend distribution, the analysis compares arm means of `min(X, c)`, so the quantity being estimated is the treatment effect on `E[min(X, c)]` rather than on `E[X]`. That trade is usually worth it: the tail that carried most of the variance is flattened, the confidence interval narrows enormously, and the decision becomes callable. The costs are real too. The capped mean sits below true mean revenue, so the effect must be reported in capped units and never multiplied out into an annual revenue figure. And if the feature works precisely by making big spenders spend more, the cap deletes the very effect you are testing for. Only 1% of users are touched, but those users may carry something like a third of the revenue.

go deeper

for a junior

Know that capping replaces values above a threshold with the threshold, that it is applied to both arms alike, and that its purpose is to shrink the metric's variance so the test can be called.

for a middle

Be able to write the new estimand as the effect on the mean of min(X, c), and explain the trade: a downward-biased level and possible attenuation of tail effects in exchange for a much narrower interval.

for a senior

Demonstrate reporting discipline — the cap value, the fraction of users and the fraction of revenue it touches, results at neighbouring thresholds, and a refusal to annualise a capped number into a business forecast.

for a principal

Own the question of whether the capped metric is the right basis for the launch decision at all, especially for features whose value proposition is concentrated in high spenders, and set that policy before results exist.

## What the cap actually does Pick a threshold `c` — commonly the 99th percentile of the pooled spend distribution. Capping (winsorising at the top) replaces each user's value `X` with `min(X, c)`. A user who spent far above `c` is not removed; their value is pulled down to `c` and they remain in the analysis as a full randomised unit. The analysis then compares the two arms' means of `min(X, c)`. That is a different quantity from the mean of `X`, and this is the whole point of the question. ## Estimand: what you are now measuring - Without the cap, the estimand is the average treatment effect on revenue per user: `E[X | treated] - E[X | control]`. - With the cap, the estimand is `E[min(X, c) | treated] - E[min(X, c) | control]`. Both are legitimate causal quantities and both are estimated without bias by the corresponding difference in arm means, because the transform `min(., c)` is a fixed function applied identically to both arms before averaging. What changes is *which* question the number answers. The capped comparison answers: does the treatment change spend among the population, counting nobody as spending more than `c`? A test on the capped metric also retains its type I error control under the null of no effect at all: if the treatment changes nothing, capping changes nothing about the equality of the two arms' distributions, so a difference-in-means test on the capped values is still a valid test. It is the *interpretation* of a rejection, and the estimate you attach to it, that shifts. ## The bias-variance trade, stated honestly Variance: enormous gain. The users producing most of the squared deviations are exactly the ones being pulled to `c`, so the standard deviation of the capped metric can fall by a large multiple, and the confidence interval on the difference narrows with it. This is often the difference between a decidable experiment and an undecidable one. Bias, in two distinct senses: 1. **Level bias.** `E[min(X, c)] < E[X]` whenever any mass sits above `c`. The capped average understates revenue per user, so the capped number must never be presented as revenue and never annualised into a business figure. 2. **Effect bias.** If part of the true treatment effect lives above `c`, the cap flattens it. Attenuation toward zero is the usual direction, but the sign can genuinely flip: a treatment that raises spend far above the cap for a few users while slightly reducing spend in the body will read negative on the capped metric and positive on the raw one. ## The 1% that is not 1% The most common interview mistake is arithmetic intuition: 'only one user in a hundred is touched, so the effect on the metric must be tiny.' On a heavy-tailed revenue metric the capped 1% may carry something like a third of total revenue. Capping is a large intervention on the *mass* of the metric even though it is a small intervention on the *count* of users. Always report both numbers: fraction of users capped, and fraction of total revenue removed by the cap. ## Practical discipline - Fix the threshold before looking at the results, and derive it from data that cannot have been shaped by the treatment — the pooled pre-period distribution is the usual choice. - Apply the same numeric cap to both arms. A cap computed separately within each arm is a per-arm transform and drags the arms toward each other. - Report the capped effect in capped units, with the cap value and the affected fractions stated alongside. - Run the analysis at a couple of neighbouring thresholds as a sensitivity check. If the conclusion is stable across them, the cap is doing variance control; if the conclusion flips, the result is being decided by a handful of users and you do not have an answer yet. ## The sentence that scores 'Capping does not clean the data, it changes the estimand: I am now estimating the effect on mean capped spend, I trade a downward-biased level and possible attenuation of a tail effect for a large reduction in variance, and I say so in the write-up.'

  • Is a test on the capped metric still a valid test?
    Yes. Under the null of no effect at all, applying one fixed threshold identically to both arms leaves the two arms' distributions equal, so the difference-in-means test keeps its type I error rate. What changes is the meaning of a rejection: it says capped spend differs, not that revenue differs, and the point estimate is in capped units.
  • Which way does the cap bias the estimated treatment effect?
    Usually toward zero, because any part of the effect above the threshold is flattened. It is not guaranteed, though: if a treatment lifts a few users far above the cap while slightly depressing the body of the distribution, the capped estimate can be negative while the raw one is positive. That is a signal to look at where the difference is coming from, not to switch metrics after the fact.
  • Only 1% of users are capped — surely the impact on the metric is small?
    No. On a heavy-tailed revenue metric that 1% can carry something like a third of total spend, so a p99 cap removes a large share of the metric's mass. Always report the fraction of users capped and the fraction of revenue removed side by side; they are very different numbers.

saying these in an interview costs you the question

  • Describes capping as removing bad or invalid data
  • Claims the cap leaves the estimand unchanged
  • Reports the capped effect as a dollar revenue lift
  • Chooses the cap after seeing the treatment effect
  • Assumes capping 1% of users touches 1% of the metric

context