Why does revenue per user make an A/B test hard to call when a few users spend far more than the rest?
answer
- the problem is spread, not the average
- a few users carry most of the total
- standard error scales with sigma
- variance is built from squared deviations
- one whale can flip the sign
basics
~10 sAlmost all the variance in revenue per user comes from a handful of very large spenders, so the interval around the treatment-control difference stays wide. Ordinary experiment sizes then cannot resolve realistic revenue effects.
solid answer
~50 sRevenue per user is heavy-tailed: most users spend nothing, most payers spend a little, and it is not unusual for the top 0.1% of users to account for something like a third of all revenue. The standard error of a mean is `sigma / sqrt(n)`, and here `sigma` is several times the mean itself, so the interval on the difference between arms stays wide even with millions of users. The estimate is also unstable: which arm a few whales happen to land in can move the point estimate more than the feature does, so a rerun would report a very different number. That is why teams decide on a capped version of the metric, or on a bounded proxy such as whether the user purchased at all, and treat raw revenue per user as a directional readout.
go deeper
Be ready to say in one breath that a few very large spenders inflate the metric's variance, that the standard error of a mean is sigma over root n, and that this is why the confidence interval stays wide.
Explain the mechanics: variance sums squared deviations so the tail dominates, and the detectable effect scales with sigma over root n, meaning four times the traffic only halves the interval.
Show the diagnostic instinct — recompute the arm difference with the top few spenders dropped, check how much of the movement survives, and explain why day-to-day swings in the metric are expected rather than alarming.
Own the consequence for how the organisation decides: name in advance which bounded or capped metric a launch hangs on, and keep raw revenue per user as a reported number that never by itself authorises a ship.
## What 'heavy-tailed' means for an experiment metric A metric is heavy-tailed when a small fraction of the units contributes a large share of the total. Revenue per user is the standard example on consumer products: the majority of users spend nothing at all, the users who do spend mostly spend small amounts, and a tiny minority spend enormously. A common shape is that the top 0.1% of users account for roughly a third of total revenue. Watch-time per user behaves the same way — a handful of always-on sessions can sit an order of magnitude above the typical user. Nothing about this is a data-quality problem. Those users are real, their spend is real, and the business genuinely wants their money. The difficulty is statistical, not about validity. ## Why the tail hurts, mechanically An A/B test on a mean compares two sample averages. The uncertainty in one arm's average is `SE = sigma / sqrt(n)`, where `sigma` is the standard deviation of the per-user metric, and the uncertainty in the difference between arms combines both arms' standard errors. The size of effect you can resolve therefore scales with `sigma / sqrt(n)`: halving the width of the interval means quadrupling the traffic. For revenue per user, `sigma` is frequently several times the mean. A metric whose standard deviation is, say, eight times its mean needs a sample size two orders of magnitude larger than a well-behaved bounded metric to detect the same relative change. Traffic is finite and experiments have to end, so in practice the interval simply never closes enough for a realistic effect. The reason `sigma` is so large is that variance is built from *squared* deviations. A user who spends a hundred times the average contributes on the order of ten thousand times the squared deviation of a user who is one unit away from the average. A few such users dominate the sum, so the sample standard deviation is essentially a statement about the tail rather than about the typical user. ## The instability you can see with your own eyes Randomisation guarantees the two arms are comparable *in expectation*, not in any particular draw. With a heavy tail, a single very large spender landing in treatment rather than control can flip the sign of the observed difference. Practitioners see this as: the metric swings day to day, the effect looks large one week and reverses the next, and re-slicing the data changes the story. That is not a bug in the pipeline; it is the sampling distribution of a mean with a fat tail. A useful diagnostic is to recompute the arm difference with the single largest spender removed, then the top five, and see how much of the observed movement survives. If most of it does not, the number is not measuring the feature. ## What teams do about it Three responses are standard, and they are choices about the estimand, not just cleanup: - **Cap or winsorise** the per-user value at a high quantile, so the metric being averaged is bounded. This cuts variance dramatically at the price of estimating the mean of a capped quantity rather than mean revenue. - **Substitute a bounded proxy**, such as whether a user purchased at all, or watch-time capped at three hours. A bounded metric has small variance by construction and is far easier to power. - **Keep uncapped revenue as a reported, non-deciding number**, alongside its (wide) interval, so nobody mistakes a noisy swing for a result. ## What not to conclude Do not say the extreme users are outliers to be deleted — deleting them changes what the business is being asked about, and dropping units by their outcome breaks the clean randomised comparison. Do not say the test is invalid; a difference-in-means comparison of two randomised arms is still a legitimate comparison, it is simply underpowered. And do not promise that more traffic will fix it: it will, eventually, but the required multiple is often far beyond what the product can supply in a sane experiment window, and every extra day brings its own whales. The interview answer that lands is: the mean is fine, the *variance* is the problem, the variance lives in the tail, and the fix is to change the metric on purpose and say out loud what the new metric estimates.
- Can you just run the experiment longer until the interval closes?In principle yes, since the interval narrows with `1 / sqrt(n)` — but quadrupling traffic only halves it, and for a metric whose standard deviation is many times its mean the required sample is usually far past what the product can supply in a reasonable window. Longer runs also accumulate more extreme spenders, so the sample standard deviation does not conveniently shrink.
- Which users actually contribute the variance?The far upper tail. Variance sums squared deviations from the mean, so a user a hundred times above average contributes on the order of ten thousand times as much as a user one unit above it. That is why a metric can look stable in its median and be violently unstable in its mean.
- Does this only affect money metrics?No. Any per-user total with a long tail behaves the same way — watch-time per user is the classic engagement case, where a small number of always-on sessions dominate the average. The same treatment applies: cap the per-user value or use a bounded version of the metric for the decision.
Weighing a room of people to detect a one-gram diet effect works fine until someone drives a truck onto the scale. The truck is real, but it decides the reading.
saying these in an interview costs you the question
- Claims the extreme spenders are data-quality bugs to be deleted
- Says more traffic always makes the effect detectable in practice
- Focuses on the skewed mean and ignores that variance drives power
- Treats one arm getting a whale as evidence the feature worked
- Calls the randomised comparison invalid rather than underpowered