How do you choose the cap for a winsorised revenue metric so the A/B comparison stays fair?
answer
- pre-register the threshold
- derive it from the pre-period
- one number, both arms
- never a per-arm quantile
- sensitivity ladder across nearby caps
basics
~20 sFix one numeric threshold before the experiment reads out, derive it from pre-period data the treatment cannot have touched, and apply that same number to both arms. Report results at neighbouring thresholds as a sensitivity check.
solid answer
~50 sThree rules make a cap defensible. First, it must be **pre-specified**: written into the metric definition before results exist, otherwise the threshold is another analyst degree of freedom and the false-positive rate is no longer what you think. Second, it must be **treatment-independent**: derive it from the pooled pre-period distribution, never from a quantile computed inside each arm separately, since a per-arm quantile is a different transform per arm and drags the two means together. Third, it must be **one number applied identically to both arms**, at the level you average over, which for a user-randomised test is the per-user total. Then document what the cap costs: the fraction of users clipped and the fraction of revenue removed, which on a heavy-tailed metric differ wildly. Finally rerun at nearby thresholds; if the conclusion flips, you do not have a result.
go deeper
Remember the two hard rules: the threshold is fixed before the results are seen, and the same number is applied to both arms.
Explain why a per-arm quantile breaks the comparison, and why a threshold taken from the pre-period is safer than one computed on the experiment's own outcomes.
Demonstrate the operating discipline: choose the loosest cap that makes the metric decidable, apply it at the per-user total over a stated window, and read a sensitivity ladder across neighbouring thresholds before calling a result.
Own the cap as shared platform policy rather than per-team choice — where the number lives, who may change it, how a change is logged so historical results stay interpretable, and what evidence lets a team deviate.
## The threshold is a decision, and it must be made in advance A cap is a parameter of the metric, not of the analysis. Choosing it after seeing the treatment effect is a garden of forking paths: you could try p95, p99, p99.5 and 'a round number', and one of them will flatter the result. That inflates the false-positive rate beyond the nominal level in a way no p-value on the final choice will reveal. So the threshold goes into the metric definition, ideally shared across all experiments on that metric, before any experiment reads out. ## Where the number should come from **Pooled pre-period data is the cleanest source.** Take the spend distribution over the eligible population in a window before the experiment started and read off, say, the 99th percentile. Because that window precedes assignment, the treatment cannot have influenced it, so the transform is guaranteed independent of the outcomes being compared. **Pooled in-experiment data is acceptable but weaker.** A quantile computed on both arms combined is only mildly dependent on the treatment — the dependence vanishes under the null and is usually negligible with large samples — but it makes the transform a function of the data you are analysing, which is harder to defend and harder to reproduce. **Per-arm quantiles are wrong.** If you cap the treatment arm at its own p99 and the control arm at its own p99, you have applied two different transforms. Suppose the treatment genuinely lifts the tail: its p99 is higher, so its cap is looser, and the comparison is no longer like-for-like in either direction you can reason about cleanly. The general rule is that any transform applied to the outcome must be identical across arms and, ideally, fixed before the outcomes exist. ## Which quantile, and at which unit There is no theorem that hands you p99. What you are trading is bias against variance, and the practical way to pick is to look, on historical data, at where the variance reduction saturates. Moving from no cap to p99 usually removes most of the variance; going further to p95 buys much less variance while removing far more of the metric's mass and pushing the estimand further from revenue. Pick the loosest cap that makes the metric decidable at your traffic. Apply the cap at the unit you average over. For a user-randomised experiment analysed on revenue per user, that is each user's total across the experiment window, not each individual transaction — capping transactions and then summing lets a user with many capped purchases still dominate the mean, which defeats the purpose. Also decide the window explicitly: a per-user total over a two-week experiment and a per-user total over six weeks have different distributions, so the same numeric cap means something different. Recompute the threshold when the metric window changes, and refresh it periodically as the spend distribution drifts — with a change log, since a moved cap makes the metric non-comparable to its own history. ## What to report Always publish, next to the result: - the cap value and its provenance (which window, which quantile); - the percentage of users whose value was clipped; - the percentage of total revenue removed by clipping — on a heavy-tailed metric this can be a third even when only 1% of users are touched; - the same clipped fractions per arm, as a sanity check that the two arms are being affected comparably; - the result at neighbouring thresholds. ## Reading the sensitivity ladder Run the analysis at, for example, p99.5, p99 and p98. Three outcomes: 1. **Same sign, similar magnitude, intervals overlapping.** The conclusion is robust; the cap is doing variance control and nothing else. 2. **Magnitude grows steadily as the cap loosens.** The effect has a real tail component; the capped number is a conservative floor, and it is worth saying so explicitly rather than pretending the capped effect is the whole story. 3. **The sign flips.** A handful of users are deciding the answer. Report that you cannot call it, and look at the individual large movers rather than picking the threshold that agrees with you. ## The sentence that scores 'One threshold, pre-registered, derived from data the treatment could not have touched, applied identically to both arms, reported with the share of users and the share of revenue it removes — plus a sensitivity ladder so nobody has to trust the exact quantile.'
- What goes wrong if each arm is capped at its own 99th percentile?You have applied two different transforms and lost the like-for-like comparison. If the treatment lifts the tail, its own p99 is higher, so its cap is looser, and the difference you measure mixes the effect with the difference in thresholds. Any outcome transform must be one fixed function applied identically to both arms.
- The conclusion flips between a p98 and a p99.5 cap — what do you do?Report that the experiment does not support a call. A flip means a small number of users near or above the threshold are deciding the answer. Inspect those users directly, check whether the movement survives dropping the top few, and consider a longer run or a design aimed at the tail — but do not select the threshold that produces the answer you like.
- How often should the threshold be recomputed?On a schedule, not per experiment: the spend distribution drifts with pricing, seasonality and mix, so an old cap slowly stops being the 99th percentile of anything. Recompute periodically, log the change, and note that results before and after a cap change are not directly comparable in magnitude.
saying these in an interview costs you the question
- Picks the threshold after seeing the treatment effect
- Computes the quantile separately within each arm
- Caps individual transactions instead of the per-user total
- Reports the capped result with no sensitivity at other caps
- Never states what fraction of revenue the cap removes