Why winsorise rather than trim extreme spenders when analysing an A/B test on revenue per user?
answer
- transform the value, do not select the unit
- trimming conditions on a post-treatment outcome
- arm sizes stay equal under winsorising
- dropped counts can differ by arm
- trimmed-mean standard error needs care
basics
~10 sWinsorising clips extreme values to a threshold but keeps every randomised user, so the arms stay comparable. Trimming deletes users for their own outcome, conditioning on something the treatment could have caused.
solid answer
~50 sBoth procedures tame the tail, but they leave you estimating different things and only one respects the randomisation. Winsorising replaces `X` with `min(X, c)`: everyone assigned is still counted, arm sizes are unchanged, and the estimand is the effect on `E[min(X, c)]`. Trimming drops users whose spend exceeds `c`, so the analysis conditions on a post-treatment outcome — the treatment itself may have decided who crosses the threshold — and the retained subpopulations are no longer guaranteed comparable. Arm sizes drift apart too, and if the treatment pushes more users over the line you have quietly removed the effect. Trimming has a further trap: the standard error of a trimmed mean is not the naive `s / sqrt(n_kept)`, which understates uncertainty. Use winsorising for the primary analysis and trimming, if at all, as a sensitivity check.
go deeper
Know the difference in one line: winsorising clips a value to a threshold and keeps the user, trimming deletes the user. Know that experiments prefer the version that keeps everyone.
Explain that trimming conditions on a post-treatment outcome so the surviving groups are no longer comparable, and that winsorising keeps arm sizes equal and gives a stateable estimand over the full randomised population.
Show the operational checks: drop counts per arm, a trimmed sensitivity run alongside the winsorised primary, and a read on what a disagreement between them says about where the effect lives.
Decide the platform default and defend it — which transform is standard for revenue metrics, what evidence lets a team deviate, and how disagreement between trimmed and winsorised readings is escalated rather than quietly picked between.
## The two operations Given a threshold `c` and a user's metric value `X`: - **Winsorising (capping)** replaces `X` with `min(X, c)`. The user stays in the analysis, contributing `c`. - **Trimming** removes the user from the analysis entirely when `X > c`. They look like variations of the same idea. In a randomised experiment they are not. ## Estimands Winsorising leaves you estimating the treatment effect on `E[min(X, c)]` — a well-defined quantity over the whole randomised population. Every user assigned is a user counted. Trimming leaves you estimating a difference between conditional means: `E[X | X <= c, treated] - E[X | X <= c, control]`. That conditions on the outcome, and the outcome is post-treatment. This is the classic post-treatment conditioning error. The treatment may itself determine who ends up above `c`, so the two surviving subpopulations are not the same kind of people any more, and the difference between them is not a causal effect on any pre-specified population. ## Why arm sizes matter Randomisation gives you two arms that are comparable in expectation. Trimming breaks the accounting: if the treatment makes users spend more, more treatment users cross `c` and are deleted, so the treatment arm loses proportionally more of its best outcomes. The observed difference can then be pushed toward zero or below it purely by the deletion rule — the very effect you are hunting is being selectively removed from one arm. A useful check when someone insists on trimming: report the number and proportion dropped in each arm. If those proportions differ by more than noise, the trim is itself reacting to the treatment and the comparison is contaminated. ## The standard-error trap Even outside a causal setting, a trimmed mean does not have standard error `s / sqrt(n_kept)` computed on the surviving values. The surviving sample has had its most variable members removed, so its sample standard deviation understates the variability of the estimator, and naive intervals come out too narrow. Correct inference for a trimmed mean uses a winsorised variance in the standard error rather than the trimmed sample's own variance. Winsorising avoids this entirely: you keep all `n` units and compute the usual standard error on the transformed values. ## When trimming is nevertheless defensible Trimming has honest uses in this context: - **Removing non-users.** If a value is impossible — a test account, a currency error, a duplicated charge — that is a data-quality fix, not tail control, and the unit should be excluded from both arms by a rule that has nothing to do with the treatment. - **Sensitivity reporting.** Publishing a trimmed comparison next to the winsorised primary, with the drop counts per arm, is a legitimate robustness exhibit. It just does not get to be the number the launch decision hangs on. ## Choosing between them in practice Default to winsorising for the decision metric because it preserves the intent-to-treat population, keeps the arms balanced by construction, gives you a clean estimand you can state in one sentence, and keeps standard-error computation ordinary. Reserve trimming for sensitivity, and if a trimmed and winsorised analysis disagree, that disagreement is information: it says the conclusion is being driven by users near or above the threshold, and that is a question about the feature, not about the estimator. ## The sentence that scores 'Winsorising transforms an outcome; trimming selects on an outcome. In a randomised experiment I will transform, not select, because selection on a post-treatment quantity breaks the comparability that randomisation bought me.'
- What estimand does each procedure leave you with?Winsorising estimates the treatment effect on the mean of `min(X, c)`, defined over everyone randomised. Trimming estimates a difference of conditional means among users whose spend fell below `c` — a subpopulation defined by a post-treatment outcome, which is not a causal effect on any population you specified in advance.
- Is there any case where deleting a user is the right call?Yes, when the row is not a real observation: a test account, a duplicate charge, a currency-conversion error. That is a data-quality exclusion decided by a rule unrelated to the treatment, applied identically to both arms and ideally identifiable before assignment. It is a different act from trimming the top of a legitimate distribution.
- How would you show that a trim went wrong?Report the count and proportion of users dropped in each arm. If the treatment arm loses a materially larger share, the deletion rule is responding to the treatment, and the surviving groups are no longer comparable. Pair that with the winsorised result to show how much of the conclusion depends on the deleted users.
saying these in an interview costs you the question
- Treats winsorising and trimming as interchangeable
- Trims by the top one percent within each arm separately
- Never checks how many users each arm lost
- Uses the surviving sample's variance for a trimmed mean interval
- Calls deleting the top spenders standard data cleaning