What changes when you trim 10% from each tail instead of winsorising at the 5th and 95th percentiles?
answer
- one deletes rows, the other rewrites values
- sample size and the rest of the row
- duplicate values pile up at the cutoffs
- the fraction is the robustness dial
- trimmed inference borrows the winsorised variance
basics
~10 sTrimming removes the extreme observations, shrinking the sample. Winsorising keeps every row but replaces extreme values with the percentile cutoffs, so the count is unchanged and ties pile up at both caps.
solid answer
~50 sTrimming 10% from each tail of a column of 100 values discards 20 rows and averages the remaining 80. Winsorising at the 5th and 95th percentiles keeps all 100 rows but overwrites the lowest 5 with the 5th-percentile value and the highest 5 with the 95th, creating ties at both caps. Trimming gives a breakdown point equal to the trimming fraction; winsorising the same way gives a breakdown point equal to the winsorised fraction. The subtle part is the standard error: after trimming you cannot compute the trimmed mean's standard error from the trimmed sample as if it were an ordinary sample of that size — the accepted approach uses the winsorised sample's variance. And the winsorised sample's own variance is compressed, because real spread was replaced by cap values, so it understates dispersion if you read it naively.
go deeper
Know the definitions apart: trimming deletes the extreme observations, winsorising replaces them with the percentile cutoff values while keeping every row.
Explain the mechanics on a concrete column: how many rows each touches, what happens to n, and that the fraction chosen sets the breakdown point.
Show you know the inference trap — a trimmed mean's standard error comes from the winsorised variance — and that a winsorised column's own variance is compressed.
Own the framing that these are estimator choices rather than verdicts on individual points, and decide when the tails are the phenomenon and neither treatment is acceptable.
## The two operations Both techniques bound how much a single extreme observation can influence a summary, and they do it in opposite ways. **Trimming** at a fraction from each tail sorts the data, deletes that fraction from the bottom and the same fraction from the top, and computes the statistic on what remains. Trimming 10% from each tail of 100 values leaves 80 observations. The trimmed mean is the mean of those 80. **Winsorising** at the 5th and 95th percentiles sorts the data and replaces — not removes — every value below the 5th percentile with the 5th-percentile value, and every value above the 95th with the 95th-percentile value. All 100 rows survive; the five smallest now share one number and the five largest share another. On a right-tailed column both pull the mean down toward the bulk. Trimming does it by dropping the tail entirely; winsorising does it by flattening the tail onto its boundary. ## Sample size, ties, and what happens to other columns Trimming changes `n`. That matters if the statistic is fed into anything that assumes a sample size, and it matters operationally when the column belongs to a wider table: dropping rows because one field is extreme discards every other field on those rows too. Winsorising preserves `n` and preserves the rest of the row, at the cost of manufacturing duplicate values. Those ties are real consequences — the column now has a spike of identical values at each cap, its empirical distribution has point masses at the boundaries, and any downstream procedure sensitive to repeated values will notice. ## Breakdown points Both are tunable in the same way. A mean trimmed by a fraction from each tail has a breakdown point equal to that fraction: 10% trimming survives up to 10% arbitrary contamination on a side. A winsorised mean at the 5th and 95th percentiles has a breakdown point of 5%. The trimming or winsorising fraction is exactly the robustness dial, which lets you sit anywhere between the ordinary mean (0% breakdown, maximum efficiency on clean data) and the median (50% breakdown, lowest efficiency among these). Note the comparison in the question is not like for like: 10% trimming at each tail touches twice as much of the data as winsorising at the 5th and 95th percentiles, and buys twice the breakdown point. ## The variance trap This is where candidates go wrong. After trimming, you hold 80 numbers, and it is tempting to treat them as an ordinary sample of size 80 and compute a standard error as the sample standard deviation over the square root of 80. That is wrong. The trimmed sample is not a random sample from anything — its extremes were removed by construction, so its spread understates the variability of the estimator. The accepted approach estimates the trimmed mean's standard error from the **winsorised** sample's variance, dividing by the untrimmed sample size and by a factor reflecting how much was trimmed. The two operations are complementary here: winsorising provides the variance input that makes trimming's inference valid. The winsorised sample carries a mirror trap. Its own sample variance is compressed, because genuine spread in the tails has been replaced by two boundary values. Reported as if it were the column's dispersion, it will understate spread — and the more you winsorise, the more it understates. ## Bias and symmetry Both estimators target the centre of a symmetric distribution and are unbiased for it there. On an asymmetric distribution neither estimates the population mean: symmetric trimming or winsorising removes more weight from the long side, so the result sits closer to the median than the mean. That may be exactly what you want, but it must be a deliberate choice, and it must be reported as such rather than described as the mean. If a genuine total is the target quantity, both are the wrong tools, because neither multiplies back into a sum. ## Which to choose - Prefer **trimming** when the extreme values carry no information you trust and the estimator is what you care about, and when you have the sample size to afford the loss. - Prefer **winsorising** when rows must be preserved because other fields on them matter, when sample size is scarce, or when you want to keep the fact that a value was extreme while removing how extreme it was. - Prefer **neither** when the extremes are the phenomenon under study. Both operations silence exactly the observations a tail analysis exists to describe. ## The framing that matters Neither operation is a judgment about whether a value is real. Both are estimator choices applied to the whole column mechanically, including to points nobody suspects. That distinction is what interviewers probe: applying a robust estimator is a statement about how much influence any single observation should have, not a claim that the extreme values were errors. ## What interviewers listen for The correct mechanics of each operation, the effect on `n` and on ties, the breakdown-point dial, the standard-error trap, and the discipline of reporting the choice and its fraction rather than quietly reshaping the column.
- Why can you not compute a trimmed mean's standard error from the trimmed sample as if it were an ordinary sample?Because the trimmed sample was not drawn at random — its extremes were removed by construction, so its spread systematically understates the estimator's variability and the standard error comes out too small. The accepted approach uses the winsorised sample's variance, scaled by the untrimmed sample size and the trimming fraction. Treating the leftover values as a fresh sample produces confidence intervals that are too narrow.
- Does a symmetric trimmed mean estimate the population mean of a right-tailed distribution?No. Symmetric trimming removes equal counts from both tails, but on an asymmetric distribution the long side carries more of the mean's weight, so the trimmed mean sits below the population mean and closer to the median. That can be the desired quantity, but it must be named as a trimmed mean with its fraction stated. Anyone who needs a true total or expectation should not be handed it.
- When would you winsorise rather than trim?When rows must survive because their other fields matter, when the sample is small enough that discarding 20% of it hurts, or when you want to preserve the information that an observation was extreme while removing how extreme. The cost is manufactured ties at both caps and a compressed sample variance for that column, so anyone reading its spread afterwards needs to know the column was winsorised and at what percentiles.
Trimming sends the tallest and shortest players home; winsorising makes them stand in a doorway and records everyone taller than the frame as exactly frame height.
saying these in an interview costs you the question
- Uses trimming and winsorising as interchangeable words
- Computes a trimmed mean's standard error from the trimmed sample
- Reports a winsorised column's variance as the real dispersion
- Calls a symmetrically trimmed result the population mean
- Applies either silently without recording the fraction used