Why must winsorising cap bounds be computed inside each training fold rather than on the full dataset?
answer
- a cap is a fitted parameter
- who got to see the held-out rows
- the 99th percentile is a few rows
- optimistic validation scores
- store the number, reuse it at serving
basics
~20 sPercentile caps are parameters estimated from data. Computing them over every row lets held-out rows shape their own preprocessing, so validation scores turn optimistic. Fit the bounds on training rows and apply those stored numbers unchanged everywhere else.
solid answer
~50 sA cap bound is a fitted quantity, exactly like a mean used for centring. If you take the 1st and 99th percentiles of the whole table before splitting, the held-out rows have influenced the transformation that will later be applied to them, which is leakage and makes cross-validation slightly optimistic. The effect is worst where it matters most: an extreme quantile is decided by a handful of rows, and on small or heavy-tailed columns those are precisely the rows you were not supposed to see. The discipline also has a production payoff. Fitting on the training fold forces you to persist the two numbers, and those same numbers are what you apply to live rows, so a warehouse pick time of 250 minutes is clipped to the stored cap rather than to a bound recomputed from today's traffic. Track how often the cap binds; a rising share is a drift signal.
code
python · 14 lines# Warehouse pick times in minutes; one cart got stuck for 400.
train = [4, 5, 6, 6, 7, 8, 9, 11, 12, 400]
s = sorted(train)
lo = s[int(0.10 * (len(s) - 1))] # nearest-rank 10th percentile
hi = s[int(0.90 * (len(s) - 1))] # nearest-rank 90th percentile
print(lo, hi) # 4 12
def clip(x, lo=lo, hi=hi):
return min(max(x, lo), hi)
# The held-out fold never contributes to the bounds; it only receives them.
val = [3, 7, 250]
print([clip(x) for x in val]) # [4, 7, 12]go deeper
Be ready to state that the cap values are learned from data, so they must come from the training rows and then be applied to validation and test data unchanged.
Explain why fitting on everything leaks: an extreme percentile is set by a few rows, and if those sit in the held-out fold it has influenced its own preprocessing, biasing the score optimistically.
Show the operational side: persist the bounds with the model, clip live values against them, and monitor the share of clipped rows as a distribution-shift alarm rather than recomputing on the fly.
Own the rule as policy, that every data-derived transformation parameter is fitted per fold and versioned with the model, so that leakage cannot be reintroduced one convenient shortcut at a time.
## A cap is a parameter, not a constant Winsorising replaces every value above an upper bound with that bound, and every value below a lower bound with that bound. Nothing is removed; the row survives with a flattened value. The bounds are usually percentiles, for instance the 1st and 99th of the column. The crucial observation is that those two numbers are **estimated from data**. They are as much a fitted parameter as the mean and standard deviation used for standardisation, or the median used for imputation. Everything you know about fitting therefore applies to them: they must be learned from training rows only, stored, and reapplied unchanged everywhere else. ## What goes wrong when you fit on everything Suppose you compute the 1st and 99th percentiles of warehouse pick time over the entire table, then run 5-fold cross-validation. In each fold, the rows you are about to evaluate on have already contributed to choosing the bound that will be applied to them. The held-out fold is no longer held out with respect to preprocessing. Why this bites harder than it looks: - **Extreme quantiles are estimated from very few rows.** The 99th percentile of a 2,000-row column is determined by about the top twenty values. If a handful of those sit in the validation fold, that fold has effectively negotiated its own cap. - **Heavy tails amplify it.** In a light-tailed, large sample the 99th percentile is stable and barely moves when you drop a fold, so the leakage is small. In a heavy-tailed column, or with a few thousand rows, the bound can shift materially, and the validation error becomes a biased estimate of true performance. - **It hides a real production failure mode.** If the bound is fitted on everything, you never confront the question of what to do with a live value larger than anything in training, because during development there wasn't one. The bias here is usually modest, which is why people talk themselves out of the discipline. That argument is backwards: doing it correctly costs almost nothing once the transformation is part of a pipeline that is fitted per fold, and doing it incorrectly means your model-selection numbers are systematically, if slightly, wrong in the optimistic direction. Small optimistic biases compound when you use those numbers to choose between many candidate models. ## The correct procedure 1. Split first: folds, or train and validation and test. 2. Inside each training portion, compute the bounds from the training rows alone. 3. Apply those bounds to the training rows, and then to the held-out rows without recomputing anything. 4. Refit the bounds on the full training data when you fit the final model, and **persist them** with the model artefact. 5. At serving time, clip incoming values with the persisted bounds. The same rule governs any other outlier treatment that estimates something from data, including a fence computed from a spread statistic or a threshold from a fitted distribution. If a number came out of the data, it must come out of the training data. ## Serving-time behaviour is part of the contract Once the bounds are persisted, three operational questions have clean answers. **A live value exceeds the stored upper bound.** Clip it to that bound and record the event. Do not recompute the percentile from the current batch: that would make the same raw input map to different feature values on different days, which is a silent and very hard-to-debug source of prediction drift. **The clip rate is rising.** Suppose 0.9% of rows were clipped in training, by construction, and last week 12% were clipped. That is not an outlier problem, it is distribution shift: the input population has moved beyond the range you fitted for. Treat it as an alarm and a trigger to retrain, and note that the model is now flattening a large fraction of a feature it once treated as informative. **Segments behave differently.** If pick time distributions differ sharply between warehouses, a single global 99th percentile may bind constantly at a slow site and never bind at a fast one. Either cap per segment, with each bound still fitted on that segment's training rows, or include the segment as a feature so the model can account for the difference itself. ## The one-line version Anything you learn from data belongs to the training fold. A cap is something you learn from data. Therefore the cap belongs to the training fold, and the number you learned is the number you ship.
- The leakage sounds tiny. Is the extra pipeline work worth it?On a large, light-tailed column it is small. But an extreme quantile is set by a handful of rows, so on small or heavy-tailed data the bias is real, and it always points the same optimistic way, which matters when you compare many candidate models on those numbers. The work is also not extra: you need the persisted bound at serving time regardless.
- In production, what should happen when a live row exceeds the stored 99th-percentile cap?Clip it to the stored bound and count the event. Recomputing the percentile from the live batch would make the same raw input map to different feature values day to day. A rising share of clipped rows is a drift alarm and a trigger to refit, not something to patch by moving the bound on the fly.
- Should the same cap bounds apply to every segment of the data?Not necessarily. If pick times differ sharply by warehouse, a global 99th percentile binds constantly at a slow site and never binds at a fast one, which flattens a real signal in one place and does nothing in another. Either cap per segment, each bound fitted on that segment's training rows, or add the segment as a feature and let the model handle it.
saying these in an interview costs you the question
- Computes percentile caps on the full table before splitting
- Recomputes the bound on the validation fold
- Recomputes the cap at serving time from the live batch
- Says leakage cannot matter because capping is simple
- Ships a model without persisting its cap bounds