skip to content

A $180,000 B2B order sits in a consumer marketplace's GMV column; do you drop, cap or keep it?

level: seniorimportance: must knowfreq 57%

answer

  1. start with the population, not the number
  2. will you score rows like it
  3. corrupt, out of scope, or just rare
  4. exclude by business rule, not percentile
  5. capping the target caps the prediction

basics

~20 s

Ask whether B2B orders will be scored in production. If so, keep the row and represent that segment explicitly; deleting it trains a model blind to your most valuable orders. Drop only corrupt or out-of-scope rows.

solid answer

~50 s

Run three checks in order. First, is it real? Verify against the order record, because a stray zero or a currency mix-up is a data bug and gets fixed, not capped. Second, is it in scope? If B2B orders will be scored by this model, the row is signal, and removing it means the model has never seen the range it must predict; keep it and give the model a way to distinguish the segment, such as a buyer-type feature or a separate model. If B2B is genuinely out of scope, filter it by the business rule, not by a percentile, because a percentile filter also deletes genuine large consumer orders. Third, if the row stays but destabilises fitting, cap the feature rather than the target: capping the target makes the model structurally incapable of predicting a large order. And never remove extremes from the validation or test set.

go deeper

for a junior

Be ready to say that you check whether the value is genuine before doing anything, and that removing real rows is a decision with consequences rather than a default cleaning step.

for a middle

Explain the three options concretely: drop, cap or keep, what each does to the fitted model, and why capping a target is different in kind from capping a feature.

for a senior

Demonstrate that you frame it as a scoping question about the production population, exclude by business rule rather than percentile, and refuse to clean the evaluation set.

for a principal

Own the organisational risk: a percentile filter left in a pipeline can silently remove a growing, valuable segment, so the exclusion rule needs an owner, a documented reason and a monitored row count.

## The decision is about population, not about distance The common instinct is to reach for a numeric rule: anything beyond some percentile goes. That instinct answers the wrong question. The real question is **whether rows like this one belong to the population your model will be asked to score**. A value can be a hundred times the median and be perfectly legitimate; a value can be well within the bulk and still be corrupt. So work through four questions in order. ### 1. Is the value real? Check the source of truth. A GMV of $180,000 in a marketplace where the median order is $40 has three plausible explanations: a genuine bulk purchase, a unit or currency error, or a test transaction that leaked into production data. Only the first is an outlier. The second is a bug to fix upstream, and the third is a row to exclude by its own flag. This step costs one query and settles the question more often than people expect. ### 2. Will rows like it be scored in production? This is the pivotal question. If the model will see B2B orders at serving time, deleting them from training produces a model that has never observed the range in which your revenue concentrates. It will predict badly on precisely the rows worth the most money, and the offline metric will not show it, because you deleted those rows from the evaluation data too. That is the failure mode that turns *cleaning* into *silently scoping the model down*. If the extremes are in scope, the useful moves are about **representation** rather than removal: - Add the feature that explains the difference. Buyer type, contract flag, channel. Once the model can tell a B2B order from a consumer one, the value is no longer inexplicable. - Consider a separate model for the segment if it behaves differently enough and there is sample to support one. - Accept that a handful of rows cannot teach a model much on their own, and be explicit about that limitation rather than hiding it by deleting them. ### 3. If they are out of scope, exclude them by rule If the product genuinely only serves consumers and B2B traffic will never be scored, then excluding those rows is correct. But exclude them **by the business rule that defines the boundary** (buyer type is business), not by a numeric threshold. A percentile filter is a proxy that misfires in both directions: it deletes the genuine large consumer orders and keeps the small B2B ones. Write the rule down, count the rows it removes, and monitor that count over time. ### 4. If they stay but destabilise fitting, cap the feature, not the target Capping is the middle path: the row is kept, its value is flattened to a bound. It is a reasonable choice when the value is genuine but you only need the model to behave sensibly across the bulk range and its accuracy on the far tail is not what you are optimising. One asymmetry decides most of these cases: **capping a feature limits the influence of an input; capping the target changes what the model is allowed to say**. If you winsorise GMV as a target at the 99th percentile, the model becomes structurally unable to predict anything above that bound, and every large-order prediction is biased low. It is not merely inaccurate, it is incapable. That is only acceptable when the business explicitly does not care about the tail, which is rarer than people assume. ## The evaluation set is not yours to clean Whatever you do to the training data, the validation and test sets must keep the extremes they naturally contain. If you remove them from evaluation as well, the reported error describes a population you will never face, and it will look better than production every time. Preprocessing does apply to evaluation rows, but only using bounds fitted on training data; **row removal is a different act from transformation** and does not carry over. A good habit is to report error separately for the bulk and the tail. That makes the cost of a cap explicit: you can see exactly how much accuracy on large orders you traded for stability on the ordinary ones, and let the business decide whether that trade is acceptable. ## Documenting the decision Whatever you choose, record it as part of the model definition: the rule, the number of rows affected, and the reason. Two failure modes follow from skipping this. Somebody re-runs the pipeline six months later without the filter and cannot reproduce your numbers. Or the filter silently starts removing 8% of rows instead of 0.2% because the business changed, and nobody notices because it was a percentile with no alarm attached to it. ## The short version Drop when corrupt or provably out of scope, by rule. Cap when genuine but disruptive, and cap inputs before you ever consider capping the target. Keep, and represent, when the row is genuine and in scope, which is more often than the reflex suggests. And never touch the evaluation set.

  • Two executive compensation packages dominate the salary column of an HR attrition model. Same call?
    Similar test, different answer, because salary is a feature here and only two rows carry it. If executives are in scope, keep them and either cap the salary feature so distance-based and coefficient-based models are not dominated, or add a job-level feature that explains the gap. If executives are outside the retention programme, filter them by role, not by salary. Note too that two rows are individually identifiable, which is its own concern.
  • When is dropping the row genuinely the right answer?
    When it is impossible, such as a negative order value or a timestamp before launch; when a business rule proves it is outside the scored population; or when it is a duplicate or an internal test transaction. In every case the criterion is a stated rule rather than a distance from the median, and you count the removals and watch that count over time.
  • How do you show that capping actually helped rather than just felt tidy?
    Train with and without the cap and compare on the same untouched validation set, then report error separately for the bulk and for the tail. That makes the trade explicit: stability across ordinary orders bought at some cost in accuracy on large ones. If the tail cost is invisible in your headline metric, you have not measured the thing the cap actually changed.

saying these in an interview costs you the question

  • Deletes anything beyond a percentile as routine cleaning
  • Removes extreme rows from the test set as well
  • Caps the target and still expects large predictions
  • Filters a business segment using a numeric threshold
  • Treats a rare but genuine row as an error to erase

context