Your demand targets carry rare 50x sensor spikes — how do you choose between Huber, a target transform, and a data fix?
answer
- Classify the spikes before changing anything
- Corruption belongs upstream, not in the loss
- Robustness moves the statistic you estimate
- Transforms change what you predict
- Score spike periods separately from typical ones
basics
~20 sDecide what the spikes are before touching the loss. Corruption belongs in the data pipeline, where the fix is auditable. A robust loss like Huber stabilises training but shifts the statistic you estimate, so it under-predicts spikes that turn out to be real.
solid answer
~50 sThe first question is diagnostic, not modelling: are the spikes corruption (a stuck meter, a unit change, a duplicated ingestion) or genuine rare events? Fixing corruption upstream is the strongest option because it is auditable, applies to every model reading the table, and leaves the objective honest — a robust loss applied to bad data hides the defect instead of reporting it. Huber is right when large residuals are noise you cannot remove: its capped tail gradient stops one row owning an update, but its optimum sits between the conditional mean and median, so genuine extremes get systematically under-predicted where the cost is highest. A monotone target transform is a different move again — back-transforming a squared-error fit on log targets gives roughly a conditional median, not a mean, which quietly under-states totals. Whatever you choose, evaluate spike periods separately rather than letting an average hide them.
go deeper
Know that extreme target values need explaining before they are handled, and that deleting or capping them is a decision with consequences rather than routine cleaning.
Be able to state each lever's mechanism: a capped tail gradient for Huber, a compressed tail for a transform, and removal at source for corruption, plus what each does to the training update.
Show the diagnosis path from extreme residuals to their source, the train/serve symmetry trap, and the back-transform bias that makes log-target totals come out low.
Own the framing that a robust loss trades peak-case fidelity for stability, name who owns the upstream defect you would otherwise absorb permanently, and mandate evaluation that reports spike periods separately.
## The decision has three candidate levers and one prerequisite The prerequisite is classification of the spikes. Everything else follows from it, and skipping it is the most common senior-level mistake — reaching for Huber because the training curve looks bad is treating a symptom. Inspect the extreme rows as records, not as numbers: - **Corruption.** They cluster on one device, one ingestion window, one upstream job. The value is physically impossible for the entity. A retry duplicated a row. A unit changed. Corroborating sources disagree. - **Genuine rare events.** They recur with an explanation — a cold snap, a promotion, an outage elsewhere shifting load. Independent sources agree. Domain experts recognise them. - **Scale or definition drift.** The target quietly changed meaning: an aggregation window widened, a hierarchy level changed, a new market with a different magnitude was folded in. ## Lever one: fix the data For corruption this dominates the alternatives, and the argument is organisational as much as statistical. A pipeline-level fix — validation rules with physical bounds, an explicit quarantine table, a nullable flag rather than a silent overwrite — is visible, testable, and shared by every consumer of that table. A robust loss achieves a similar training curve by *absorbing* the defect, which means the defect keeps arriving, nobody is paged, and the next model built on the same table rediscovers it. Prefer masking or excluding a corrupted target over imputing a plausible-looking value: an imputed number becomes indistinguishable from a real one within a week. The operational trap is train/serve asymmetry. If you clean targets in training but the same faults appear in features at serving time, you have built a model that has never seen the conditions it will meet. Clean both sides or neither. ## Lever two: change the loss Huber earns its place when the large residuals are irreducible noise rather than removable defects — noisy labels, an inherently heavy-tailed process, an aggregation you do not control. Mechanically it keeps squared error's behaviour for small residuals and caps the gradient magnitude in the tail, so a single extreme row cannot own a mini-batch update, and unlike absolute error it still shrinks the gradient to zero as the fit tightens. The cost is a change of estimand. Squared error's optimum is the conditional mean; absolute error's is the conditional median; Huber's sits between them. On a right-skewed demand target that means predictions move down. If the extremes were real, the model now misses exactly the days that matter most, and — worse — its aggregate predictions under-shoot totals, which is the failure a downstream planning system notices last and complains about loudest. So the honest framing to a stakeholder is: a robust loss buys training stability and typical-case accuracy at the price of peak-case fidelity. That is a business tradeoff, not a technical detail. ## Lever three: transform the target A monotone transform (log-like on a positive target, or a per-entity scaling) compresses the tail so ordinary squared error stops being dominated by it. It is attractive for multiplicative processes, where relative error is the natural notion and a fixed absolute error means something quite different at 10 and at 10,000. Two hazards. First, you are now optimising error in transformed space, so the model is tuned to relative rather than absolute misses — fine if that matches the decision, misleading if it does not. Second, back-transforming is not neutral: exponentiating a fitted conditional mean of log-targets gives roughly the conditional *median* of the original target, not its mean, because the exponential of an average is not the average of exponentials. Predictions and especially summed totals come out low, and this is frequently mistaken for a modelling failure. If you need means, you need an explicit correction or a loss defined on the original scale. ## Combining, and what to hold the model to These are not exclusive. A defensible design is often: fix and quarantine known corruption upstream; use a robust or transformed objective for the residual heavy tail that remains; and, where the extremes are real and consequential, treat them as their own problem — a feature that flags the driving condition, a separate model for event periods, or an objective that predicts a spread rather than a point. Whatever you pick, evaluation must change with it. An average error over all periods lets a model that ignores spikes look excellent, because spikes are rare by definition. Report typical-period and spike-period performance separately, keep a held-out set of known real events, and state which statistic the model is now estimating so that downstream consumers of the number are not silently misled. ## The organisational question Finally, decide who owns the defect. If the sensor pipeline belongs to another team, a robust loss is the modelling team unilaterally deciding to tolerate their bug forever, and that decision should be explicit and revisited rather than buried in a config. Data quality that is absorbed by models stops being measured.
- How do you distinguish a corrupted target from a genuine rare event in practice?Treat the extreme rows as records. Check whether they concentrate on one device, one ingestion window, or one job; whether the value is physically possible for that entity; whether an independent source agrees; and whether the pattern recurs with a plausible driver such as weather or a promotion. Corruption clusters on infrastructure; real events cluster on explanations, and domain experts recognise them immediately.
- What breaks if you clean targets in training but leave the serving pipeline untouched?You get a model trained on a distribution that never occurs at serving time. If the same fault also perturbs features, the model meets inputs it has never seen and behaves unpredictably exactly when the pipeline is misbehaving. Cleaning must be a shared transformation applied identically on both sides, or the fault must be detectable at serving time and handled explicitly rather than silently scored.
- The extremes turn out to be real and business-critical. What do you change besides the loss?Give the model the information that explains them — the weather front, the promotion calendar, the upstream outage — so the spike is predictable rather than an unmodelled shock. Consider a separate treatment for event periods, and evaluate on a held-out set of known events instead of an all-periods average. If the timing is genuinely unknowable, predicting a spread rather than a point is more honest than a point estimate that is always wrong.
saying these in an interview costs you the question
- Reaches for a robust loss without diagnosing the spikes
- Deletes extreme rows without checking whether they are real events
- Assumes exponentiating a log-target prediction recovers the mean
- Ignores that robustness changes which statistic is estimated
- Evaluates only an all-period average that hides spike-day failures
- Cleans targets in training but not at serving time