skip to content

When is discretising a continuous predictor into bins worth the information it destroys?

level: principalimportance: should knowfreq 36%

answer

  1. lossy compression of a feature
  2. who reads it, a person or a model
  3. step function with edges you chose
  4. trees already search their own cutpoints
  5. buy explainability, pay in resolution

basics

~20 s

Rarely for accuracy, often for everything else. Binning discards within-bin variation and imposes arbitrary cutpoints, so it usually costs predictive power. It earns its place when bins buy interpretability, stable reporting, or let a linear model express a non-monotone effect.

solid answer

~50 s

Binning is a lossy compression of a feature, so start from what it costs: all variation inside a bin is thrown away, two customers either side of an edge are treated as different while two at opposite ends of one bin are treated as identical, and the resulting step function is discontinuous at boundaries you chose. That normally reduces predictive power, and it buys nothing for a gradient-boosted tree ensemble, which already searches cutpoints and will pick better ones. The legitimate reasons are not accuracy. Bins give a rule a non-technical audience can read and sign off, make cohort reporting stable across periods, cap the influence of a wandering tail, and let a plain linear model represent a U-shaped or threshold effect a single coefficient cannot. My rule: bin when a human consumes or approves the output, stay continuous when a model is the only consumer.

go deeper

for a junior

Know that binning throws information away: everything inside a bin becomes identical, and the edges are choices someone made rather than facts in the data.

for a middle

Explain the concrete costs — lost within-bin variation, a discontinuity at each edge, extra hyperparameters — and why a tree ensemble that finds its own splits gains nothing from your bins.

for a senior

Show you measure the tradeoff: same folds, continuous versus binned, and a number for what the bins cost. Then justify the choice by who consumes the output and how stable it must be.

for a principal

Own the policy. Decide when the organisation trades accuracy for a rule a committee can approve, who owns bin definitions once several teams report against them, and how edges are versioned when the population drifts.

## What binning actually does Discretising replaces a continuous feature with a small set of levels: `age` becomes `under 25`, `25-34`, `35-49`, `50+`. Formally the model is now a step function of that feature — constant within each bin, jumping at each edge. That is a strong structural assumption, and unlike most modelling assumptions it is imposed by hand. ## The costs, stated precisely **Within-bin variation is discarded.** If the true effect varies smoothly across a bin, that variation is gone and cannot be recovered downstream. A 25-year-old and a 34-year-old become the same input. **Boundaries are arbitrary and discontinuous.** Two rows a fraction apart, straddling an edge, receive different treatment; the model has a cliff at a threshold that the data did not necessarily put there. In anything consumer-facing this is also a fairness and appeals surface: 'why did my rate change when I turned 35?' **Statistical power falls.** Replacing a continuous predictor with a handful of indicators typically weakens the measured relationship, because you are estimating a coarse approximation of it with the same amount of data. **The edges are hyperparameters.** Every cutpoint you choose is a decision fitted, formally or informally, to data. Choosing edges by scanning many candidate splits while watching the outcome overfits, and the resulting in-sample lift will not survive a fresh sample unless the search happened inside cross-validation folds. **Tree ensembles gain nothing.** A gradient-boosted or bagged tree already selects thresholds greedily and can place many of them, adapting depth by depth. Feeding it your ten fixed bins replaces a search over all cutpoints with a search over nine, and the resolution you removed is unrecoverable. Pre-binning to 'help' a tree signals a missing model of how trees split. ## The benefits worth paying for **Human consumability.** A pricing tier, an eligibility rule, an underwriting guideline or a segment definition has to be read, argued about and signed off by people who will never see a coefficient. Bins are the format that conversation happens in, and a slightly less accurate model that gets approved beats a better one that does not. **Stable reporting.** Cohort reporting needs categories that mean the same thing across periods and teams. Fixed bins give that; a continuous feature does not, and quantile bins recomputed each period do not either. **Robustness to extremes and drift.** A top bin defined as 'over 50 GB' behaves the same whether the maximum this month is 80 GB or 800 GB. The model's exposure to a wandering tail is capped by construction, which also makes monitoring simple: watch the share of the population in each bin. **Non-monotone and threshold effects in a linear model.** A single coefficient forces a monotone linear effect. Bin indicators let the model say 'risk is high for the very young, low in the middle, high again for the very old' with no functional form imposed. If the underlying process really does have a legal or contractual threshold — a plan allowance, an age of majority, a deductible — the step function is not an approximation, it is the truth. **Missing gets a home.** A bin scheme absorbs 'missing' as an explicit level rather than requiring an imputed number, which is often more honest. ## How to decide, in practice 1. **Who consumes the output?** A person who must approve or explain a rule pushes toward bins; a downstream model pushes away from them. 2. **What model family?** Tree ensembles: keep it continuous. Linear or additive scoring models: bins are a legitimate way to buy flexibility and readability. 3. **Do domain thresholds already exist?** If the business has real cutpoints, use them. Domain edges cost no statistical power to choose and need no defence. 4. **Can you afford the accuracy?** Measure it. Fit both versions under the same cross-validation and put a number on what the bins cost. 'Two points of AUC for a rule the credit committee will approve' is a decision someone can make; 'we binned everything because it is tidier' is not. 5. **If you must choose edges from data,** choose them from marginal quantiles or domain meaning rather than by scanning against the outcome, and treat any tuned edges as parameters fitted inside the training folds. ## The failure mode to name The default-binning habit — discretising every continuous feature as a preprocessing reflex — is the version of this that shows up in review. It costs accuracy silently, adds a pile of hyperparameters nobody tuned, and is usually inherited from a workflow where the model family made it necessary. Binning should be a deliberate, justified choice on specific features, with a stated consumer.

  • How do you choose the number of bins without overfitting?
    Prefer edges that come from the domain — contractual thresholds, plan allowances, regulatory ages — since those cost no degrees of freedom. Failing that, use marginal quantiles of the training fold. If you genuinely tune the count, tune it inside cross-validation and treat both the count and the edges as fitted parameters, never chosen on the full dataset.
  • Why does pre-binning a continuous feature usually make a gradient-boosted tree slightly worse?
    Because the model already searches cutpoints itself and can place far more of them than you will. Handing it ten fixed bins restricts it to nine possible thresholds on that feature and permanently discards the resolution in between. Trees are also invariant to monotone rescaling, so none of the usual shape arguments apply to them.
  • How do you keep binned features stable when the underlying distribution drifts?
    Freeze the edges from a reference period and monitor the population share landing in each bin. Recomputing quantiles every period silently redefines the feature, so a model reading it sees a different variable over time and drift becomes invisible — the very thing you needed the monitoring to catch.
  • What would you measure before agreeing to bin a feature for a production model?
    The accuracy delta under identical cross-validation: same folds, same model, continuous versus binned. That turns an aesthetic argument into a number a decision-maker can weigh against the interpretability or stability benefit being bought. If the delta is negligible and a human must approve the rule, bin it.

saying these in an interview costs you the question

  • Bins every continuous feature by default as a preprocessing habit
  • Believes binning removes outlier influence at no cost
  • Chooses cutpoints by scanning many splits against the outcome, then reports in-sample lift
  • Assumes a boosted tree benefits from hand-crafted bins
  • Ignores the discontinuity a bin edge creates for rows just either side
  • Cannot quantify what the binned version costs in accuracy

context