skip to content

How do you stratify cross-validation folds when the regression target is continuous?

level: middleimportance: nice to knowfreq 30%

answer

  1. borrow the trick from classification
  2. you need a discrete key
  3. quantile bins of the target
  4. check the tail reaches every fold

basics

~20 s

Turn the continuous target into a temporary discrete label — quantile bins such as quartiles — and stratify the folds on that bin. Every fold then carries a similar spread of target values, steadying fold metrics on small or skewed data.

solid answer

~50 s

Stratification needs a discrete key, so for regression you manufacture one: cut the target into quantile bins and stratify on the bin label. On a right-skewed student test-score target, split into quartiles and deal each quartile across the folds. Without it, a small dataset can hand one fold most of the top-tail rows and another almost none, and the fold metrics move for reasons that have nothing to do with the model — per-fold R-squared is especially sensitive, because it measures error against that fold's own target variance, so a fold with a compressed target range looks bad at identical absolute error. Choose bins by quantile rather than equal width on skewed data, and keep the bin count low enough that every bin can supply rows to every fold. With a large sample and a mild distribution, random folds already match and the technique adds little.

go deeper

for a junior

Know that the idea transfers from classification: you can make a discrete key out of a continuous target by binning it. Recognising the term quantile bin is enough here.

for a middle

Explain the full construction — quantile cuts, stratify on the bin index, discard the bin afterwards — and why quantile beats equal-width on a skewed target. Be ready to justify a bin count.

for a senior

Show you diagnose the need rather than applying it reflexively: noticing fold metrics moving more than the model should, knowing R-squared is the most exposed metric, and answering the leakage objection cleanly.

for a principal

Frame it as an estimate-quality decision: how much of the observed variation between candidate models is split noise, whether the tail of the target is where the business value sits, and when the sample is simply too small for any k-fold comparison to settle the question.

## Why regression folds need it at all Stratified k-fold is usually taught as a classification device, and the reason is mechanical: stratification balances a *discrete* key across folds, and regression has no discrete key. The underlying motivation carries over unchanged though. A validation fold whose target distribution does not resemble the population produces a metric that does not resemble the population metric, and the fold-to-fold scatter you then see is split noise rather than model behaviour. The symptom is loudest with a right-skewed target — student test scores where most of the mass sits mid-range and a thin tail runs high, house prices, claim amounts, session durations. With 300 rows and 10 folds, the twenty highest-value rows are dealt at random across the folds; one fold may get five of them and another none. The fold that got five is asked to predict values the training data barely covers, and its error is large; the fold that got none is scored only on the easy middle. ## The construction Bin the target, then stratify on the bin. Concretely: sort the target, cut it at the quartiles (or deciles, depending on sample size), attach the bin index to each row as a temporary label, and run stratified k-fold on that label. Each fold then holds roughly a quarter of its rows from each quartile of the target distribution. Use *quantile* bins — equal numbers of rows per bin — rather than equal-width bins. On a skewed target, equal-width bins put almost every row into the first bin and a handful into the last, which reproduces the problem you were trying to solve. Quantile bins are equal-count by construction, so each bin has enough rows to be spread across the folds. The bin label exists only for the splitter. It is discarded afterwards; the model still trains on the continuous target with a regression loss, and the metrics are still regression metrics. ## How many bins More bins mean a tighter match between the fold and population target distributions, and fewer rows per bin. The binding constraint is the same one that limits k for a rare class: a bin with fewer rows than there are folds cannot be represented in every fold. As a rule of thumb, keep the number of bins small enough that each bin holds several times k rows — with 300 rows and 10 folds, quartiles give 75 rows per bin, comfortably enough; twenty bins would give 15 per bin, which is already marginal. Quartiles or quintiles are a good default; deciles are worth it only on larger samples. ## Why per-fold R-squared is the most sensitive metric R-squared on a fold is one minus the ratio of the fold's squared error to the fold's total squared deviation from its own mean target. The denominator is the fold's target variance. Two folds with exactly the same absolute prediction error will report different R-squared values if one has a wider spread of targets than the other, and the narrow fold will look worse. Balancing the target distribution across folds removes that source of movement. Metrics defined purely on absolute error — mean absolute error, root mean squared error — are less exposed to it but still move when one fold holds a disproportionate share of extreme values, because large targets tend to carry large errors. ## Is using the target to build the folds a form of leakage? No, and this is worth being able to answer crisply. Leakage means information from the held-out rows reaches the model — through a feature, a fitted parameter, a preprocessing statistic. Binning the target uses the labels only to decide which row goes into which fold. No feature is derived from it, no parameter is fitted on held-out values, and the model never sees a validation row's target. It is exactly the same operation stratified classification folds perform, and nobody calls that leakage. The rule to hold onto is that the split may look at the label; the model may not. ## When to skip it With a large sample and a target that is not badly skewed, random folds already match the population closely and stratifying on bins changes nothing measurable. It is worth reaching for when the sample is small, when the target is heavily skewed or multimodal, when a rare high-value region is what the business actually cares about, or when you have noticed fold metrics moving more than the model plausibly does.

  • How many bins should you use, and what limits the number?
    Quartiles or quintiles are a sensible default; deciles need a larger sample. The limit is the same one that caps k for a rare class: a bin holding fewer rows than there are folds cannot appear in every fold. Aim for each bin to hold several times k rows. Use quantile cuts rather than equal-width cuts, since equal-width bins on a skewed target dump nearly everything into the first bin.
  • Isn't using the target to construct the folds a form of leakage?
    No. Leakage is held-out information reaching the model through a feature, a fitted parameter or a preprocessing statistic. Here the target is used only to decide which row goes into which fold; nothing is derived from it and the model never sees a validation row's target. It is the same thing stratified classification folds do. The split may look at the label; the model may not.
  • Which regression metric suffers most from unbalanced target distributions across folds?
    R-squared, because it divides the fold's squared error by the fold's own total variance around its mean target. A fold with a compressed target range reports a poor R-squared even at identical absolute error. Mean absolute error and root mean squared error are steadier, though they still move when one fold collects a disproportionate share of extreme targets.

saying these in an interview costs you the question

  • Says stratification only applies to classification problems
  • Uses equal-width bins on a heavily right-skewed target
  • Bins so finely that some bins hold fewer rows than folds
  • Believes binning the target for splitting leaks it into the model
  • Keeps the bin label as a feature after the split is built

context