skip to content

In gradient boosting, what do row subsampling and column subsampling per round buy you?

level: middleimportance: should knowfreq 46%

answer

  1. each round sees only part of the data
  2. trees become less alike
  3. a noisier gradient is the point
  4. cheaper rounds come for free
  5. too little data makes stopping unreliable

basics

~20 s

Fitting each round's tree on a random fraction of rows and columns decorrelates successive trees, adds a regularising noise to the gradient estimate, and cuts per-round cost. It usually validates better than the full-data fit, until the subsample is so small the fit turns unstable.

solid answer

~50 s

Stochastic boosting fits each round's tree on a random subsample of the rows — commonly 0.5 to 0.8, drawn without replacement — and often on a random subset of the columns as well. Row subsampling makes successive trees see different data, so they correlate less and the ensemble's variance drops; it also injects noise into the gradient estimate, which acts as a regulariser, and it cuts the per-round training cost roughly in proportion. Column subsampling stops every tree keying on the same one or two dominant features and speeds up the split search. On a two-million-row impression log, 0.8 by row and 0.6 by column is a typical setting that trains faster and usually validates better than the full-data fit. Pushed too far, the gradient estimate becomes so noisy that the validation curve wobbles, the stopping round becomes unreliable, and rare-positive problems suffer.

go deeper

for a junior

Know that boosting can fit each tree on a random slice of the rows and columns rather than all of them, and that this is normal practice rather than an exotic option.

for a middle

Explain both effects — decorrelating successive trees and adding noise to the gradient estimate — and note that the per-round cost falls roughly in proportion to the fraction kept.

for a senior

Show judgment about how far to push it: rare positives, jagged held-out curves and an unreliable stopping round are the failure modes you should be able to name from experience.

for a principal

Frame subsampling as one of several interacting regularisers and set team defaults that keep training affordable without making the stopping signal too noisy to trust.

## The change to the algorithm Ordinary gradient boosting fits round `m`'s tree to the pseudo-residuals of **every** training row. Stochastic boosting changes one line: before fitting, draw a random fraction of the rows without replacement — say 80% — and fit the tree to the residuals of that subsample only. The resulting tree is then shrunk by the learning rate and added to the ensemble exactly as before, and it scores all rows on the next round. Column subsampling is the same idea on the other axis: each tree (or, in some implementations, each level or each split) may only consider a random subset of the features, say 60% of them. Both are cheap, both are on by default in most modern boosting practice, and both were introduced as regularisers rather than as speed tricks — though they are also speed tricks. ## What row subsampling actually buys **Decorrelation.** Successive trees in a boosted ensemble are highly dependent by construction: each one is fitted to what its predecessors left behind. If every round sees the identical data, every round chases the identical quirks in it. Resampling the rows means that a particular unlucky cluster of points only drives some of the trees, so the errors the trees make are less aligned and the summed model has lower variance. **A noisy gradient, deliberately.** The residuals of an 80% subsample are an unbiased but noisy estimate of the residuals of the whole set. Fitting to a noisy target prevents the ensemble from resolving fine detail that is not stable across resamples. This is the same reasoning that makes noisy optimisation generalise better in other settings, and it is why stochastic boosting frequently beats the full-data version on validation, not merely ties it. **Cost.** Fitting on 80% of the rows costs about 80% as much per round. On a two-million-row impression log that is a straightforward saving on every one of thousands of rounds. ## What column subsampling buys When one or two features dominate, an unconstrained split search will pick them near the root of nearly every tree, and the ensemble ends up as many near-copies of the same story. Restricting each tree to a random 60% of the columns forces some trees to find the second-best structure, which is exactly the structure that would otherwise be masked by a correlated dominant feature. The ensemble covers more of the signal, and the split search is proportionally cheaper because there are fewer candidate features to score. ## How it interacts with the other knobs - **With the learning rate.** Subsampling and a small rate are complementary regularisers, and both are usually applied together. Neither substitutes for the other: subsampling changes *what* each tree sees, shrinkage changes *how much* of each tree is admitted. - **With the stopping decision.** This is the interaction people miss. Aggressive subsampling makes the validation curve noticeably jagged, because the model is a sum of noisier increments. A jagged curve makes the round with the best held-out score less trustworthy — the apparent minimum may be a wiggle. Milder subsampling, or averaging the curve across folds, restores a clean signal. - **With rare positives.** A 30% row subsample of a dataset with a 0.5% positive rate can leave a round's tree with very few positives to learn from. When positives are scarce, subsample less aggressively. ## Choosing values There is no universal setting, but the shape of the tradeoff is predictable. Rates in the 0.5-0.9 range for rows are the normal territory; below roughly 0.5 the gradient noise usually starts costing more than the regularisation buys, and at 1.0 you have plain boosting with none of the benefit. Column fractions are chosen against how correlated and how numerous the features are: many redundant features tolerate an aggressive fraction, a handful of carefully engineered features do not. The honest way to set both is empirical: they are cheap to vary, their effect shows up clearly on a held-out curve, and unlike the round count they transfer reasonably across nearby learning rates. ## The common misconceptions Subsampling here is drawn **without** replacement per round, and its purpose is decorrelation and regularisation of a sequential fit — it is not an averaging scheme over independently trained models. It also does not guarantee that every row is used, nor that each row is used equally often; over thousands of rounds the coverage evens out statistically, but no round makes any promise about a particular row.

  • Would you set the row subsample to 0.3 on a dataset with a 0.5% positive rate?
    No. A 30% subsample of an already tiny positive class can leave a round with almost nothing to learn from, so the trees fit noise and the validation curve becomes erratic. With rare positives I keep the row fraction high — 0.8 or more — and lean on the learning rate and the stopping round for regularisation instead.
  • How does subsampling change how you read the held-out curve?
    It makes it jagged. Each increment is noisier, so the curve wiggles and the single best-scoring round may be a fluctuation rather than the true optimum. I either soften the subsampling, average the curve over several folds, or refuse to trust a minimum that is not visibly separated from its neighbours.
  • Is column subsampling doing the same job as row subsampling?
    Related but not the same. Row subsampling varies which examples a tree explains; column subsampling varies which explanations it is allowed to reach for. Columns specifically attack the case where one dominant feature is picked at the root of every tree and masks correlated alternatives — row resampling would not break that.

saying these in an interview costs you the question

  • Says every row is guaranteed to be used exactly once
  • Treats subsampling as a pure speed trick with no effect on fit
  • Claims subsampling replaces the need for a small learning rate
  • Uses an aggressive row fraction with very rare positives
  • Ignores that a noisier curve makes the stopping round less reliable

context