skip to content

Why does gradient boosting fit pseudo-residuals instead of plain residuals?

level: middleimportance: should knowfreq 58%

answer

  1. The name of the algorithm is a hint
  2. Derivative with respect to what?
  3. Squared error is the lucky special case
  4. For log loss the score is log-odds
  5. Label minus predicted probability

basics

~20 s

Each tree fits the negative gradient of the loss at the current prediction, not the raw error. For squared error that equals the plain residual; for log loss it is the label minus the predicted probability.

solid answer

~50 s

The pseudo-residual for a row is `-dL/dF` evaluated at the model's current score for that row -- boosting is gradient descent in function space, so the derivative is taken with respect to the *prediction*, not with respect to split thresholds or features. With squared error written as `0.5*(y - F)^2` this is exactly `y - F`, which is why the folk description -fit the residuals- survives. Change the loss and the quantity changes with it. Under log loss the score `F` lives on the log-odds scale, `p` is the logistic of `F`, and the pseudo-residual is `y - p` -- a number in `(-1, 1)`, and the tree's output is an increment in log-odds, not a probability. Under absolute error it is just the sign of the error. That generality is what makes it *gradient* boosting rather than residual boosting.

go deeper

for a junior

Know that -fit the residuals- is shorthand and that the real target depends on the loss you chose. Being able to say that squared error is the case where the two coincide already puts you ahead at this level.

for a middle

Be ready to derive the two standard cases on a whiteboard: squared error gives y minus prediction, log loss gives label minus predicted probability. State clearly that the derivative is with respect to the prediction for that row.

for a senior

Demonstrate scale discipline in a real problem: name the constant start, the scale the score accumulates on, and the form of the pseudo-residual before you touch anything else. Show why a count or rate target needs a loss whose model lives on the log scale.

for a principal

The strategic point is that a boosting stack is a loss-swappable framework, so the business objective can be encoded in the loss rather than patched afterwards. Be ready to argue when that is worth the extra explanation cost versus fitting a familiar loss and post-processing.

## From residual to pseudo-residual The textbook one-liner for gradient boosting -- *each tree is fitted to the residuals* -- is true for exactly one loss function and misleading for every other. What each tree is actually fitted to is the **negative gradient of the loss with respect to the current prediction**, evaluated row by row. That quantity is called the *pseudo-residual*, and only for squared error does it coincide with the plain arithmetic residual. ### The derivative that matters For a single training row with target `y`, the model produces a score `F`. The loss `L(y, F)` is a function of that score, and the pseudo-residual is ``` r = -dL/dF evaluated at F = F_{m-1}(x) ``` Note carefully what the derivative is taken with respect to: the model's **output value** for that row, not the tree's split thresholds and not the feature values. Boosting is gradient descent in *function space* -- the algorithm asks -- in which direction should the prediction for this row move to reduce its loss -- and then fits a tree to that field of directions so it can move whole regions of feature space at once. ### Four losses, four different -residuals- **Squared error.** Write the loss as `L = 0.5 * (y - F)^2`. Then `dL/dF = -(y - F)`, so `r = y - F`. This is the ordinary residual, and it is where the folk description comes from. **Log loss (binary classification).** Here the model score `F` lives on the **log-odds** scale and the probability is `p = 1 / (1 + exp(-F))`. With `L = -[y*log(p) + (1-y)*log(1-p)]` the derivative works out to `dL/dF = p - y`, so the pseudo-residual is `r = y - p`. For a clinic no-show model where a patient did attend (`y = 0`) but the model predicted a 0.7 chance of no-show, the pseudo-residual is `-0.7`: the next tree is asked to push that region's log-odds down. Two things follow. First, the target of the tree is a number in the open interval `(-1, 1)`, not a price or a count. Second, the tree's output is an increment in **log-odds**, and probabilities only appear when the accumulated score is squashed through the logistic function at the end. **Absolute error.** With `L = |y - F|`, the derivative is `-sign(y - F)`, so the pseudo-residual is `+1` or `-1` -- the *direction* of the error with its magnitude thrown away. Every row votes with equal strength regardless of how badly it is missed, which is precisely what makes this loss insensitive to extreme targets. **A count-style target.** For an insurance pure-premium or claim-count model fitted with a Poisson deviance, the score is again on a log scale with mean `mu = exp(F)`. The pseudo-residual is `r = y - mu`: the observed count minus the predicted mean, on the count scale, while the model itself accumulates on the log scale. A policy with three claims against a predicted mean of 0.4 produces a much larger pull than one with three claims against a predicted mean of 2.8 -- and the log-scale accumulation guarantees the fitted mean can never go negative, which a squared-error model on the same data would happily do. ### Why this generality is the point Because the only thing the algorithm needs from a loss is its first derivative with respect to the prediction, gradient boosting is a *framework*, not a single model. Swap the loss and everything else -- the trees, the stagewise loop, the sum -- stays identical. That is why the same machinery serves regression, binary classification, count targets and ranking, and it is the answer to the interview question -- what makes it *gradient* boosting rather than just *residual* boosting. ### Getting the scales straight The most common confusion in this material is mixing the scale the model accumulates on with the scale the target lives on. A useful discipline: name three things for any boosted model -- what `F0` is, what scale `F` lives on, and what `-dL/dF` looks like. For log loss that is: log-odds of the base rate; the log-odds scale; `y - p`. For Poisson: log of the mean count; the log scale; `y - mu`. For squared error all three collapse onto the target scale, which is why the special case is so easy to mistake for the general rule. ### A caution about the word residual Some presentations keep calling the fitted quantity -the residual- for every loss. That is harmless shorthand once you know the derivation and dangerous before, because it hides the fact that the units change. If an interviewer asks what a boosted classifier's second tree predicts and you answer -the errors-, expect the follow-up: errors measured how, on which scale, bounded by what.

  • For a binary clinic no-show model trained with log loss, what does each tree actually output?
    An increment on the log-odds scale, not a probability. The constant start is the log-odds of the training no-show rate, every tree adds to that score, and only at the end is the accumulated score pushed through the logistic function to give a probability. That is also why the tree's target, `y - p`, is bounded between minus one and one.
  • Your target is claim counts and you keep squared-error loss -- what goes wrong?
    Squared error treats an error of three identically whether the expected count is 0.4 or 28, and it lets fitted values go negative, which is meaningless for counts. A count-appropriate loss puts the model on a log scale so the fitted mean stays positive, and makes the pseudo-residual the observed count minus the predicted mean.
  • Is the gradient taken with respect to the tree's parameters?
    No. It is taken with respect to the model's output value for each row, which is why boosting is described as gradient descent in function space. The tree is then fitted to that field of desired directions, which is how a per-row derivative becomes an update over whole regions of feature space.

The loss tells each row which way its prediction should lean; the pseudo-residual is that lean written down as a number, and the tree learns to lean whole regions at once.

saying these in an interview costs you the question

  • Says the tree always fits y minus the prediction
  • Thinks boosting only works with squared-error loss
  • Confuses the log-odds scale with the probability scale
  • Claims the gradient is taken over split thresholds
  • Cannot say what the tree predicts for a classifier

context