skip to content

When would you choose isotonic regression over Platt scaling to recalibrate a model's scores?

level: middleimportance: should knowfreq 52%

answer

  1. one is parametric, the other is not
  2. two parameters versus a step function
  3. a sigmoid can only make sigmoids
  4. flexibility costs held-out rows
  5. steps create ties and cannot extrapolate

basics

~20 s

Choose isotonic regression when the calibration set is large, a few thousand rows or more, and the distortion is not sigmoid-shaped. Platt scaling fits only a two-parameter logistic curve, so it is the safer choice on small calibration sets.

solid answer

~40 s

Both learn a monotone map from the model's raw score to a probability, fitted on data the model never saw during training, and they differ in flexibility. Platt scaling fits a two-parameter logistic curve, `p = 1 / (1 + exp(a*s + b))`, by maximum likelihood - it assumes the distortion is sigmoid-shaped, has almost no variance, and works with a few hundred rows. Isotonic regression fits a non-decreasing step function with no shape assumption, so it straightens any monotone distortion, such as a random forest whose scores never leave 0.2 to 0.8. The price is variance: under about a thousand held-out rows it fits noise, and its piecewise-constant output creates ties and cannot extrapolate. So: plenty of data plus a non-sigmoid curve means isotonic; scarce data or a raw decision value means Platt.

go deeper

for a junior

Recall that both methods learn a monotone map from raw score to probability and neither retrains the model. Knowing that the mapping must be fitted on held-out data is the point most often missed at this level.

for a middle

Explain the mechanics: the two-parameter logistic fitted by maximum likelihood versus the non-decreasing step function from pooling adjacent violators, and the data-size threshold where the flexible method starts fitting noise.

for a senior

Demonstrate the pipeline discipline - a dedicated split or out-of-fold predictions, calibration measured on a further untouched slice, and awareness that isotonic ties and its clamped ends bite exactly in the high-score region where decisions are taken.

for a principal

Own the tradeoff between a smooth, extrapolatable but assumption-heavy map and a flexible, data-hungry one, and set the team's standard for how much data is reserved for calibration versus honest evaluation.

## The shared setup Recalibration takes an existing model as given and learns a second, one-dimensional function that maps its raw score `s` to a probability. The model is never retrained. Two properties follow immediately: the map is fitted on **labelled data the model did not train on**, and the map is **monotone**, so the ranking the model produced survives. The input `s` need not be a probability. An SVM decision value is a signed distance from the separating boundary and can be any real number; recalibration is precisely how such a score is turned into something a business can consume. This is the setting Platt scaling was introduced for. ## Platt scaling Fit a logistic curve with two parameters to the held-out scores and their labels: `p(s) = 1 / (1 + exp(a*s + b))` with `a` and `b` chosen by maximum likelihood on `(score, label)` pairs. Because a is typically negative, the curve rises with the score. **Strengths.** Two parameters is almost nothing to estimate, so the fit is stable on a few hundred rows. It is defined over the whole real line, so it extrapolates naturally beyond the scores it saw. It is smooth and strictly increasing, which means no ties are introduced and AUC is preserved exactly. **Weakness.** It can only produce a sigmoid. If the true distortion is not sigmoid-shaped - if the reliability curve wiggles, or is flat in the middle and steep at both ends - Platt scaling will straighten part of the curve and leave the rest bent. It is a strong assumption dressed up as a convenience. ## Isotonic regression Fit the non-decreasing function `g` that minimises the squared error between `g(s_i)` and the 0/1 labels, subject only to `g` never decreasing. The pool-adjacent-violators algorithm computes it: sort the rows by score, and wherever a later group has a lower mean label than an earlier one, merge the two groups and replace both by their pooled mean, repeating until the sequence is non-decreasing. The result is a **step function** - a set of intervals, each with one constant probability. **Strengths.** No shape assumption. It repairs any monotone distortion, which is what you need for a bagged model whose averaging compresses every score into 0.2 to 0.8: isotonic regression simply stretches the ends back out. Given enough data it is hard to beat. **Weaknesses.** - **Variance.** The flexibility that fixes odd shapes also fits noise. Below roughly a thousand held-out rows it tends to overfit, and the correct comparison is against Platt scaling, not against doing nothing. - **Ties.** Everything inside a step gets exactly the same probability. Because AUC counts tied pairs as half a point, collapsing many distinct scores into one value can nudge AUC down slightly - monotone but not *strictly* monotone. - **No extrapolation.** Outside the score range of the calibration set the function is clamped at its end values, and the extreme steps are estimated from the fewest rows, which is exactly where decisions are often taken. - **Blocky output.** A step function gives few distinct probabilities, which can look wrong to a consumer expecting a smooth score distribution. ## Where the calibration data comes from The single most common implementation error is fitting the calibrator on the same rows the model was trained on. In-sample scores are pushed toward the correct labels by the training process itself, so the reliability curve on those rows looks far better than reality; the calibrator then learns a near-identity map and corrects nothing, or worse, corrects in the wrong direction. Calibration looks flawless in the report and fails on the first live batch. The options, in order of data cost: 1. **A dedicated calibration split** - a third slice of data, separate from train and test. Simplest and cleanest when data is plentiful. 2. **Cross-validated calibration** - split the training data into k folds; for each fold, train a model on the other k-1 and predict the held-out fold. Every training row now has an *out-of-fold* prediction produced by a model that never saw it. Pool those predictions to fit one calibrator. Then refit the model on all the data and apply that calibrator to it. This spends no data and is the usual choice when data is tight. 3. **Reusing the test split** - workable in a pinch, but the test set is then no longer an honest estimate of anything, since it has been consumed by fitting. ## Choosing, in practice Plot the reliability diagram first; it names the distortion. A clean sigmoid bend, or a calibration set in the hundreds, or a raw decision value with unbounded range: Platt scaling. A few thousand held-out rows and a curve that no sigmoid will follow, such as a compressed forest score: isotonic regression. When in doubt, fit both on the same held-out folds and compare their calibration on a further untouched slice - the comparison is cheap, and the reliability diagram of each makes the winner obvious.

  • Why must the calibrator be fitted on rows the model did not train on?
    Because in-sample scores are already pulled toward their labels by training, the reliability curve on training rows looks much better than the truth. A calibrator fitted there learns a near-identity map, corrects nothing real, and can even bend the scores the wrong way - while reporting excellent calibration. The distortion you need to measure only appears on data the model has never seen.
  • You cannot spare a third split. How do you get calibration data anyway?
    Cross-validated calibration. Split the training data into k folds; for each fold train on the other k-1 and predict the held-out fold, giving every training row an out-of-fold prediction from a model that never saw it. Pool those predictions and labels, fit a single calibrator on them, then refit the model on all the data and apply that calibrator to its output.
  • Does recalibration change which rows the model ranks highest?
    Not with Platt scaling: a strictly increasing logistic map preserves every pairwise ordering, so AUC is identical to the decimal. Isotonic regression is non-decreasing but not strictly increasing - rows inside one step receive the same probability, and those new ties can shave a little off AUC. Ordering across steps is untouched, so a top-N list is essentially unchanged.

Platt scaling is a tailor who can only take a garment in or let it out. Isotonic regression re-cuts the whole pattern - far better if the shape was wrong, disastrous if you gave the tailor only three measurements.

saying these in an interview costs you the question

  • Fits the calibrator on the model's own training predictions
  • Reaches for isotonic regression on a few hundred rows
  • Believes recalibration retrains or improves the model itself
  • Thinks Platt scaling can straighten any reliability curve
  • Forgets isotonic output cannot extrapolate past the calibration range
  • Expects recalibration to raise AUC

context