How do you check whether a classifier's predicted probabilities are calibrated?
answer
- bucket the scores, then compare
- predicted rate versus observed rate
- distance from the 45-degree line
- one weighted average of the gaps
- ranking metrics are blind to it
basics
~20 sBucket the predictions by score, then compare each bucket's mean predicted probability with the fraction of rows in it that were actually positive. Plotted, that is a reliability diagram: a calibrated model sits on the diagonal.
solid answer
~40 sI sort the held-out predictions into bins - either equal-width bins on the score, or equal-frequency bins so each holds the same number of rows - and for each bin I compute two numbers: the mean predicted score and the observed fraction of positives. Plotting observed against predicted gives a reliability diagram, and a perfectly calibrated model lies on the 45-degree line. Points below the line mean the model is overconfident there, points above mean it is underconfident. I summarise the picture with expected calibration error, the bin-count-weighted average of the absolute gaps, but I always keep the diagram and the per-bin counts, because a sparse high-score bin is noise. Crucially, ranking metrics say nothing about this: any monotone rescaling of the scores leaves AUC untouched while wrecking the numbers.
go deeper
Be ready to state the definition in one line - among rows scored 0.3, about 30% should be positive - and to describe the bin-and-plot procedure. Knowing that a good AUC does not imply good probabilities already puts you ahead.
Explain the mechanics: how bins are formed, what each axis of the reliability diagram holds, the expected calibration error formula, and why equal-width binning misbehaves when the positive rate is low.
Show you check calibration on held-out data at production prevalence, read the shape of the curve as a diagnosis rather than a pass or fail, and inspect the score region where decisions are actually taken.
Own the question of whether calibration is worth buying at all. Argue when a ranked list is sufficient and when downstream money calculations make the absolute number a hard requirement, and set what the team reports alongside AUC.
## What calibration actually claims A binary classifier gives each row a score between 0 and 1. The model is **calibrated** if the score can be read as a probability: among all the rows scored near 0.30, about 30% really turn out positive; among the rows scored near 0.70, about 70% do. Notice what is *not* being claimed. Nothing about accuracy, nothing about which rows are ranked above which. Calibration is a property of the *numbers*; discrimination is a property of the *ordering*. A model can be excellent at one and hopeless at the other. This matters the moment the score is consumed as a number rather than as a rank. A subscription-churn model whose 0.30 bucket must really churn 30% of the time before finance sizes a retention offer is the standard case: if the true churn rate in that bucket is 12%, every expected-value calculation built on the score - offer budget, expected saved revenue, expected number of churners next month - is wrong by more than a factor of two, even though the model may rank customers beautifully. ## The reliability diagram The diagnostic is mechanical: 1. Take predictions on data the model was **not** trained on. 2. Split the score range into bins. *Equal-width* bins cut [0, 1] into, say, 10 slices of 0.1. *Equal-frequency* bins put the same number of rows in each. 3. In every bin compute the mean predicted score (the x coordinate) and the observed positive rate (the y coordinate). 4. Plot the points, add the 45-degree diagonal, and plot a histogram of bin counts underneath. Reading it: - **On the diagonal** - calibrated in that region. - **Below the diagonal** - observed rate is lower than predicted; the model is **overconfident** there, promising more positives than materialise. - **Above the diagonal** - observed rate exceeds predicted; the model is **underconfident**, understating the risk. - **A curve steeper than the diagonal, crossing it in the middle** - the classic signature of scores squeezed toward 0.5. A model that averages many weak votes, such as a forest of trees, often never emits a score below 0.2 or above 0.8, so at the low end it overshoots and at the high end it undershoots. - **A flatter S, hugging the extremes** - scores pushed toward 0 and 1; predictions of 0.97 that come true only 70% of the time. ## Expected calibration error The diagram condensed to one number: `ECE = sum over bins b of (n_b / N) * | mean_predicted_b - observed_rate_b |` That is the average vertical gap in the diagram, weighted by how many rows fall in each bin. Zero is perfect. It is useful and it is fragile, for reasons worth knowing: - **It depends on the binning.** The same model scored over 10 equal-width bins and over 50 bins gives different numbers, usually a larger one with more bins, because fine bins stop opposite-signed errors from cancelling inside a wide bin. Only compare ECE across models computed with identical bins on identical data, and report the bin count alongside the number. - **Equal-width binning collapses under a low base rate.** On an insurance-claim propensity score where almost every row lands under 0.1, nine of ten equal-width bins may hold a handful of rows each; their observed rates are pure noise, and the weighting mercifully shrinks their influence but does not make the picture readable. Equal-frequency bins fix the counts at the cost of uneven widths. - **It throws away direction.** The absolute value means a model that is overconfident at the top and underconfident at the bottom can score the same as one uniformly off in one direction, though the fixes differ. Keep the diagram. - **It is an estimate.** With a few thousand rows the per-bin observed rates carry real sampling error; do not read a difference of 0.002 as an improvement. ## Why AUC cannot answer this question ROC AUC is the probability that a random positive scores above a random negative. Apply any strictly increasing transform to every score - halve them, square them, push them through a sigmoid - and every pairwise comparison is preserved, so AUC is *exactly* unchanged while the probabilities become arbitrary. This is why a churn model can sit at AUC 0.82 and still be unusable for pricing, and equally why recalibrating a model is a safe operation for its ranking: a monotone reshaping cannot damage the ordering. ## Practical hygiene Always evaluate calibration on held-out data, at the class prevalence the model will actually meet in production. Show the counts per bin. Look at the region you actually act on - if decisions are made above 0.8, the calibration of the 0.1 bucket is a curiosity. And treat a bad diagram as diagnosis, not verdict: the shape tells you which recalibration method is likely to fix it.
- Your expected calibration error is computed over 10 equal-width bins. How much should that single number be trusted?Only as a comparison against other models measured with the identical binning. Change to 50 bins and the number usually rises, because narrow bins stop opposite errors cancelling. With a low base rate, equal-width bins leave the high-score bins nearly empty and their observed rates are noise. Report the bin count and the per-bin counts, and read the diagram rather than the scalar.
- A churn model has AUC 0.82 but its 0.30 bucket churns at only 12%. Is the model useless?No. AUC 0.82 says the ordering carries real signal, so a top-N retention list would work. What is broken is the scale, and the scale is exactly what finance needs to size an offer. A monotone recalibration fitted on held-out data can straighten the numbers without disturbing the ranking that already works.
- The reliability curve sits above the diagonal across the whole score range. What is the model doing?Understating risk everywhere: the observed positive rate exceeds the predicted score in every bin, so it is systematically underconfident. Common causes are a training set whose positive rate is lower than the evaluation set's, heavy regularisation shrinking the scores toward the middle, or averaging over many weak learners. The fix is a monotone recalibration, or checking that the two prevalences actually match.
A bathroom scale that reads two kilos heavy still ranks a family by weight perfectly. It is simply wrong about every number it prints, and you only notice when someone tries to use the number for something.
saying these in an interview costs you the question
- Claims a high AUC proves the probabilities are trustworthy
- Reads overall accuracy as evidence of calibration
- Ignores how many rows fall in each bin
- Treats a score of 0.9 as 90% confidence without ever checking
- Measures calibration on the same data the model was trained on
- Quotes one expected calibration error without saying how many bins