skip to content

Probability Quality

Whether a predicted 0.7 really happens about 70% of the time: proper scoring rules, reliability diagrams, and the fixes that reshape scores. Ranking well and being calibrated are not the same.

on this pageshow

explore

questions

11

How do you check whether a classifier's predicted probabilities are calibrated?

level: juniorimportance: must knowfreq 62%

answer

  1. bucket the scores, then compare
  2. predicted rate versus observed rate
  3. distance from the 45-degree line
  4. one weighted average of the gaps
  5. ranking metrics are blind to it

basics

~20 s

Bucket the predictions by score, then compare each bucket's mean predicted probability with the fraction of rows in it that were actually positive. Plotted, that is a reliability diagram: a calibrated model sits on the diagonal.

solid answer

~40 s

I sort the held-out predictions into bins - either equal-width bins on the score, or equal-frequency bins so each holds the same number of rows - and for each bin I compute two numbers: the mean predicted score and the observed fraction of positives. Plotting observed against predicted gives a reliability diagram, and a perfectly calibrated model lies on the 45-degree line. Points below the line mean the model is overconfident there, points above mean it is underconfident. I summarise the picture with expected calibration error, the bin-count-weighted average of the absolute gaps, but I always keep the diagram and the per-bin counts, because a sparse high-score bin is noise. Crucially, ranking metrics say nothing about this: any monotone rescaling of the scores leaves AUC untouched while wrecking the numbers.

go deeper

for a junior

Be ready to state the definition in one line - among rows scored 0.3, about 30% should be positive - and to describe the bin-and-plot procedure. Knowing that a good AUC does not imply good probabilities already puts you ahead.

for a middle

Explain the mechanics: how bins are formed, what each axis of the reliability diagram holds, the expected calibration error formula, and why equal-width binning misbehaves when the positive rate is low.

for a senior

Show you check calibration on held-out data at production prevalence, read the shape of the curve as a diagnosis rather than a pass or fail, and inspect the score region where decisions are actually taken.

for a principal

Own the question of whether calibration is worth buying at all. Argue when a ranked list is sufficient and when downstream money calculations make the absolute number a hard requirement, and set what the team reports alongside AUC.

## What calibration actually claims A binary classifier gives each row a score between 0 and 1. The model is **calibrated** if the score can be read as a probability: among all the rows scored near 0.30, about 30% really turn out positive; among the rows scored near 0.70, about 70% do. Notice what is *not* being claimed. Nothing about accuracy, nothing about which rows are ranked above which. Calibration is a property of the *numbers*; discrimination is a property of the *ordering*. A model can be excellent at one and hopeless at the other. This matters the moment the score is consumed as a number rather than as a rank. A subscription-churn model whose 0.30 bucket must really churn 30% of the time before finance sizes a retention offer is the standard case: if the true churn rate in that bucket is 12%, every expected-value calculation built on the score - offer budget, expected saved revenue, expected number of churners next month - is wrong by more than a factor of two, even though the model may rank customers beautifully. ## The reliability diagram The diagnostic is mechanical: 1. Take predictions on data the model was **not** trained on. 2. Split the score range into bins. *Equal-width* bins cut [0, 1] into, say, 10 slices of 0.1. *Equal-frequency* bins put the same number of rows in each. 3. In every bin compute the mean predicted score (the x coordinate) and the observed positive rate (the y coordinate). 4. Plot the points, add the 45-degree diagonal, and plot a histogram of bin counts underneath. Reading it: - **On the diagonal** - calibrated in that region. - **Below the diagonal** - observed rate is lower than predicted; the model is **overconfident** there, promising more positives than materialise. - **Above the diagonal** - observed rate exceeds predicted; the model is **underconfident**, understating the risk. - **A curve steeper than the diagonal, crossing it in the middle** - the classic signature of scores squeezed toward 0.5. A model that averages many weak votes, such as a forest of trees, often never emits a score below 0.2 or above 0.8, so at the low end it overshoots and at the high end it undershoots. - **A flatter S, hugging the extremes** - scores pushed toward 0 and 1; predictions of 0.97 that come true only 70% of the time. ## Expected calibration error The diagram condensed to one number: `ECE = sum over bins b of (n_b / N) * | mean_predicted_b - observed_rate_b |` That is the average vertical gap in the diagram, weighted by how many rows fall in each bin. Zero is perfect. It is useful and it is fragile, for reasons worth knowing: - **It depends on the binning.** The same model scored over 10 equal-width bins and over 50 bins gives different numbers, usually a larger one with more bins, because fine bins stop opposite-signed errors from cancelling inside a wide bin. Only compare ECE across models computed with identical bins on identical data, and report the bin count alongside the number. - **Equal-width binning collapses under a low base rate.** On an insurance-claim propensity score where almost every row lands under 0.1, nine of ten equal-width bins may hold a handful of rows each; their observed rates are pure noise, and the weighting mercifully shrinks their influence but does not make the picture readable. Equal-frequency bins fix the counts at the cost of uneven widths. - **It throws away direction.** The absolute value means a model that is overconfident at the top and underconfident at the bottom can score the same as one uniformly off in one direction, though the fixes differ. Keep the diagram. - **It is an estimate.** With a few thousand rows the per-bin observed rates carry real sampling error; do not read a difference of 0.002 as an improvement. ## Why AUC cannot answer this question ROC AUC is the probability that a random positive scores above a random negative. Apply any strictly increasing transform to every score - halve them, square them, push them through a sigmoid - and every pairwise comparison is preserved, so AUC is *exactly* unchanged while the probabilities become arbitrary. This is why a churn model can sit at AUC 0.82 and still be unusable for pricing, and equally why recalibrating a model is a safe operation for its ranking: a monotone reshaping cannot damage the ordering. ## Practical hygiene Always evaluate calibration on held-out data, at the class prevalence the model will actually meet in production. Show the counts per bin. Look at the region you actually act on - if decisions are made above 0.8, the calibration of the 0.1 bucket is a curiosity. And treat a bad diagram as diagnosis, not verdict: the shape tells you which recalibration method is likely to fix it.

  • Your expected calibration error is computed over 10 equal-width bins. How much should that single number be trusted?
    Only as a comparison against other models measured with the identical binning. Change to 50 bins and the number usually rises, because narrow bins stop opposite errors cancelling. With a low base rate, equal-width bins leave the high-score bins nearly empty and their observed rates are noise. Report the bin count and the per-bin counts, and read the diagram rather than the scalar.
  • A churn model has AUC 0.82 but its 0.30 bucket churns at only 12%. Is the model useless?
    No. AUC 0.82 says the ordering carries real signal, so a top-N retention list would work. What is broken is the scale, and the scale is exactly what finance needs to size an offer. A monotone recalibration fitted on held-out data can straighten the numbers without disturbing the ranking that already works.
  • The reliability curve sits above the diagonal across the whole score range. What is the model doing?
    Understating risk everywhere: the observed positive rate exceeds the predicted score in every bin, so it is systematically underconfident. Common causes are a training set whose positive rate is lower than the evaluation set's, heavy regularisation shrinking the scores toward the middle, or averaging over many weak learners. The fix is a monotone recalibration, or checking that the two prevalences actually match.

A bathroom scale that reads two kilos heavy still ranks a family by weight perfectly. It is simply wrong about every number it prints, and you only notice when someone tries to use the number for something.

saying these in an interview costs you the question

  • Claims a high AUC proves the probabilities are trustworthy
  • Reads overall accuracy as evidence of calibration
  • Ignores how many rows fall in each bin
  • Treats a score of 0.9 as 90% confidence without ever checking
  • Measures calibration on the same data the model was trained on
  • Quotes one expected calibration error without saying how many bins

context

open as a page

How does split conformal prediction turn a trained regression model into 90% prediction intervals?

level: middleimportance: must knowfreq 40%

basics

~20 s

Split conformal holds out a calibration set the model never trained on, scores every row by its absolute residual, and takes the residual sitting at the 90% rank. Each new prediction then gets that one value added and subtracted.

open as a page

How do log-loss and the Brier score differ in punishing a confident wrong probability?

level: middleimportance: must knowfreq 68%

basics

~20 s

Log-loss charges minus the log of the probability the model gave to the outcome that actually happened, so a confident miss costs unboundedly much. The Brier score squares the probability error, so no single row can cost more than 1.

open as a page

What does it mean for a probability scoring rule like the Brier score to be proper?

level: middleimportance: should knowfreq 41%

basics

~20 s

A scoring rule is proper when a forecaster gets their best expected score by reporting the probability they genuinely believe, rather than by shading it. Log-loss and the Brier score are strictly proper; thresholded accuracy is not.

open as a page

When would you choose isotonic regression over Platt scaling to recalibrate a model's scores?

level: middleimportance: should knowfreq 52%

basics

~20 s

Choose isotonic regression when the calibration set is large, a few thousand rows or more, and the distortion is not sigmoid-shaped. Platt scaling fits only a two-parameter logistic curve, so it is the safer choice on small calibration sets.

open as a page

Your conformal intervals hit 90% coverage overall but only 74% on the oldest patient cohort — why?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Conformal prediction guarantees marginal coverage: 90% averaged over the whole population, not inside every subgroup. One global quantile under-covers groups whose errors are larger and over-covers the easy ones. Group-wise calibration or an input-adaptive score is the fix.

open as a page

Your conformal calibration split was collected before a pricing change — does the 90% guarantee still hold?

level: seniorimportance: should knowfreq 26%

basics

~20 s

No. Conformal coverage rests on exchangeability between the calibration rows and the rows you predict on. A pricing change makes post-change traffic a different population, so the stored quantile is stale and realised coverage can drift below 90%.

open as a page

Is a log-loss of 0.31 good when the positive class occurs 12% of the time?

level: seniorimportance: should knowfreq 47%

basics

~20 s

On its own the number means nothing. Compare it with the constant forecaster that predicts 0.12 on every row, which scores about 0.367. A log-loss of 0.31 is therefore roughly a 15% improvement on that reference: real, but modest.

open as a page

A model trained on 1:1 undersampled data outputs 0.5 when the live positive rate is 4%. How do you fix the probabilities?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Balancing the training data shifted the log-odds by a constant. Undo it: multiply the predicted odds by the ratio of true prior odds to training prior odds. A 0.5 score becomes 0.04 at a 4% base rate.

open as a page

How do you set the coverage level for a skin-lesion classifier's conformal prediction sets in a triage workflow?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

Choose the miscoverage level from the cost of missing the true diagnosis, not from convention. Higher coverage means larger label sets and more escalations, and the workflow must handle sets of any size, including two labels or none.

open as a page

How would you decide whether log-loss or the Brier score is your team's headline probability metric?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

Choose by the decision the probabilities feed. Log-loss when an overconfident mistake is expensive and you want extremes punished hard; the Brier score when you need a bounded, stakeholder-legible number that no single row can dominate.

open as a page