skip to content

Evaluation Metrics

You will learn which metric fits which problem: precision against recall, PR-AUC against ROC-AUC, calibration, RMSE against MAE, plus where to cut. 'Precision or recall here?' opens most DS screens.

on this pageshow

explore

questions

page 1 of 2

How do you check whether a classifier's predicted probabilities are calibrated?

level: juniorimportance: must knowfreq 62%

answer

  1. bucket the scores, then compare
  2. predicted rate versus observed rate
  3. distance from the 45-degree line
  4. one weighted average of the gaps
  5. ranking metrics are blind to it

basics

~20 s

Bucket the predictions by score, then compare each bucket's mean predicted probability with the fraction of rows in it that were actually positive. Plotted, that is a reliability diagram: a calibrated model sits on the diagonal.

solid answer

~40 s

I sort the held-out predictions into bins - either equal-width bins on the score, or equal-frequency bins so each holds the same number of rows - and for each bin I compute two numbers: the mean predicted score and the observed fraction of positives. Plotting observed against predicted gives a reliability diagram, and a perfectly calibrated model lies on the 45-degree line. Points below the line mean the model is overconfident there, points above mean it is underconfident. I summarise the picture with expected calibration error, the bin-count-weighted average of the absolute gaps, but I always keep the diagram and the per-bin counts, because a sparse high-score bin is noise. Crucially, ranking metrics say nothing about this: any monotone rescaling of the scores leaves AUC untouched while wrecking the numbers.

go deeper

for a junior

Be ready to state the definition in one line - among rows scored 0.3, about 30% should be positive - and to describe the bin-and-plot procedure. Knowing that a good AUC does not imply good probabilities already puts you ahead.

for a middle

Explain the mechanics: how bins are formed, what each axis of the reliability diagram holds, the expected calibration error formula, and why equal-width binning misbehaves when the positive rate is low.

for a senior

Show you check calibration on held-out data at production prevalence, read the shape of the curve as a diagnosis rather than a pass or fail, and inspect the score region where decisions are actually taken.

for a principal

Own the question of whether calibration is worth buying at all. Argue when a ranked list is sufficient and when downstream money calculations make the absolute number a hard requirement, and set what the team reports alongside AUC.

## What calibration actually claims A binary classifier gives each row a score between 0 and 1. The model is **calibrated** if the score can be read as a probability: among all the rows scored near 0.30, about 30% really turn out positive; among the rows scored near 0.70, about 70% do. Notice what is *not* being claimed. Nothing about accuracy, nothing about which rows are ranked above which. Calibration is a property of the *numbers*; discrimination is a property of the *ordering*. A model can be excellent at one and hopeless at the other. This matters the moment the score is consumed as a number rather than as a rank. A subscription-churn model whose 0.30 bucket must really churn 30% of the time before finance sizes a retention offer is the standard case: if the true churn rate in that bucket is 12%, every expected-value calculation built on the score - offer budget, expected saved revenue, expected number of churners next month - is wrong by more than a factor of two, even though the model may rank customers beautifully. ## The reliability diagram The diagnostic is mechanical: 1. Take predictions on data the model was **not** trained on. 2. Split the score range into bins. *Equal-width* bins cut [0, 1] into, say, 10 slices of 0.1. *Equal-frequency* bins put the same number of rows in each. 3. In every bin compute the mean predicted score (the x coordinate) and the observed positive rate (the y coordinate). 4. Plot the points, add the 45-degree diagonal, and plot a histogram of bin counts underneath. Reading it: - **On the diagonal** - calibrated in that region. - **Below the diagonal** - observed rate is lower than predicted; the model is **overconfident** there, promising more positives than materialise. - **Above the diagonal** - observed rate exceeds predicted; the model is **underconfident**, understating the risk. - **A curve steeper than the diagonal, crossing it in the middle** - the classic signature of scores squeezed toward 0.5. A model that averages many weak votes, such as a forest of trees, often never emits a score below 0.2 or above 0.8, so at the low end it overshoots and at the high end it undershoots. - **A flatter S, hugging the extremes** - scores pushed toward 0 and 1; predictions of 0.97 that come true only 70% of the time. ## Expected calibration error The diagram condensed to one number: `ECE = sum over bins b of (n_b / N) * | mean_predicted_b - observed_rate_b |` That is the average vertical gap in the diagram, weighted by how many rows fall in each bin. Zero is perfect. It is useful and it is fragile, for reasons worth knowing: - **It depends on the binning.** The same model scored over 10 equal-width bins and over 50 bins gives different numbers, usually a larger one with more bins, because fine bins stop opposite-signed errors from cancelling inside a wide bin. Only compare ECE across models computed with identical bins on identical data, and report the bin count alongside the number. - **Equal-width binning collapses under a low base rate.** On an insurance-claim propensity score where almost every row lands under 0.1, nine of ten equal-width bins may hold a handful of rows each; their observed rates are pure noise, and the weighting mercifully shrinks their influence but does not make the picture readable. Equal-frequency bins fix the counts at the cost of uneven widths. - **It throws away direction.** The absolute value means a model that is overconfident at the top and underconfident at the bottom can score the same as one uniformly off in one direction, though the fixes differ. Keep the diagram. - **It is an estimate.** With a few thousand rows the per-bin observed rates carry real sampling error; do not read a difference of 0.002 as an improvement. ## Why AUC cannot answer this question ROC AUC is the probability that a random positive scores above a random negative. Apply any strictly increasing transform to every score - halve them, square them, push them through a sigmoid - and every pairwise comparison is preserved, so AUC is *exactly* unchanged while the probabilities become arbitrary. This is why a churn model can sit at AUC 0.82 and still be unusable for pricing, and equally why recalibrating a model is a safe operation for its ranking: a monotone reshaping cannot damage the ordering. ## Practical hygiene Always evaluate calibration on held-out data, at the class prevalence the model will actually meet in production. Show the counts per bin. Look at the region you actually act on - if decisions are made above 0.8, the calibration of the 0.1 bucket is a curiosity. And treat a bad diagram as diagnosis, not verdict: the shape tells you which recalibration method is likely to fix it.

  • Your expected calibration error is computed over 10 equal-width bins. How much should that single number be trusted?
    Only as a comparison against other models measured with the identical binning. Change to 50 bins and the number usually rises, because narrow bins stop opposite errors cancelling. With a low base rate, equal-width bins leave the high-score bins nearly empty and their observed rates are noise. Report the bin count and the per-bin counts, and read the diagram rather than the scalar.
  • A churn model has AUC 0.82 but its 0.30 bucket churns at only 12%. Is the model useless?
    No. AUC 0.82 says the ordering carries real signal, so a top-N retention list would work. What is broken is the scale, and the scale is exactly what finance needs to size an offer. A monotone recalibration fitted on held-out data can straighten the numbers without disturbing the ranking that already works.
  • The reliability curve sits above the diagonal across the whole score range. What is the model doing?
    Understating risk everywhere: the observed positive rate exceeds the predicted score in every bin, so it is systematically underconfident. Common causes are a training set whose positive rate is lower than the evaluation set's, heavy regularisation shrinking the scores toward the middle, or averaging over many weak learners. The fix is a monotone recalibration, or checking that the two prevalences actually match.

A bathroom scale that reads two kilos heavy still ranks a family by weight perfectly. It is simply wrong about every number it prints, and you only notice when someone tries to use the number for something.

saying these in an interview costs you the question

  • Claims a high AUC proves the probabilities are trustworthy
  • Reads overall accuracy as evidence of calibration
  • Ignores how many rows fall in each bin
  • Treats a score of 0.9 as 90% confidence without ever checking
  • Measures calibration on the same data the model was trained on
  • Quotes one expected calibration error without saying how many bins

context

open as a page

In a binary classifier's confusion matrix, what are the four cells and how is accuracy computed?

level: juniorimportance: must knowfreq 84%

basics

~20 s

The four cells count true positives, false positives, true negatives and false negatives - one per pairing of predicted and actual label. Accuracy is (TP + TN) divided by all four counts; the error rate is 1 minus accuracy.

open as a page

Why is 99.7% accuracy uninformative for a defect detector when 0.3% of units are defective?

level: juniorimportance: must knowfreq 88%

basics

~20 s

Because a model that never predicts 'defective' already scores 99.7% on that data. The majority class supplies almost every row, so accuracy measures performance on good units and says nothing about whether a single defect was ever caught.

open as a page

How do you read per-class precision and recall off a 10-class confusion matrix?

level: juniorimportance: must knowfreq 62%

basics

~20 s

With true classes as rows and predicted classes as columns, a class's recall is its diagonal cell divided by its row total, and its precision is that same diagonal cell divided by its column total.

open as a page

What is the difference between precision and recall, and why do they trade off?

level: juniorimportance: must knowfreq 88%

basics

~20 s

Precision is the share of predicted positives that are truly positive; recall is the share of actual positives the model catches. Loosening the decision rule labels more examples positive, which raises recall and usually lowers precision.

open as a page

What does a cumulative gains chart show, and how does a lift chart differ?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A cumulative gains chart plots the share of all responders captured against the share of the file contacted, after sorting by score. A lift chart shows the same result as a multiple of random targeting, where random equals 1.

open as a page

What does a precision-recall curve plot, and what does average precision measure?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A precision-recall curve plots precision against recall as the decision threshold sweeps from strict to permissive. Average precision compresses that curve into one number: the recall-weighted mean of the precisions along it, estimating the area beneath.

open as a page

What do the two axes of a binary classifier's ROC curve show, and what traces the curve?

level: juniorimportance: must knowfreq 88%

basics

~20 s

An ROC curve plots true positive rate on the y-axis against false positive rate on the x-axis. Each point is one decision threshold, and sweeping the threshold from strictest to loosest traces the curve from (0,0) to (1,1).

open as a page

Lowering a trained classifier's decision threshold from 0.5 to 0.3 changes what, and what stays fixed?

level: juniorimportance: must knowfreq 78%

basics

~20 s

More items are labelled positive, so true positives and false positives can only rise, recall can only rise, and precision usually falls. The fitted model, its scores and its ranking of items do not change at all.

open as a page

Before shipping a regression model, what baseline must it beat and why?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Two: the best constant prediction — the target's median for MAE, its mean for squared error — and the rule of thumb the business already runs. Both on the same held-out rows, they turn a bare error number into a decision.

open as a page

Why report a regression model's R-squared on held-out rows instead of the rows it was fitted on?

level: juniorimportance: must knowfreq 70%

basics

~20 s

R-squared on the fitting rows is measured on data whose residuals the fit already minimised, so it flatters the model. Held-out R-squared scores rows the model never saw, the only honest estimate of future performance.

open as a page

Why does RMSE react far more strongly than MAE to a single large prediction error?

level: juniorimportance: must knowfreq 85%

basics

~20 s

RMSE squares every error before averaging, so one huge miss contributes out of all proportion. MAE averages the absolute errors, giving each mistake weight in line with its size. RMSE tracks the worst errors, MAE the typical one.

open as a page

How does split conformal prediction turn a trained regression model into 90% prediction intervals?

level: middleimportance: must knowfreq 40%

basics

~20 s

Split conformal holds out a calibration set the model never trained on, scores every row by its absolute residual, and takes the residual sitting at the 90% rank. Each new prediction then gets that one value added and subtracted.

open as a page

How do log-loss and the Brier score differ in punishing a confident wrong probability?

level: middleimportance: must knowfreq 68%

basics

~20 s

Log-loss charges minus the log of the probability the model gave to the outcome that actually happened, so a confident miss costs unboundedly much. The Brier score squares the probability error, so no single row can cost more than 1.

open as a page

How do macro, micro and weighted averaging of multiclass F1 differ?

level: middleimportance: must knowfreq 78%

basics

~20 s

Macro averages per-class F1 scores with equal weight, so a rare class counts as much as a common one. Micro pools every class's true positives and errors first, so frequent classes dominate. Weighted averages per-class scores by class support.

open as a page

A fraud model shows ROC-AUC 0.97 but average precision 0.15. Why the gap?

level: middleimportance: must knowfreq 66%

basics

~20 s

Both are right; they use different denominators. With 0.2% fraud, tens of thousands of false positives barely move the false positive rate, but they dominate the flagged set and crush precision. Read average precision against a 0.002 baseline, not 0.5.

open as a page

What does an ROC-AUC of 0.78 mean in probability terms for a display-ad click model?

level: middleimportance: must knowfreq 76%

basics

~20 s

Draw one clicker and one non-clicker at random: the model scores the clicker higher 78% of the time, counting ties as half a win. ROC-AUC is a statement about ranking order, not about the numeric size of the scores.

open as a page

How do you turn a cost matrix over false positives and false negatives into a decision threshold?

level: middleimportance: must knowfreq 62%

basics

~20 s

Flag an item when the expected cost of flagging is below the expected cost of not flagging. With false-positive and false-negative costs only, the break-even is cost_FP / (cost_FP + cost_FN), which is 0.5 only for equal costs.

open as a page

Why can a regression model's good overall MAE hide a cohort where it is badly broken?

level: middleimportance: must knowfreq 68%

basics

~20 s

Overall MAE is a row-weighted average, so a small cohort can run three times worse and barely move it: a cohort holding 3% of rows at triple the error lifts overall MAE by only about 6%. Cut the metric by cohort.

open as a page

What does a negative R-squared on a held-out test set tell you about the model?

level: middleimportance: must knowfreq 62%

basics

~10 s

A negative held-out R-squared means the model's squared error on those rows exceeds what predicting one number, the split's mean, for every row would give. The model lost to a constant.

open as a page

What does it mean for a probability scoring rule like the Brier score to be proper?

level: middleimportance: should knowfreq 41%

basics

~20 s

A scoring rule is proper when a forecaster gets their best expected score by reporting the probability they genuinely believe, rather than by shading it. Log-loss and the Brier score are strictly proper; thresholded accuracy is not.

open as a page

When would you choose isotonic regression over Platt scaling to recalibrate a model's scores?

level: middleimportance: should knowfreq 52%

basics

~20 s

Choose isotonic regression when the calibration set is large, a few thousand rows or more, and the distortion is not sigmoid-shaped. Platt scaling fits only a two-parameter logistic curve, so it is the safer choice on small calibration sets.

open as a page

How does balanced accuracy differ from plain accuracy on an imbalanced test set?

level: middleimportance: should knowfreq 60%

basics

~20 s

Balanced accuracy averages the per-class recalls, giving every class equal weight, while plain accuracy weights each class by how many rows it contributes. On skewed data the two diverge sharply, and a constant predictor scores 1/k instead of the majority share.

open as a page

Why is F1 the harmonic mean of precision and recall rather than the arithmetic mean?

level: middleimportance: should knowfreq 62%

basics

~20 s

The harmonic mean is dominated by the smaller of the two values, so F1 stays low unless precision and recall are both decent. Precision 0.90 with recall 0.10 gives F1 0.18, while the arithmetic mean would flatter it at 0.50.

open as a page

Why can a sepsis alert with 95% specificity still be wrong most times it fires?

level: middleimportance: should knowfreq 48%

basics

~20 s

Specificity is measured only among patients who do not have sepsis, so it ignores how rare sepsis is. At a 2% rate, 5% of the huge non-sepsis group yields far more false alarms than true cases, leaving precision near 25%.

open as a page

In credit scoring, what do a KS of 38 and a Gini of 0.52 mean?

level: middleimportance: should knowfreq 48%

basics

~20 s

KS 38 means that at its best cut-off the scorecard separates 38 percentage points more of the bads than of the goods. Gini 0.52 is a whole-curve summary equal to two times AUC minus one, so AUC is 0.76.

open as a page

In a decile lift table, what does non-monotonic lift across deciles tell you?

level: middleimportance: should knowfreq 40%

basics

~20 s

It means the score stops rank-ordering cleanly in that region. Usually it is sampling noise in small deciles; sometimes it is a genuinely unstable model, a shifted population, or a scoring bug. Check decile sizes before you diagnose anything else.

open as a page

In a residual-versus-fitted plot, what does a funnel widening with the predicted value mean?

level: middleimportance: should knowfreq 48%

basics

~20 s

Residual spread grows with the size of the prediction, so error is roughly proportional to the quantity rather than constant. Any absolute-error headline is then mostly a statement about the largest predictions. Report error banded by predicted magnitude.

open as a page

Why is MAPE a poor scoring metric for a call-centre wait-time model where some waits are under a minute?

level: middleimportance: should knowfreq 58%

basics

~10 s

MAPE divides each error by the actual value, so a 30-second miss on a one-minute wait scores 50% and near-zero waits explode. It is also asymmetric: over-prediction is unbounded, under-prediction capped at 100%.

open as a page

Your conformal intervals hit 90% coverage overall but only 74% on the oldest patient cohort — why?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Conformal prediction guarantees marginal coverage: 90% averaged over the whole population, not inside every subgroup. One global quantile under-covers groups whose errors are larger and over-covers the easy ones. Group-wise calibration or an input-adaptive score is the fix.

open as a page

showing 1–30 of 50