skip to content

Evaluation Metrics

You will learn which metric fits which problem: precision against recall, PR-AUC against ROC-AUC, calibration, RMSE against MAE, plus where to cut. 'Precision or recall here?' opens most DS screens.

on this pageshow

explore

questions

page 2 of 2

Your conformal calibration split was collected before a pricing change — does the 90% guarantee still hold?

level: seniorimportance: should knowfreq 26%

basics

~20 s

No. Conformal coverage rests on exchangeability between the calibration rows and the rows you predict on. A pricing change makes post-change traffic a different population, so the stored quantile is stale and realised coverage can drift below 90%.

open as a page

Is a log-loss of 0.31 good when the positive class occurs 12% of the time?

level: seniorimportance: should knowfreq 47%

basics

~20 s

On its own the number means nothing. Compare it with the constant forecaster that predicts 0.12 on every row, which scores about 0.367. A log-loss of 0.31 is therefore roughly a 15% improvement on that reference: real, but modest.

open as a page

A model trained on 1:1 undersampled data outputs 0.5 when the live positive rate is 4%. How do you fix the probabilities?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Balancing the training data shifted the log-odds by a constant. Undo it: multiply the predicted odds by the ratio of true prior odds to training prior odds. A 0.5 score becomes 0.04 at a 4% base rate.

open as a page

A defect model reads 94% accuracy on a test set rebalanced to 50/50, but live defects run at 0.3%. What does that number tell you?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Almost nothing about production. Accuracy depends on the class mix it was measured on, so a 50/50 test set reports balanced accuracy, not the live figure at a 0.3% defect rate. Recombine the per-class rates at the true base rate.

open as a page

Why report Cohen's kappa instead of raw agreement for a 5-grade severity classifier?

level: seniorimportance: should knowfreq 38%

basics

~10 s

Cohen's kappa subtracts the agreement expected by chance from the observed agreement and rescales the remainder, so a model that merely mimics the grade distribution scores near zero even when raw agreement looks high.

open as a page

Why is ROC-AUC unchanged after an evaluation set is downsampled from 3% to 50% positives?

level: seniorimportance: should knowfreq 52%

basics

~20 s

True positive rate is computed only among positives and false positive rate only among negatives, so randomly discarding negatives leaves both rates unchanged in expectation. ROC-AUC measures ranking quality and is therefore blind to the class mix.

open as a page

A promo-abuse detector's positive rate triples during a January sale. What happens at its fixed threshold?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Flagged volume rises sharply. If only the base rate moved, precision at the fixed cut rises while recall holds, so the queue floods rather than degrades. If the traffic itself changed shape, false positives rise and precision can fall.

open as a page

A threshold picked to hit 85% precision on validation delivers less in production. Why?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Picking the cut that first clears 85% in a sweep selects the point where sampling noise flattered you, and precision at a strict cut rests on few items. The estimate is optimistic; aim above the target on untouched data.

open as a page

Why can the same regression model post very different held-out R-squared on two test splits?

level: seniorimportance: should knowfreq 46%

basics

~20 s

R-squared divides the model's squared error by the target's spread on the rows being scored. Narrow that spread and the same absolute errors give a much lower score, so the number is not comparable across splits.

open as a page

Your model is tuned to minimise MAE and its predicted totals fall short of actual totals on a right-skewed target. Why?

level: seniorimportance: should knowfreq 45%

basics

~10 s

Minimising absolute error drives predictions toward the conditional median, and on a right-skewed target the median sits below the mean. Median-like predictions therefore sum low, even though each individual row looks accurate.

open as a page

What does Hamming loss measure for a news tagger that assigns several topics per article?

level: middleimportance: nice to knowfreq 27%

basics

~20 s

Hamming loss is the fraction of individual label decisions that are wrong across all articles and all candidate topics. Each missed tag and each spurious tag counts once, lower is better, and partly-correct tag sets earn partial credit.

open as a page

What do Youden's J and maximising F1 each assume when used to pick a classification threshold?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Youden's J, sensitivity plus specificity minus one, weights both error rates equally and ignores class sizes. Maximising F1 ignores true negatives entirely and implies a trade-off that shifts with the operating point rather than a fixed cost ratio.

open as a page

What does RMSLE measure that RMSE does not when scoring skewed positive order values?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

RMSLE is the root mean squared error of log(1 + value), so it scores the ratio between prediction and actual, not the currency gap. Being twice too high on a $20 order costs what it costs on a $2,000 one.

open as a page

How do you choose beta in F-beta when a missed case costs far more than a false alarm?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Beta sets how many times more recall matters than precision: F2 weights recall twice as heavily, F0.5 half as heavily, F1 equally. Derive beta from the cost ratio of a missed case to a false alarm.

open as a page

A telemarketing model shows 4x lift in decile 1 — why might the campaign not lift response 4x?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Lift measures who is likely to respond, not who responds because you called. Much of the top decile would have bought anyway, so targeting them redistributes credit rather than creating sales. Measuring the campaign's real effect needs a randomised control group.

open as a page

A precision-recall curve built on only 60 positives is jagged. How much do you trust its average precision?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Not much as a point value. The curve's resolution comes from the positive count, not the row count: with 60 positives each is a 1.7% recall step, so average precision carries sampling variance. Report a resampled interval, not three decimals.

open as a page

When reporting a regression model's worst error slice, how do you know it is not noise?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

A maximum over many slices is biased: cut finely enough and one always looks terrible by chance. Fix the slice list and minimum size in advance, attach a resampling interval to each slice's error, and require it to repeat on fresh data.

open as a page

How do you set the coverage level for a skin-lesion classifier's conformal prediction sets in a triage workflow?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

Choose the miscoverage level from the cost of missing the true diagnosis, not from convention. Higher coverage means larger label sets and more escalations, and the workflow must handle sets of any size, including two labels or none.

open as a page

How would you decide whether log-loss or the Brier score is your team's headline probability metric?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

Choose by the decision the probabilities feed. Log-loss when an overconfident mistake is expensive and you want extremes punished hard; the Brier score when you need a bounded, stakeholder-legible number that no single row can dominate.

open as a page

Two candidate models' ROC curves cross - how do you decide which one to ship?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Crossing curves mean neither model dominates - each wins over a different range of false positive rates. Compare them where you will actually operate, such as true positive rate at a fixed low false positive rate, not by total area.

open as a page

showing 31–50 of 50