Evaluation Metrics
You will learn which metric fits which problem: precision against recall, PR-AUC against ROC-AUC, calibration, RMSE against MAE, plus where to cut. 'Precision or recall here?' opens most DS screens.
on this pageshowhide
explore
- Confusion Matrix Measures12 questions
- Accuracy and Its Limits4 questions
- Precision Versus Recall4 questions
- Multiclass Averaging4 questions
- Ranking and Thresholds16 questions
- ROC Curve and AUC4 questions
- Precision-Recall Curve3 questions
- Choosing a Threshold5 questions
- Lift and Gains Charts4 questions
- Probability Quality11 questions
- Log-Loss and Brier Score4 questions
- Reliability and Recalibration3 questions
- Conformal Prediction Intervals4 questions
- Regression Errors11 questions
- RMSE Versus MAE4 questions
- Out-of-Sample R-Squared3 questions
- Error Slices and Baselines4 questions
questions
page 2 of 2Your conformal calibration split was collected before a pricing change — does the 90% guarantee still hold?
basics
~20 sNo. Conformal coverage rests on exchangeability between the calibration rows and the rows you predict on. A pricing change makes post-change traffic a different population, so the stored quantile is stale and realised coverage can drift below 90%.
Is a log-loss of 0.31 good when the positive class occurs 12% of the time?
basics
~20 sOn its own the number means nothing. Compare it with the constant forecaster that predicts 0.12 on every row, which scores about 0.367. A log-loss of 0.31 is therefore roughly a 15% improvement on that reference: real, but modest.
A model trained on 1:1 undersampled data outputs 0.5 when the live positive rate is 4%. How do you fix the probabilities?
basics
~20 sBalancing the training data shifted the log-odds by a constant. Undo it: multiply the predicted odds by the ratio of true prior odds to training prior odds. A 0.5 score becomes 0.04 at a 4% base rate.
A defect model reads 94% accuracy on a test set rebalanced to 50/50, but live defects run at 0.3%. What does that number tell you?
basics
~20 sAlmost nothing about production. Accuracy depends on the class mix it was measured on, so a 50/50 test set reports balanced accuracy, not the live figure at a 0.3% defect rate. Recombine the per-class rates at the true base rate.
Why report Cohen's kappa instead of raw agreement for a 5-grade severity classifier?
basics
~10 sCohen's kappa subtracts the agreement expected by chance from the observed agreement and rescales the remainder, so a model that merely mimics the grade distribution scores near zero even when raw agreement looks high.
Why is ROC-AUC unchanged after an evaluation set is downsampled from 3% to 50% positives?
basics
~20 sTrue positive rate is computed only among positives and false positive rate only among negatives, so randomly discarding negatives leaves both rates unchanged in expectation. ROC-AUC measures ranking quality and is therefore blind to the class mix.
A promo-abuse detector's positive rate triples during a January sale. What happens at its fixed threshold?
basics
~20 sFlagged volume rises sharply. If only the base rate moved, precision at the fixed cut rises while recall holds, so the queue floods rather than degrades. If the traffic itself changed shape, false positives rise and precision can fall.
A threshold picked to hit 85% precision on validation delivers less in production. Why?
basics
~20 sPicking the cut that first clears 85% in a sweep selects the point where sampling noise flattered you, and precision at a strict cut rests on few items. The estimate is optimistic; aim above the target on untouched data.
Why can the same regression model post very different held-out R-squared on two test splits?
basics
~20 sR-squared divides the model's squared error by the target's spread on the rows being scored. Narrow that spread and the same absolute errors give a much lower score, so the number is not comparable across splits.
Your model is tuned to minimise MAE and its predicted totals fall short of actual totals on a right-skewed target. Why?
basics
~10 sMinimising absolute error drives predictions toward the conditional median, and on a right-skewed target the median sits below the mean. Median-like predictions therefore sum low, even though each individual row looks accurate.
What does Hamming loss measure for a news tagger that assigns several topics per article?
basics
~20 sHamming loss is the fraction of individual label decisions that are wrong across all articles and all candidate topics. Each missed tag and each spurious tag counts once, lower is better, and partly-correct tag sets earn partial credit.
What do Youden's J and maximising F1 each assume when used to pick a classification threshold?
basics
~20 sYouden's J, sensitivity plus specificity minus one, weights both error rates equally and ignores class sizes. Maximising F1 ignores true negatives entirely and implies a trade-off that shifts with the operating point rather than a fixed cost ratio.
What does RMSLE measure that RMSE does not when scoring skewed positive order values?
basics
~20 sRMSLE is the root mean squared error of log(1 + value), so it scores the ratio between prediction and actual, not the currency gap. Being twice too high on a $20 order costs what it costs on a $2,000 one.
How do you choose beta in F-beta when a missed case costs far more than a false alarm?
basics
~20 sBeta sets how many times more recall matters than precision: F2 weights recall twice as heavily, F0.5 half as heavily, F1 equally. Derive beta from the cost ratio of a missed case to a false alarm.
A telemarketing model shows 4x lift in decile 1 — why might the campaign not lift response 4x?
basics
~20 sLift measures who is likely to respond, not who responds because you called. Much of the top decile would have bought anyway, so targeting them redistributes credit rather than creating sales. Measuring the campaign's real effect needs a randomised control group.
A precision-recall curve built on only 60 positives is jagged. How much do you trust its average precision?
basics
~20 sNot much as a point value. The curve's resolution comes from the positive count, not the row count: with 60 positives each is a 1.7% recall step, so average precision carries sampling variance. Report a resampled interval, not three decimals.
When reporting a regression model's worst error slice, how do you know it is not noise?
basics
~20 sA maximum over many slices is biased: cut finely enough and one always looks terrible by chance. Fix the slice list and minimum size in advance, attach a resampling interval to each slice's error, and require it to repeat on fresh data.
How do you set the coverage level for a skin-lesion classifier's conformal prediction sets in a triage workflow?
basics
~20 sChoose the miscoverage level from the cost of missing the true diagnosis, not from convention. Higher coverage means larger label sets and more escalations, and the workflow must handle sets of any size, including two labels or none.
How would you decide whether log-loss or the Brier score is your team's headline probability metric?
basics
~20 sChoose by the decision the probabilities feed. Log-loss when an overconfident mistake is expensive and you want extremes punished hard; the Brier score when you need a bounded, stakeholder-legible number that no single row can dominate.
Two candidate models' ROC curves cross - how do you decide which one to ship?
basics
~20 sCrossing curves mean neither model dominates - each wins over a different range of false positive rates. Compare them where you will actually operate, such as true positive rate at a fixed low false positive rate, not by total area.
showing 31–50 of 50