skip to content

Ranking and Thresholds

How a model scores before any cut-off is fixed: ROC and precision-recall curves, the area under each, and picking an operating point from asymmetric costs. Rare-positive problems live here.

on this pageshow

explore

questions

16

What does a cumulative gains chart show, and how does a lift chart differ?

level: juniorimportance: must knowfreq 62%

answer

  1. Sort by score, then slice
  2. Two axes: budget spent, value captured
  3. Random is a diagonal on one chart
  4. Random is a flat 1.0 on the other
  5. Capture rate divided by population share

basics

~20 s

A cumulative gains chart plots the share of all responders captured against the share of the file contacted, after sorting by score. A lift chart shows the same result as a multiple of random targeting, where random equals 1.

solid answer

~50 s

Both charts start the same way: score every record, sort descending, and cut the file into equal slices, usually deciles. The cumulative gains chart puts the fraction of the population contacted on the x-axis and the fraction of all positives captured on the y-axis, so random targeting is the 45-degree diagonal and a good model bows above it. If decile 1 holds 41% of all responders, that 41% capture rate is the first point on the curve. The lift chart divides rather than accumulates: cumulative lift at 10% of the file is 41% / 10% = 4.1x, meaning the top decile responds 4.1 times better than an untargeted mailing. Random targeting is a flat line at 1.0 instead of a diagonal. Gains answers "how much of the value do I get"; lift answers "how much better than nothing is this".

go deeper

for a junior

Be ready to state both axes of the gains chart without hesitating, and to say where random targeting sits on each chart: the diagonal on gains, a flat 1.0 on lift.

for a middle

Expect to compute the numbers live. Given decile counts, produce capture rate, per-decile lift and cumulative lift, and explain why cumulative lift must decay to 1.0 at the full file.

for a senior

Show that you insist on a holdout, sanity-check decile sizes for sampling noise, and warn stakeholders that a lift number is meaningless without the base rate beside it.

for a principal

Own the reporting standard: which chart the business sees, whether lift is quoted cumulatively or per decile, and how a campaign's claimed value is reconciled against realised results after the fact.

## What the charts are for A cumulative gains chart and a lift chart are the two standard ways of showing a business audience what a *ranking* is worth, without asking them to read a curve in true-positive-rate/false-positive-rate coordinates. Both assume a scored file: every customer, applicant or lead has a model score, and the plan is to work down the list from the highest score until money, time or list size runs out. ## Building them 1. Score a **holdout** sample the model never trained on. Every number below is meaningless on training data, where a flexible model can memorise who responded. 2. Sort records by score, highest first. 3. Split into equal-sized slices. Ten slices gives deciles, the near-universal convention; some teams use twenty ("demi-deciles") or a hundred. 4. For each slice count how many positives (responders, defaulters, churners) it contains. From those counts everything follows. Let `p` be the overall positive rate in the file. - **Per-decile capture rate** = positives in that decile / all positives. - **Cumulative capture (gains)** at decile d = positives in deciles 1..d / all positives. - **Per-decile lift** = (positives in the decile / records in the decile) / `p`. - **Cumulative lift** at decile d = cumulative capture / cumulative share of the population = e.g. `0.41 / 0.10 = 4.1`. ## Reading the gains chart The x-axis is the fraction of the population contacted, the y-axis the fraction of all positives captured. Three reference shapes matter: - **The diagonal.** Contact 30% at random, capture roughly 30% of responders. Any useful model sits above it. - **The model curve.** It is concave when the score rank-orders well, and it always ends at (100%, 100%) because contacting everyone captures everyone. - **The perfect curve.** A model that put every positive above every negative would climb a straight line to (p, 100%) and then flatten. That is the ceiling, and it is set by the base rate, not by the algorithm. The practical read is a single sentence: "contacting the top 30% of the file reaches 68% of the responders." That is why marketing and risk teams like it — the axis is budget, not statistics. ## Reading the lift chart The same information as a ratio. Cumulative lift starts high on the left and decays monotonically toward 1.0 at 100% of the file, because contacting everyone *is* random targeting. Per-decile (non-cumulative) lift is the noisier, more diagnostic version: it should decline across deciles and dip below 1.0 in the tail. Mixing the two up is the most common error in a review meeting — a slide reading "lift 4.1x" means very different things as decile-1 lift and as cumulative lift at 30%. ## The base rate sets a hard ceiling The top 10% of the file cannot hold more than 10% of the population, so it cannot hold more than `min(1, 0.10 / p)` of the positives. Top-decile lift is therefore capped at `min(10, 1/p)`. If 25% of customers respond, top-decile lift can never exceed 4.0 no matter how good the model is, while a 1%-response campaign has headroom to 10x. This is why lift numbers are **not comparable across campaigns with different base rates**, and why "we got 7x lift" is only impressive once you know `p`. ## What these charts do and do not tell you The y-axis of the gains chart at a given cut is exactly recall (the share of positives found) at that cut, so the chart is a rank-quality summary. It is invariant to any monotone transform of the score: rescaling scores, or squashing them into a 300-850 band, moves no point on the curve. That invariance is the point — and also the limitation. These charts say nothing about whether a predicted probability of 0.30 corresponds to a 30% real-world rate; that is a calibration question, and a badly calibrated model can have a perfect gains curve. Two further cautions. First, the curve is estimated from a finite sample: with 2,000 records, a decile holds 200, and the capture rate for that decile carries real sampling noise. Second, the shape depends on the mix of the population; if the scored file is a different mix from the one the campaign will actually go to, the curve travels badly. ## Common vocabulary You will hear "capture rate", "cumulative response rate", "decile analysis" and "gains table" for essentially the same artefact. In a direct-marketing shop the deliverable is usually the *table*, ten rows of counts, response rates, lift and cumulative capture, with the chart as decoration.

  • A direct-mail cross-sell budget only covers the top 3 deciles. What does the gains curve tell the marketer?
    Read the curve at 30% of the file. If it sits at 68%, the campaign is expected to reach 68% of everyone who would have responded, at 30% of the postage of a full mailing — a cumulative lift of about 2.3x. The slope between deciles 3 and 4 also shows what the next tranche of budget would buy, which is the argument for or against extending it.
  • Why is a top-decile lift of 8x impossible if 25% of the file responds?
    The top decile is 10% of the population, so it can hold at most 10/25 = 40% of the responders. That caps its capture rate at 40% and its lift at 0.40 / 0.10 = 4.0. The ceiling on top-decile lift is min(10, 1/base rate), so high-prevalence problems produce small lift numbers even with an excellent model.
  • Does a good gains curve mean the model's predicted probabilities are trustworthy?
    No. The curve depends only on the ordering of scores, so any monotone transform leaves it unchanged. A model that outputs 0.9 for everyone who responds and 0.8 for everyone who does not has a superb gains curve and useless probabilities. Whether a score can be read as a rate is a separate calibration question.

Gains is the odometer: how far you have got. Lift is the speedometer reading relative to walking: how much faster than random you are covering ground.

saying these in an interview costs you the question

  • Says the gains chart's x-axis is the score threshold value
  • Reports lift and gains from the training sample
  • Quotes decile-1 lift as if it were cumulative lift
  • Claims lift above 1 proves the campaign caused the responses
  • Compares lift across campaigns with different base rates
  • Thinks a good gains curve implies well-calibrated probabilities

context

open as a page

What does a precision-recall curve plot, and what does average precision measure?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A precision-recall curve plots precision against recall as the decision threshold sweeps from strict to permissive. Average precision compresses that curve into one number: the recall-weighted mean of the precisions along it, estimating the area beneath.

open as a page

What do the two axes of a binary classifier's ROC curve show, and what traces the curve?

level: juniorimportance: must knowfreq 88%

basics

~20 s

An ROC curve plots true positive rate on the y-axis against false positive rate on the x-axis. Each point is one decision threshold, and sweeping the threshold from strictest to loosest traces the curve from (0,0) to (1,1).

open as a page

Lowering a trained classifier's decision threshold from 0.5 to 0.3 changes what, and what stays fixed?

level: juniorimportance: must knowfreq 78%

basics

~20 s

More items are labelled positive, so true positives and false positives can only rise, recall can only rise, and precision usually falls. The fitted model, its scores and its ranking of items do not change at all.

open as a page

A fraud model shows ROC-AUC 0.97 but average precision 0.15. Why the gap?

level: middleimportance: must knowfreq 66%

basics

~20 s

Both are right; they use different denominators. With 0.2% fraud, tens of thousands of false positives barely move the false positive rate, but they dominate the flagged set and crush precision. Read average precision against a 0.002 baseline, not 0.5.

open as a page

What does an ROC-AUC of 0.78 mean in probability terms for a display-ad click model?

level: middleimportance: must knowfreq 76%

basics

~20 s

Draw one clicker and one non-clicker at random: the model scores the clicker higher 78% of the time, counting ties as half a win. ROC-AUC is a statement about ranking order, not about the numeric size of the scores.

open as a page

How do you turn a cost matrix over false positives and false negatives into a decision threshold?

level: middleimportance: must knowfreq 62%

basics

~20 s

Flag an item when the expected cost of flagging is below the expected cost of not flagging. With false-positive and false-negative costs only, the break-even is cost_FP / (cost_FP + cost_FN), which is 0.5 only for equal costs.

open as a page

In credit scoring, what do a KS of 38 and a Gini of 0.52 mean?

level: middleimportance: should knowfreq 48%

basics

~20 s

KS 38 means that at its best cut-off the scorecard separates 38 percentage points more of the bads than of the goods. Gini 0.52 is a whole-curve summary equal to two times AUC minus one, so AUC is 0.76.

open as a page

In a decile lift table, what does non-monotonic lift across deciles tell you?

level: middleimportance: should knowfreq 40%

basics

~20 s

It means the score stops rank-ordering cleanly in that region. Usually it is sampling noise in small deciles; sometimes it is a genuinely unstable model, a shifted population, or a scoring bug. Check decile sizes before you diagnose anything else.

open as a page

Why is ROC-AUC unchanged after an evaluation set is downsampled from 3% to 50% positives?

level: seniorimportance: should knowfreq 52%

basics

~20 s

True positive rate is computed only among positives and false positive rate only among negatives, so randomly discarding negatives leaves both rates unchanged in expectation. ROC-AUC measures ranking quality and is therefore blind to the class mix.

open as a page

A promo-abuse detector's positive rate triples during a January sale. What happens at its fixed threshold?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Flagged volume rises sharply. If only the base rate moved, precision at the fixed cut rises while recall holds, so the queue floods rather than degrades. If the traffic itself changed shape, false positives rise and precision can fall.

open as a page

A threshold picked to hit 85% precision on validation delivers less in production. Why?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Picking the cut that first clears 85% in a sweep selects the point where sampling noise flattered you, and precision at a strict cut rests on few items. The estimate is optimistic; aim above the target on untouched data.

open as a page

What do Youden's J and maximising F1 each assume when used to pick a classification threshold?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Youden's J, sensitivity plus specificity minus one, weights both error rates equally and ignores class sizes. Maximising F1 ignores true negatives entirely and implies a trade-off that shifts with the operating point rather than a fixed cost ratio.

open as a page

A telemarketing model shows 4x lift in decile 1 — why might the campaign not lift response 4x?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Lift measures who is likely to respond, not who responds because you called. Much of the top decile would have bought anyway, so targeting them redistributes credit rather than creating sales. Measuring the campaign's real effect needs a randomised control group.

open as a page

A precision-recall curve built on only 60 positives is jagged. How much do you trust its average precision?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Not much as a point value. The curve's resolution comes from the positive count, not the row count: with 60 positives each is a 1.7% recall step, so average precision carries sampling variance. Report a resampled interval, not three decimals.

open as a page

Two candidate models' ROC curves cross - how do you decide which one to ship?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Crossing curves mean neither model dominates - each wins over a different range of false positive rates. Compare them where you will actually operate, such as true positive rate at a fixed low false positive rate, not by total area.

open as a page