What does a cumulative gains chart show, and how does a lift chart differ?
answer
- Sort by score, then slice
- Two axes: budget spent, value captured
- Random is a diagonal on one chart
- Random is a flat 1.0 on the other
- Capture rate divided by population share
basics
~20 sA cumulative gains chart plots the share of all responders captured against the share of the file contacted, after sorting by score. A lift chart shows the same result as a multiple of random targeting, where random equals 1.
solid answer
~50 sBoth charts start the same way: score every record, sort descending, and cut the file into equal slices, usually deciles. The cumulative gains chart puts the fraction of the population contacted on the x-axis and the fraction of all positives captured on the y-axis, so random targeting is the 45-degree diagonal and a good model bows above it. If decile 1 holds 41% of all responders, that 41% capture rate is the first point on the curve. The lift chart divides rather than accumulates: cumulative lift at 10% of the file is 41% / 10% = 4.1x, meaning the top decile responds 4.1 times better than an untargeted mailing. Random targeting is a flat line at 1.0 instead of a diagonal. Gains answers "how much of the value do I get"; lift answers "how much better than nothing is this".
go deeper
Be ready to state both axes of the gains chart without hesitating, and to say where random targeting sits on each chart: the diagonal on gains, a flat 1.0 on lift.
Expect to compute the numbers live. Given decile counts, produce capture rate, per-decile lift and cumulative lift, and explain why cumulative lift must decay to 1.0 at the full file.
Show that you insist on a holdout, sanity-check decile sizes for sampling noise, and warn stakeholders that a lift number is meaningless without the base rate beside it.
Own the reporting standard: which chart the business sees, whether lift is quoted cumulatively or per decile, and how a campaign's claimed value is reconciled against realised results after the fact.
## What the charts are for A cumulative gains chart and a lift chart are the two standard ways of showing a business audience what a *ranking* is worth, without asking them to read a curve in true-positive-rate/false-positive-rate coordinates. Both assume a scored file: every customer, applicant or lead has a model score, and the plan is to work down the list from the highest score until money, time or list size runs out. ## Building them 1. Score a **holdout** sample the model never trained on. Every number below is meaningless on training data, where a flexible model can memorise who responded. 2. Sort records by score, highest first. 3. Split into equal-sized slices. Ten slices gives deciles, the near-universal convention; some teams use twenty ("demi-deciles") or a hundred. 4. For each slice count how many positives (responders, defaulters, churners) it contains. From those counts everything follows. Let `p` be the overall positive rate in the file. - **Per-decile capture rate** = positives in that decile / all positives. - **Cumulative capture (gains)** at decile d = positives in deciles 1..d / all positives. - **Per-decile lift** = (positives in the decile / records in the decile) / `p`. - **Cumulative lift** at decile d = cumulative capture / cumulative share of the population = e.g. `0.41 / 0.10 = 4.1`. ## Reading the gains chart The x-axis is the fraction of the population contacted, the y-axis the fraction of all positives captured. Three reference shapes matter: - **The diagonal.** Contact 30% at random, capture roughly 30% of responders. Any useful model sits above it. - **The model curve.** It is concave when the score rank-orders well, and it always ends at (100%, 100%) because contacting everyone captures everyone. - **The perfect curve.** A model that put every positive above every negative would climb a straight line to (p, 100%) and then flatten. That is the ceiling, and it is set by the base rate, not by the algorithm. The practical read is a single sentence: "contacting the top 30% of the file reaches 68% of the responders." That is why marketing and risk teams like it — the axis is budget, not statistics. ## Reading the lift chart The same information as a ratio. Cumulative lift starts high on the left and decays monotonically toward 1.0 at 100% of the file, because contacting everyone *is* random targeting. Per-decile (non-cumulative) lift is the noisier, more diagnostic version: it should decline across deciles and dip below 1.0 in the tail. Mixing the two up is the most common error in a review meeting — a slide reading "lift 4.1x" means very different things as decile-1 lift and as cumulative lift at 30%. ## The base rate sets a hard ceiling The top 10% of the file cannot hold more than 10% of the population, so it cannot hold more than `min(1, 0.10 / p)` of the positives. Top-decile lift is therefore capped at `min(10, 1/p)`. If 25% of customers respond, top-decile lift can never exceed 4.0 no matter how good the model is, while a 1%-response campaign has headroom to 10x. This is why lift numbers are **not comparable across campaigns with different base rates**, and why "we got 7x lift" is only impressive once you know `p`. ## What these charts do and do not tell you The y-axis of the gains chart at a given cut is exactly recall (the share of positives found) at that cut, so the chart is a rank-quality summary. It is invariant to any monotone transform of the score: rescaling scores, or squashing them into a 300-850 band, moves no point on the curve. That invariance is the point — and also the limitation. These charts say nothing about whether a predicted probability of 0.30 corresponds to a 30% real-world rate; that is a calibration question, and a badly calibrated model can have a perfect gains curve. Two further cautions. First, the curve is estimated from a finite sample: with 2,000 records, a decile holds 200, and the capture rate for that decile carries real sampling noise. Second, the shape depends on the mix of the population; if the scored file is a different mix from the one the campaign will actually go to, the curve travels badly. ## Common vocabulary You will hear "capture rate", "cumulative response rate", "decile analysis" and "gains table" for essentially the same artefact. In a direct-marketing shop the deliverable is usually the *table*, ten rows of counts, response rates, lift and cumulative capture, with the chart as decoration.
- A direct-mail cross-sell budget only covers the top 3 deciles. What does the gains curve tell the marketer?Read the curve at 30% of the file. If it sits at 68%, the campaign is expected to reach 68% of everyone who would have responded, at 30% of the postage of a full mailing — a cumulative lift of about 2.3x. The slope between deciles 3 and 4 also shows what the next tranche of budget would buy, which is the argument for or against extending it.
- Why is a top-decile lift of 8x impossible if 25% of the file responds?The top decile is 10% of the population, so it can hold at most 10/25 = 40% of the responders. That caps its capture rate at 40% and its lift at 0.40 / 0.10 = 4.0. The ceiling on top-decile lift is min(10, 1/base rate), so high-prevalence problems produce small lift numbers even with an excellent model.
- Does a good gains curve mean the model's predicted probabilities are trustworthy?No. The curve depends only on the ordering of scores, so any monotone transform leaves it unchanged. A model that outputs 0.9 for everyone who responds and 0.8 for everyone who does not has a superb gains curve and useless probabilities. Whether a score can be read as a rate is a separate calibration question.
Gains is the odometer: how far you have got. Lift is the speedometer reading relative to walking: how much faster than random you are covering ground.
saying these in an interview costs you the question
- Says the gains chart's x-axis is the score threshold value
- Reports lift and gains from the training sample
- Quotes decile-1 lift as if it were cumulative lift
- Claims lift above 1 proves the campaign caused the responses
- Compares lift across campaigns with different base rates
- Thinks a good gains curve implies well-calibrated probabilities