In a decile lift table, what does non-monotonic lift across deciles tell you?
answer
- Expected shape: steady decline
- Count positives, not records
- Small counts wobble a lot
- Ties mean an arbitrary split
- Top deciles matter, tail rarely does
basics
~20 sIt means the score stops rank-ordering cleanly in that region. Usually it is sampling noise in small deciles; sometimes it is a genuinely unstable model, a shifted population, or a scoring bug. Check decile sizes before you diagnose anything else.
solid answer
~50 sA healthy decile table falls steadily — say 4.1x in decile 1 down to 0.6x by decile 8 — because each slice holds lower-scoring records. A reversal, where decile 3 lifts higher than decile 2, says the ordering inside that band is no better than a coin flip. My first check is arithmetic: how many positives are actually in each decile? With 200 records and a 3% response rate a decile holds six responders, and a swing of two is nothing. If the counts make the reversal real, I look for a structural cause: a segment the model scores badly, a feature that broke for part of the file, or overfitting so the holdout ordering does not hold. A wobbly mid-decile rarely deserves action; reversals in the top two deciles do, because that is where the campaign spends.
code
python · 22 linesimport random
random.seed(7)
rows = []
for _ in range(2000):
score = random.random()
responded = 1 if random.random() < 0.02 + 0.16 * score else 0
rows.append((score, responded))
rows.sort(key=lambda r: -r[0])
total_resp = sum(y for _, y in rows)
base_rate = total_resp / len(rows)
size = len(rows) // 10
cum = 0
for d in range(10):
chunk = rows[d * size:(d + 1) * size]
resp = sum(y for _, y in chunk)
cum += resp
lift = (resp / len(chunk)) / base_rate
capture = 100 * cum / total_resp
print("decile", d + 1, "lift", round(lift, 2), "cum capture %", round(capture, 1))go deeper
Know what the table's columns are and that lift should generally fall from decile 1 downward. Being able to say a reversal means the score is not ordering well there is enough at this level.
Explain the mechanics: equal-sized deciles mean unequal positive counts, so mid-table lift is noisy. Be ready to estimate that noise roughly from the positive count and to name ties as a second cause.
Show a diagnostic order — counts, then score plateaus, then segment mixing, then pipeline faults — and judge severity by whether the affected deciles are the ones the campaign will actually use.
Set the standard for how model performance is reported: positive counts shown beside lift, two holdout periods, and a stated rule for when a reversal triggers model work versus a footnote.
## The expected shape A decile lift table is built by sorting a scored holdout descending, cutting it into ten equal slices, and reporting for each: records, positives, response rate, lift against the overall rate, and cumulative capture. Because the slices are ordered by predicted propensity, the response rate should decline as you go down. A typical well-behaved table falls from around 4.1x in decile 1 to about 0.6x by decile 8 and lower in the tail. Monotone decline is the visual proof that the score does what it claims: separate likely positives from unlikely ones across the whole range, not only at the top. ## What a reversal actually means When decile 5 lifts above decile 4, the model has assigned higher scores to a group that responded less. Within that band the ranking carries no information. There are four families of cause, and they are worth working through in order of likelihood. **1. Sampling noise — nearly always check this first.** Deciles are equal in *records*, not in *positives*. With a 2% base rate and 5,000 holdout records, each decile has 500 records and the middle deciles hold perhaps 8-12 positives each. A difference of three responders swings the lift by a third. Compute a rough interval: the count of positives in a decile behaves like a binomial draw, so its standard deviation is near the square root of the count. Ten positives carries a plus-or-minus three of pure noise. Most mid-table reversals evaporate at this step, and the honest reporting fix is to use quintiles instead of deciles, or to pool more holdout data. **2. Genuine flatness in the score distribution.** Many models pile a large mass of records at nearly identical scores — a tree ensemble with limited depth emits a finite set of distinct values. If deciles 4 through 7 all sit inside one plateau of tied scores, the split between them is arbitrary, decided by whatever tie-break the sort used. Then the ordering across those deciles is genuinely meaningless and no amount of data will make it monotone. Look at the score range spanned by each decile: if deciles 4-7 span 0.11 to 0.12, that is the diagnosis. **3. Population mixing.** The file may be a blend of segments with different base rates — new versus tenured customers, two channels, two countries. If the model scores one segment systematically low but that segment responds well, its records land in a mid decile and push its lift up. The tell is that the reversal disappears when the table is rebuilt within each segment. This is a real model defect, not a reporting artefact, and it usually means a missing interaction or a missing segment feature. **4. Something broke.** A feature that is null for part of the file, a scoring pipeline that applied the wrong transform, a label joined on the wrong date, or a leak that made the training ordering unreproducible on fresh data. Overfitting shows up as a table that is beautifully monotone on training data and ragged on the holdout — always build the table on data the model never saw, and be suspicious when only the holdout misbehaves. ## Where in the table the reversal sits Severity depends on location. A campaign that mails the top three deciles cares intensely about the ordering of deciles 1-3, because that ordering determines both who is contacted and the expected response. Reversals down at deciles 6-9 affect nobody in that campaign — the records are not being contacted, and the tail is where counts are smallest and noise largest. Do not spend a week fixing decile 7. Do escalate a table where decile 2 outperforms decile 1: that says the very top of the score is not the best part of the model, which undermines the whole targeting story. ## What to report A decile table is a point estimate and should be shown as one. Good practice is to include raw positive counts next to every lift figure so a reader can judge stability themselves, to rebuild the table on at least two disjoint holdout periods and show both, and to say plainly which deciles the campaign will use. If a reversal survives all of that on the deciles that matter, the response is model work — segment features, a monotonic constraint if the domain justifies one, or more training data in the affected region — not a prettier chart. ## The cumulative view hides it Cumulative lift is monotone almost by construction: it starts high and decays toward 1.0, and a single bad decile barely bends it. That smoothness is why cumulative curves look reassuring. Always read the per-decile column too, because that is the one that exposes where the score stops working.
- How would you decide whether a reversal between deciles 4 and 5 is real?Count the positives in each. Their counts behave roughly like binomial draws, so the noise is on the order of the square root of the count — ten positives carry about plus-or-minus three. If the gap is inside that band, it is noise. If it survives, rebuild the table on a second, disjoint holdout period and see whether the same two deciles swap again.
- Why can a table look perfectly monotone on training data and ragged on a holdout?On training data the model has partly memorised which records responded, so the ordering it produces is self-fulfilling. The holdout removes that advantage and shows the ordering the score actually generalises. A large gap between the two tables is a direct overfitting signal, and only the holdout table should ever be reported.
- What does it mean if deciles 4 through 7 all span nearly the same score range?The score distribution has a plateau, often because the model emits a small set of distinct values. The boundaries between those deciles are then set by tie-breaking rather than by the model, so their relative order is arbitrary and will not become monotone with more data. Reporting quintiles or grouping the plateau is more honest.
saying these in an interview costs you the question
- Treats any reversal as proof the model is broken
- Never checks how many positives sit in each decile
- Builds the decile table on training data
- Reorders or smooths deciles to make the slide look monotone
- Ignores that ties make decile boundaries arbitrary
- Spends effort fixing tail deciles nobody will contact