skip to content

How do you evaluate an uplift model when no customer's individual treatment effect is observed?

level: seniorimportance: nice to knowfreq 32%

answer

  1. individual effects are never observed
  2. evaluate groups, not units
  3. rank, then compare arms within slices
  4. cumulative incremental gain versus random
  5. area between curve and diagonal

basics

~20 s

Score the ranking, not the units. Sort a held-out randomized sample by predicted uplift and compare treated with untreated outcomes inside each slice: a good model shows a large incremental gap at the top that shrinks further down. The Qini curve summarises that.

solid answer

~50 s

No unit's own effect is ever observed, so per-unit error metrics do not exist and evaluation moves to the group level on held-out randomized data. Rank the holdout by predicted uplift and walk down it: at each cumulative fraction, compute incremental responders as treated responders minus control responders rescaled to the same population size. Plotting that against the fraction targeted gives the uplift or Qini curve; targeting in random order traces a straight diagonal, and the area between the curve and that diagonal — the Qini coefficient — summarises what the ranking is worth. The blunt version is a decile plot: the observed treated-minus-control gap within each predicted-uplift decile, which should decline across deciles. What none of this is: the AUC of a response model, which measures how well outcomes are ranked and can be excellent while the effect ranking is worthless.

go deeper

for a junior

Know that a unit's true treatment effect is never observed, so uplift models cannot be scored against individual labels the way ordinary predictions are.

for a middle

Explain the mechanics of an uplift or Qini curve: rank by predicted effect, compare treated with untreated cumulatively, and read the gap against random targeting.

for a senior

Demonstrate holdout discipline — randomized evaluation data, intervals on slice estimates, and a treatment depth chosen where incremental value still exceeds cost.

for a principal

Own the measurement standard the organisation is judged by: incremental outcome per unit treated, maintained as a permanent practice rather than a one-off model bake-off.

## Why the usual scoring does not apply Every unit is observed under one condition only, so its true effect `Y_i(1) - Y_i(0)` never exists as a label. There is nothing to compute a squared error against. An uplift model therefore cannot be evaluated the way a churn model is, and any workflow that reports accuracy, log loss or AUC against an outcome label is measuring a different model than the one you built. What can be evaluated is the *ranking*. If the model orders units by how much the treatment moves them, then groups high in the order should show a bigger treated-versus-untreated gap than groups low in the order — and gaps between groups are estimable whenever the holdout contains both arms. ## The evaluation setup You need a holdout that the model never saw, containing randomized treated and untreated units. Randomization inside the holdout is what makes each slice's treated-minus-untreated difference an unbiased effect estimate for that slice. Score every holdout unit, sort descending by predicted uplift, and then work cumulatively down the list. ## The uplift and Qini curves At each cumulative fraction of the ranked list, count responders among the treated units in that prefix and among the untreated units in that prefix. The two arms are rarely the same size inside a prefix, so the control count is rescaled by the ratio of treated to control units before subtracting: `incremental(k) = R_treated(k) - R_control(k) x (N_treated(k) / N_control(k))` Plot `incremental(k)` against `k`. That is the Qini curve. Two reference lines make it readable. Targeting in random order produces a straight line from the origin to the total incremental responders at 100% coverage — because a random prefix has the average effect. A perfect ranking rises steeply, plateaus once the persuadables are exhausted, and can bend downward at the end if a negatively affected group sits at the bottom. The Qini coefficient is the area between the model's curve and the random diagonal. Bigger is better; zero means the ranking is no better than shuffling; negative means it is worse than shuffling, which happens more often than teams expect when a response score is mistaken for an effect score. ## The decile plot A cruder but more legible diagnostic: split the ranked holdout into ten equal bins, and in each bin compute the treated-minus-untreated difference with an interval. A working model gives a downward staircase — largest gap in the top bin, near zero in the middle, possibly negative at the bottom. Non-monotone bins are the norm rather than the exception, because each bin estimates a difference from a tenth of the data. Widen to quintiles or halves before declaring the model dead, and always draw the intervals; a plot without them invites reading noise as structure. ## Why response-model AUC misleads here AUC on a response model asks how well predicted scores separate responders from non-responders. That is a statement about outcome levels. Targeting is judged on change. The two rankings systematically diverge, because the units most likely to respond very often respond under either condition — a model can be near-perfect at predicting conversion and still concentrate spend on people the treatment does not move. Reporting response AUC as evidence for a targeting policy is the single most common evaluation error in this area. ## Choosing the treatment depth The curve does double duty: it ranks candidate models and it sets the operating point. Keep extending down the list while the incremental outcome gained from the next slice exceeds the cost of treating that slice. The peak of the curve is where remaining units contribute nothing, and any downward slope past it is the negatively affected tail actively destroying value. Treating everyone is a decision to include that tail. ## Practical cautions The holdout has to be big enough that slice-level differences are estimable at all — a rule of thumb is to work backwards from the width of the interval you would need to distinguish the top slice from the average. Compare models on the same holdout and the same slicing, since curve shape is sensitive to both. And keep a control arm after rollout: the curve estimated before launch is a prediction about a policy, and the realised effect on live traffic is the only thing that confirms it. ## The one-line answer Individual effects are unobservable, so you evaluate the ranking with group-level treated-versus-untreated comparisons on randomized holdout data, summarised by an uplift or Qini curve against random targeting — never by a response model's AUC.

  • Why is a high AUC on a response model not evidence of a good targeting policy?
    AUC measures how well the score separates responders from non-responders, which is a statement about outcome levels. A targeting policy is judged on the change it causes, and the two rankings diverge because the likeliest responders often respond under either condition. A model can be nearly perfect at predicting who converts and still direct the entire budget at people whose behaviour the treatment never moves.
  • Your uplift decile plot is noisy and non-monotone. What does that tell you?
    Usually that the holdout is too small, since each decile estimates a treated-minus-untreated difference from a tenth of the data. Widen to quintiles or halves and put intervals on every bar before drawing conclusions. If the top group still shows no larger incremental gap than the bottom after that, the ranking genuinely carries no effect signal and the model should not be shipped.
  • How do you pick how far down the ranked list to treat?
    From the curve plus the economics: keep extending while the incremental outcome from the next slice exceeds the cost of treating it. The uplift curve typically rises, flattens once the persuadable population is exhausted, and can fall where a negatively affected group sits. The operating point is at or before the peak, and treating the whole list is a deliberate decision to absorb the flat and negative tail.

saying these in an interview costs you the question

  • Wants per-unit accuracy for an unobservable individual effect
  • Judges a targeting policy by response-model AUC
  • Reads a noisy decile plot as proof of no signal
  • Evaluates uplift using the treated arm only
  • Forgets the holdout itself must be randomized

context