skip to content

A telemarketing model shows 4x lift in decile 1 — why might the campaign not lift response 4x?

level: seniorimportance: nice to knowfreq 32%

answer

  1. Ranking claim, not a causal one
  2. The denominator is random targeting, not no contact
  3. Some buyers would have bought anyway
  4. Contact can also annoy people away
  5. Randomised holdout inside the contacted deciles

basics

~20 s

Lift measures who is likely to respond, not who responds because you called. Much of the top decile would have bought anyway, so targeting them redistributes credit rather than creating sales. Measuring the campaign's real effect needs a randomised control group.

solid answer

~40 s

The 4x is a statement about ranking: the top decile's response rate is four times the file's average under the same treatment. It is not a claim that calling those people multiplies responses by four. Two things break the leap. First, the model learned who responded in past campaigns, so it ranks propensity, and much of the top decile would have converted anyway — you are harvesting, not causing. Second, the baseline in the ratio is "a randomly targeted call", not "no call at all". Some high-propensity customers can even be annoyed into churning, a negative effect the chart cannot see. The fix is a randomised control withheld inside the contacted deciles, compared on response. That gives incremental response, and ranking by that difference is uplift modelling, which often picks a different decile 1.

go deeper

for a junior

Know that lift compares targeted contact to random contact, not to no contact at all, so it cannot by itself say how many sales a campaign created.

for a middle

Explain why a model trained on who responded learns propensity rather than effect, and how a randomised holdout inside the contacted group turns response into incremental response.

for a senior

Demonstrate that you would design the holdout before launch, report incremental rather than raw response, and expect the effect ranking to differ from the propensity ranking.

for a principal

Own the measurement contract with the business: that every campaign carries a control arm, what the holdout costs, and how incremental results rather than lift slides govern the next budget.

## What the 4x actually asserts A decile-1 lift of 4.0 says: among scored customers, the tenth of the file with the highest scores responded at four times the overall rate. It is a **ranking** claim, evaluated on historical data where everyone in the sample got broadly the same treatment. Nothing in its construction compares contacted customers to uncontacted ones, so nothing in it licenses the sentence "calling these people produces four times the sales". This matters because that sentence is exactly what gets said in the campaign readout, and it is where analytics teams lose credibility. The campaign runs, the finance team compares actual sales to the promised uplift, and the numbers do not reconcile. ## The three reasons the leap fails **1. Propensity is not causality.** The label the model learned is "responded", collected from people who were contacted. High-scoring customers are those most inclined to buy — often because they were already in-market, already loyal, or already browsing. A meaningful share of them would have bought through the website, a branch, or a later visit with no call at all. Calling them converts a sale that was going to happen anyway into a sale attributed to the campaign. Total revenue barely moves; the attribution moves a lot. **2. The baseline in the ratio is the wrong counterfactual.** Lift divides the top decile's response rate by the whole file's response rate. Both of those populations were treated. The denominator is "call someone at random", not "call nobody". So even taken entirely at face value, 4x says targeted calling beats untargeted calling by four times — a real and useful claim about *allocation* of a fixed call budget, and a claim about nothing else. If the question on the table is whether to run the campaign at all, the lift chart has no opinion. **3. Treatment effects are not uniform, and some are negative.** Customers split roughly into four groups: those who respond whether or not you call (sure things), those who respond only if called (persuadables), those who never respond (lost causes), and those who would have responded but are pushed away by the intrusion (sleeping dogs, or do-not-disturb). A propensity model scores sure things and persuadables alike at the top, and it is blind to sleeping dogs entirely — they look like good prospects right up until the call triggers a complaint or a cancellation. Only the persuadables generate incremental value. ## How to measure the effect that is real Randomise. Within the deciles you plan to contact, withhold a random sample — 5-10% is typical — and leave them alone. Then, per decile, compare the response rate of treated to that of the untreated control. The difference is the **incremental response rate**, and multiplied by decile size it is the campaign's actual contribution. Two things almost always emerge from a first such holdout. First, incremental response is much smaller than raw response. A decile responding at 12% against a control responding at 9% contributes three points, not twelve. Second, the ordering changes. Incremental response frequently peaks in a *middle* decile, not the top one, because the very top is dense with people who convert regardless. A gains chart drawn on incremental response looks nothing like the one drawn on raw response, and it is the one that should drive budget. ## Uplift modelling in one paragraph Once a randomised treated/control experiment exists, you can model the difference directly rather than modelling response. Two-model approaches fit one response model on the treated arm and one on the control arm and rank by the gap; single-model approaches include treatment as a feature and score each customer twice, with and without treatment; other formulations transform the label so that a standard classifier estimates the difference. All of them rank by estimated *effect of contact*, and all of them require experimental data — you cannot recover an incremental ranking from a file where everyone was treated. The corresponding chart is a Qini or incremental-gains curve, read the same way as an ordinary gains curve but with the y-axis measuring incremental responses captured. ## What to say in the room The defensible version of the claim is: "Under a fixed call budget, targeting the top decile should produce about four times the responses of calling the same number of people at random." Then add the second sentence that keeps you honest: "How many of those responses the campaign *created* is a separate question, and we will answer it with a randomised holdout of 5% of the contacted list." That combination — using lift correctly for allocation while refusing to use it for incrementality — is what distinguishes a senior analyst from someone reciting a chart. One practical note: the holdout costs real money in forgone contacts, and someone will ask to skip it. The counter is that without it, no campaign can ever be evaluated, only described — and every future budget argument becomes an argument about opinions.

  • How large should the randomised control group be, and where does it sit?
    It sits inside the deciles you are contacting, not in the untargeted tail, or you cannot compare like with like. Size is set by the difference you need to detect: to resolve a two-point difference on a ten-point base rate you need thousands per arm, so 5-10% of a large contacted list is typical. On a small campaign, holding out a share too small to be conclusive is worse than pooling several campaigns.
  • Why can incremental response peak in a middle decile rather than the top one?
    The top decile is crowded with customers who would convert with or without contact, so the treated and control arms both respond highly and the gap is thin. Further down sit customers whose conversion genuinely depends on the prompt, producing a larger difference. Ranking by response and ranking by effect of contact are different orderings.
  • What is a sleeping dog, and why can a propensity model never see one?
    A customer who would have responded but is driven away by being contacted — a negative treatment effect, such as a call that triggers a cancellation. A propensity model is trained on who responded when contacted, so a suppressed response looks identical to a plain non-response. Only a treated-versus-control comparison exposes the negative gap.

A weather forecast ranks which days will be sunny. Standing outside on the top-ranked days does not make the sun shine — and the forecast never claimed it would.

saying these in an interview costs you the question

  • Promises the business four times more sales from 4x lift
  • Says a control group wastes budget and can be skipped
  • Places the control group in the uncontacted deciles
  • Assumes contact never reduces the chance of a response
  • Confuses response rate with incremental response rate
  • Believes uplift can be estimated without experimental data

context