Your org pilots every programme on its worst-performing segment and reports the rebound as impact — how do you fix that?
answer
- the evaluation is broken, not the targeting
- selection on an extreme guarantees a rebound
- counterfactual chosen by the same rule
- publish the expected rebound as the null
- change the incentive on before-and-after deltas
basics
~20 sChange the evaluation, not the targeting. Selecting the worst segment guarantees a rebound whatever the programme does, so require an untreated comparison chosen by the same extreme rule and a published expected-rebound figure any claimed effect must beat.
solid answer
~50 sThe defect is structural, not analytical: when the selection rule is 'worst on the metric we will measure again', a favourable-looking result is guaranteed, so the evaluation system cannot tell a good programme from a useless one and will keep funding both. Keep targeting the bottom segment if that is where the value is — only the counterfactual has to change. In order of strength: randomise which eligible units get the programme first, or stagger rollout so a comparably extreme slice stays untreated for a period; failing that, build the baseline from several pre-periods rather than the selection period; measure the outcome on a different instrument from the one used to select; and publish, before launch, the rebound predicted by the metric's period-to-period correlation as the null the programme must clear. Finally, stop rewarding teams for raw before-and-after deltas, because that incentive is what produced the habit.
go deeper
You are unlikely to be asked to own this call, but recognise the pattern when a plan says 'pilot on the worst segment and compare after to before', and raise it as a question rather than assuming the analysis is fine.
Be able to compute the expected rebound from the metric's period-to-period correlation and present it as the number the programme has to beat. That calculation is often the most useful contribution in the room.
Design the fix under constraints: randomise within the eligible pool, stagger the rollout, or fall back to multi-period baselines, and say clearly what each option does and does not remove.
Argue it as capital allocation, not statistics: an evaluation whose result does not depend on programme quality cannot inform funding. Then change the standing review templates and the incentives that reward raw before-and-after deltas.
## Diagnose the failure at the system level The problem is not that one analysis was sloppy. It is that the organisation has a standing rule — pick the bottom decile, run the programme, compare after to before — which manufactures a positive-looking result independent of whether the programme does anything. Under that rule the evaluation has almost no diagnostic power: an excellent programme and a worthless one both produce a rebound, so budget, promotion and roadmap decisions are being made on a number that does not vary with the thing it is supposed to measure. That is the argument to make to leadership, and it is a stronger one than 'the statistics are wrong', because it lands as a capital-allocation problem. The magnitude is estimable, which is what makes the case concrete rather than academic. Take untreated units, correlate the metric in one period with the same metric in the next, and you have the shrinkage factor. Multiply the treated group's selected deviation from the mean by that factor and you have the level the group is expected to reach with no programme at all. If your last three programmes all reported gains smaller than that figure, you have evidence that the pipeline has been booking artifact as impact. ## Fix the counterfactual, not the targeting A common wrong turn is to propose that programmes stop targeting the worst segment. Usually that targeting is commercially correct — that is where the recoverable value sits — and proposing to change it will get the whole recommendation rejected. Say explicitly that who receives the programme does not change; only how it is judged does. The options, strongest first: **Randomise within the eligible pool.** Almost always more units qualify as 'bottom segment' than can be served at once, because capacity is limited. Allocating the first wave at random among qualifying units gives an untreated group that went through exactly the same extreme selection, so both regress identically and the difference is the effect. This costs nothing but sequencing discipline. **Stagger the rollout.** Where randomisation is politically impossible, treat the qualifying units in waves and use the not-yet-treated waves as the comparison for a defined window. Everyone gets the programme; the comparison is temporary. **Multi-period baselines.** Where no untreated comparison exists, ban the selection period as the baseline. Use an average over several prior periods, which contains far less of the specific fluctuation that triggered selection. **Independent remeasurement.** Where possible, do not evaluate on the same noisy instrument used to select. If reps were chosen on last quarter's closed revenue, evaluating on a differently-constructed outcome breaks part of the shared noise term and reduces the artifact. **A published null.** Before launch, require every programme brief to state the expected rebound computed from the metric's period-to-period correlation. This is the single cheapest change: it converts the artifact from an invisible bias into a number on the same slide as the target, and a result that merely matches it is read as no effect. ## Fix the incentives None of this holds if teams are still rewarded for the size of a before-and-after delta on a self-selected group. The measurement convention and the incentive have to move together: evaluation standards belong to a function that does not own the programme's success metric, review templates should require the counterfactual to be named, and 'we targeted the worst segment and it improved' should be treated in review as an unanswered question rather than as a result. ## Know the limits of your own fix Be honest about what the corrections do not solve. A multi-period baseline reduces but does not eliminate the bias, since the selection period usually sits inside the average and the selection still favoured high-variance units. Randomising within the eligible pool answers the effect question for that pool only; it says nothing about units outside it. And no design rescues an evaluation whose outcome window is a single noisy period — insist on enough post-period measurement that a one-step rebound and a sustained change are distinguishable, since the artifact returns a group once to its own level while a real effect keeps holding. ## What the interviewer is testing This is a judgment question with no single right answer, and the strong signals are: separating the targeting decision from the evaluation decision instead of conflating them; proposing something implementable under real political constraints rather than only the textbook ideal; quantifying the expected artifact rather than merely naming it; and recognising that the durable fix is to the standing process and its incentives, not to one analysis. A weak answer stops at 'that's regression to the mean' without saying what the organisation should do differently on Monday.
- How would you compute the expected rebound so teams can benchmark against it?Among untreated units, correlate the metric in the selection period with the metric in the following period. Multiply that correlation by the treated group's deviation from the overall mean, and you have the level the group is expected to reach with no programme. Publish it beside the target so a result that merely matches it is read as no effect.
- Targeting the worst segment is often commercially correct — does your fix change who is served?No, and saying so early is what makes the recommendation survivable. Keep targeting the bottom segment; change only how it is judged. Randomising which qualifying units are served first, or staggering rollout so a comparably extreme slice waits a period, preserves the targeting exactly while creating an honest untreated comparison.
- What do you do when leadership refuses any untreated group at all?Fall back to the weaker but still honest options: a multi-period baseline instead of the selection period, an outcome measured on a different instrument from the one used to select, and an explicit rebound figure reported as the null. Also require the gain to hold over several periods, since the artifact is a single return to the group's own level rather than a trend.
saying these in an interview costs you the question
- Proposes abandoning the targeting instead of the evaluation
- Says a holdout is impossible so before-and-after is acceptable
- Treats a large rebound as proof the programme worked
- Blames individual analysts rather than the standing selection rule
- Names the phenomenon without proposing a process change