skip to content

In Lord's paradox, why do a gain-score analyst and a baseline-adjusted analyst disagree?

level: principalimportance: nice to knowfreq 22%

answer

  1. two correct analyses, opposite conclusions
  2. change over time versus same starting value
  3. each assumes a different counterfactual
  4. baseline differs between the groups
  5. randomisation would make them agree

basics

~20 s

They estimate different quantities under different untestable assumptions. One treats each person's own starting value as the fair baseline; the other compares people who started alike. With groups that differ at baseline, those comparisons can point opposite ways.

solid answer

~50 s

Lord's paradox is the continuous-outcome cousin of Simpson's paradox. In Lord's setup a university weighs students at the start and end of an academic year. Neither group's mean weight changes over the year. The first analyst compares mean weight change, finds it is zero for both groups, and concludes there was no differential effect of the dining halls. The second analyst models final weight while holding initial weight fixed and finds that among students who started at the same weight, one group ended heavier — a real difference. Both computations are correct. They differ because each encodes a different counterfactual: "a student unaffected would have kept their own starting weight" versus "two students with the same starting weight would have ended alike". With groups that differ systematically at baseline, and a grouping nobody assigned, no statistic adjudicates between those assumptions. You choose the estimand and defend it.

go deeper

for a junior

It is enough to recognise that comparing change over time and comparing people who started at the same value are different comparisons that need not agree.

for a middle

Be able to state each analysis precisely — mean of the differences versus final value holding the starting value fixed — and say that both are correct computations on the same table.

for a senior

Name the counterfactual each analysis assumes, connect the divergence to baseline imbalance in a non-randomised comparison, and explain why randomisation makes the two agree.

for a principal

Own the process: estimand fixed and assumptions written down before results are seen, both analyses reported with their assumptions stated, and a standing rule that the observed result never selects the method.

## The setup Lord described a university that records every student's weight in September and again the following June, for two groups of students who eat in different dining halls. Two statisticians analyse the same table. **Analyst 1** computes each student's weight change, June minus September, and averages it within each group. In Lord's telling, neither group's mean weight changed over the year: the average gain is zero for both. The conclusion is that there is no evidence of a differential effect of the dining halls on weight. **Analyst 2** models final weight as a function of initial weight and group membership — comparing students who arrived at the same September weight. On that comparison one group finishes systematically heavier than the other, and the conclusion is that there *is* a difference. Neither analyst has made an error. This is Lord's paradox. ## Why the two answers differ The groups differ at baseline: their September weight distributions are not the same. Every analysis of change has to supply, implicitly, an answer to the question *what would this student's June weight have been in the absence of a group difference?* The two analysts answer it differently. - **Change scores** assume that, absent an effect, a student's expected June weight equals their September weight. The baseline is the person themselves. - **Adjusting for baseline** assumes that, absent an effect, two students who arrived at the same September weight would have the same expected June weight regardless of group. The baseline is other people who started alike. Those are different assumptions about an unobserved counterfactual, and they are not nested, not reconcilable, and not checkable from the data. When the groups' baseline distributions overlap poorly, the two assumptions can and do disagree in sign. Nothing in the table distinguishes them, which is precisely why the paradox is instructive rather than merely annoying. ## The relationship to Simpson's paradox Same skeleton, continuous covariate. In Simpson's paradox a comparison reverses depending on whether you condition on a categorical variable; in Lord's paradox it reverses depending on whether you condition on a continuous baseline measurement. In both cases the arithmetic is unambiguous and the disagreement lives entirely in the choice of what to hold fixed — a choice that expresses a belief about how the data were generated. ## The extra difficulty: the grouping was never assigned There is a second problem in Lord's example that a lead should name explicitly. The groups were not created by an intervention anyone controlled. Asking what a student's weight *would have been* in the other group requires imagining a manipulation that was never defined — and when the grouping is an attribute rather than an action, the counterfactual may not be well defined at all. Before arguing about which analysis is better, a principal-level answer asks what intervention the question refers to. Many "paradoxes" of this shape dissolve once someone insists on stating the intervention. ## When the two analyses agree If group membership had been randomly assigned, baseline weight would be balanced between the groups in expectation. Both analyses then target the same quantity and differ only in precision, with the baseline-adjusted version typically the more precise because it removes variance explained by the starting value. The disagreement is a symptom of baseline imbalance in a non-randomised comparison, not a defect of either method. That is worth saying out loud, because it locates the fix in the design rather than in the modelling. ## What a lead actually does about it 1. **Fix the estimand before the analysis.** Write down, in words, the quantity of interest and the intervention it refers to, and get agreement on it while the result is still unknown. 2. **State the identifying assumption in plain language.** "We assume two students who arrived at the same weight would have ended at the same weight had they eaten in the same hall" is a sentence stakeholders can argue with. A model specification is not. 3. **Run both and treat the gap as information.** If change scores and baseline adjustment disagree materially, that is a finding about baseline imbalance, and it belongs in the write-up rather than being resolved silently by whoever holds the keyboard. 4. **Do not let the result choose the method.** Pre-specify. The single most damaging version of this paradox is an analyst who tries both and reports the one matching the expected story. 5. **Push the fix upstream.** Where the comparison matters and assignment is feasible, randomise or otherwise balance the baseline; where it is not, be explicit that the conclusion rests on an assumption the data cannot test. ## The signal in an interview Weak answers pick a side — "you should always adjust for baseline" or "change scores are unbiased". Strong answers say that both estimates are correct arithmetic, name the counterfactual each assumes, observe that baseline imbalance in a non-assigned grouping is what makes them diverge, and describe the process by which the choice gets made and recorded before anyone sees the result.

  • Under what condition do the two analyses agree?
    When the groups are balanced at baseline, which random assignment delivers in expectation. Both approaches then target the same quantity and differ only in precision, with baseline adjustment usually tighter because it soaks up variance explained by the starting value. Divergence is therefore a symptom of baseline imbalance in a non-randomised comparison.
  • How is this related to Simpson's paradox?
    It is the same structure with a continuous covariate instead of a categorical one. In both, the conclusion flips depending on what you hold fixed, both computations are arithmetically correct, and the disagreement is entirely about which comparison has the causal meaning you want. Neither can be settled by inspecting the numbers.
  • As the lead, what do you do when both analyses are defensible?
    Fix the estimand and the identifying assumption in writing before results are seen, then report both estimates with the assumption each relies on stated in plain language. Treat a material gap as a finding about baseline imbalance rather than something to resolve quietly. Never let the observed result select the method.

saying these in an interview costs you the question

  • Says one analyst simply used the wrong method
  • Claims a larger sample would resolve the disagreement
  • Insists adjusting for baseline is always the safer choice
  • Treats an unassigned grouping as a treatment without comment
  • Tries both analyses and reports whichever matches expectations

context