skip to content

Why did observational studies of hormone replacement therapy overstate its heart benefit?

level: seniorimportance: should knowfreq 46%

answer

  1. who chooses the treatment, and why
  2. the treated group was healthier to begin with
  3. adjusted for what was recorded, not what mattered
  4. adherers to placebo also do better
  5. unmeasured confounding, not sampling error

basics

~20 s

Women who took hormone replacement therapy were healthier and more health-seeking than those who did not, in ways the studies never measured. Conditional exchangeability failed, so the adjusted comparison mixed the drug's effect with a healthy-user effect.

solid answer

~50 s

Large observational cohorts reported that women on hormone replacement therapy had less coronary heart disease, and the later randomized trial found no cardioprotective benefit, with harm in the combined estrogen-progestin arm. The gap is a textbook failure of conditional exchangeability from unmeasured confounding. Women who took the therapy were, on average, wealthier, better insured, more likely to be screened, more physically active and more adherent to medical advice. The studies adjusted for what they had, such as age, smoking, blood pressure and reported activity, but health-seeking behaviour itself was never recorded. The lesson is that a rich-looking covariate list is not evidence of exchangeability: what matters is whether the covariates capture the drivers of the treatment decision. Randomized trials have found that people who adhere to placebo have better outcomes than people who do not, which shows how large the adherer effect alone can be.

go deeper

for a junior

Know that people who choose a treatment can differ from those who do not in ways nobody wrote down, and that this can make a treatment look better than it is.

for a middle

Explain that bias from an omitted driver of treatment choice does not shrink with sample size, and that adjusting for correlates of that driver leaves residual confounding behind.

for a senior

Demonstrate the diagnostic habit: reconstruct who decides treatment and on what basis, judge whether the covariates measure that basis, and state the direction of the remaining bias.

for a principal

Own the call on whether an observational causal claim is publishable at all, and set the norm that the team escalates to an experiment when the key driver of treatment choice is unmeasurable.

## What happened For years, observational cohorts consistently found that women taking hormone replacement therapy had lower rates of coronary heart disease than women who did not. The finding was reproduced across studies, was biologically plausible, and shaped prescribing. When the question was finally settled by a large randomized trial, the cardioprotective effect was not there: the trial found no coronary benefit, and the combined estrogen-progestin arm was stopped early because the overall risk profile was unfavourable. This is the most instructive case in applied causal inference precisely because the observational studies were not sloppy. They were large, prospective, and adjusted for a long list of covariates. What failed was the identification assumption underneath the adjustment. ## The mechanism: healthy-user and healthy-adherer effects Treatment was not assigned; it was chosen, jointly by patients and their doctors. The women who chose it differed systematically: - **Healthy-user effect.** People who take a preventive medication are, on average, people who do other health-protective things: they exercise, they attend screenings, they have better access to care, they have more money and more education. - **Healthy-adherer effect.** Among those prescribed something, those who keep taking it are a further-selected group. Randomized trials have found that participants adhering to *placebo* have better outcomes than those who do not, which is unambiguous evidence that adherence marks something about the person rather than something about the pill. - **Prevalent-user selection.** A cohort that recruits women already on therapy has, by construction, excluded anyone who started and stopped because of an early adverse event, enriching the treated group with people who tolerated it well. Each of these makes the treated group look better regardless of the drug, which is exactly what conditional ignorability, `(Y(1), Y(0)) independent of T | X`, forbids. ## Why the adjustment did not save it A covariate set can be long and still be the wrong set. Age, smoking, blood pressure and body mass are outcomes and correlates of health-seeking behaviour, not measurements of it. Adjusting for them removes the part of the confounding they capture and leaves the rest, and residual confounding after adjustment looks identical, in the output, to a real effect. There is no statistic in the analysis that flags the problem, because the assumption concerns something absent from the data. Two warning signs were available in principle. First, the estimated effect implied benefit across many outcomes with no shared mechanism, which is the fingerprint of a generally healthier treated group rather than of a drug. Second, the treatment decision was known to be socioeconomically patterned, so the mechanism argument for exchangeability was weak from the start. ## What a strong candidate does with this - **Starts from the decision.** Who chooses this treatment, on what information, and does that information predict the outcome? If the answer is a disposition rather than a recorded variable, say so up front. - **Asks whether the covariates measure the driver or merely correlate with it.** A proxy measured with error leaves residual confounding proportional to the error. - **Prefers designs where assignment does not depend on the disposition.** An experiment where feasible; failing that, a source of variation in treatment that is plausibly unrelated to health-seeking behaviour. - **Restricts to new users.** Comparing people from the moment they start therapy avoids the survivor enrichment of a prevalent-user cohort. - **Names the direction of the residual bias.** Healthy-user selection biases a preventive treatment toward looking beneficial, so an observed protective effect should be discounted rather than taken at face value. - **Adjusts the language, not just the model.** If the assumption cannot be defended, report an association and say what would be needed to make it causal. ## The general lesson The failure was not statistical noise, and no sample size would have fixed it, because bias does not shrink with n. It was a failure of the identification assumption: the covariate set did not contain the driver of treatment choice. When an observational result and a randomized result disagree, the default explanation is that the observational study's exchangeability assumption was wrong, not that the trial was unrepresentative, though population differences are worth checking as a secondary question.

  • What would the covariate set have needed to contain for the observational estimate to hold?
    Everything driving both the decision to take the therapy and later cardiac risk: health-seeking behaviour, screening frequency, diet, physical activity, access to care, willingness to adhere, and the prescribing physician's judgment. Most of these are dispositions rather than fields in a cohort or claims dataset, which is why adding more of the usual clinical covariates would not have closed the gap.
  • If you cannot measure health-seeking behaviour, what are your options?
    Change the design so the assumption is plausible, through an experiment or a source of variation in treatment unrelated to that disposition; narrow the question to one the data can support; or report the association explicitly as an association. Piling on more routinely collected covariates gives false comfort, since they correlate with the missing driver rather than measure it.
  • Why does restricting to new users help in a study like this?
    A cohort that enrols people already on long-term therapy has silently excluded anyone who started and stopped because of an early adverse event, so the treated group is enriched with people who tolerated the drug. Following people from the moment they initiate therapy puts the observational comparison on the same footing as a trial's randomization point.

saying these in an interview costs you the question

  • Blames the trial's sample or duration rather than confounding
  • Says adding more covariates would have fixed it
  • Treats a very large cohort as evidence against a randomized result
  • Assumes statistical adjustment removes unmeasured confounding
  • Calls the discrepancy random error rather than systematic bias

context