skip to content

How do you check covariate balance after propensity score matching?

level: middleimportance: must knowfreq 82%

answer

  1. scale-free, not a hypothesis test
  2. before and after, side by side
  3. one row per covariate, dots
  4. the 0.1 rule of thumb
  5. means alone miss spread and interactions

basics

~20 s

Compare standardized mean differences for each covariate before and after matching, with an absolute value under about 0.1 as the usual bar, and read them off a love plot. Check variance ratios and interactions too, not just means.

solid answer

~50 s

The workhorse diagnostic is the standardized mean difference: the gap between the treated and matched control means of a covariate, divided by a standard deviation held fixed from the original sample. Because it is scale-free it can be compared across covariates and, unlike a p-value, it does not improve just because matching shrank the sample. The convention is that absolute SMDs below `0.1` are acceptable. A love plot puts one row per covariate with the before-matching and after-matching values side by side, so residual imbalance is impossible to miss. Means alone are not enough: also compare variance ratios, look at the score distributions in each group, and check a few interactions and squared terms. If balance fails you go back and change the score or the matching design, which is legitimate precisely because you do it blind to outcomes.

go deeper

for a junior

Know that matching must be followed by a balance check, and that the standard summary is a standardized mean difference per covariate rather than a table of p-values.

for a middle

Explain the mechanics: the difference in means divided by a fixed standard deviation, compared before and after matching, with roughly 0.1 as the acceptable bar and a love plot as the display.

for a senior

Demonstrate that you go past means, checking variance ratios, interactions and score overlap, and that you can defend iterating on the design as legitimate because outcomes stay sealed.

for a principal

Own the standard the organisation holds itself to: what balance evidence must appear in every causal write-up, and the rule that design and outcome analysis are separated so specification search cannot chase a result.

## Why balance is the deliverable of the design stage Matching does not prove anything by itself. It is a procedure that produces a matched sample, and the claim that matched sample supports is: *for the covariates I measured, the treated and control groups are now comparable*. Balance checking is how you substantiate that claim, and it is the only output of the design stage. Until balance is acceptable, there is no analysis to run. ## The standardized mean difference For a covariate `X`, the standardized mean difference is SMD = (mean_treated - mean_control) / s where `s` is a standard deviation used as the yardstick. The usual convention is to fix `s` from the original, pre-matching sample — commonly the treated group's standard deviation, or a pooled one — and to keep the same `s` for the before and after numbers. Holding the denominator fixed matters: if you recompute the spread inside the matched sample, a change in the SMD can reflect a change in variance rather than a genuine improvement in balance. For a binary covariate the same idea applies to the difference in proportions, standardized by the corresponding spread, so categorical variables sit on the same scale as continuous ones. Two properties make the SMD the right tool: - **It is scale-free.** A difference of 3 years of tenure and a difference of $400 of monthly spend become comparable numbers, so a single threshold can apply across a covariate table. - **It does not depend on sample size.** This is the decisive property, and the reason it beats hypothesis testing (below). The rule of thumb is that an absolute SMD below `0.1` counts as good balance and anything above `0.2` is a problem worth fixing. These are conventions, not theorems. A covariate that is strongly related to the outcome deserves tighter balance than one that barely matters. ## The love plot A love plot is a dot plot with one row per covariate and the absolute SMD on the horizontal axis, showing the value before matching and the value after, usually with a reference line at 0.1. Read left to right it answers the only question that matters at a glance: did every covariate move toward zero, and did any move away? Covariates that were badly imbalanced and stayed badly imbalanced are the ones to argue about; a covariate that was already balanced and got slightly worse is usually noise. Plotting before and after together also guards against a subtle self-deception. An after-matching table alone looks reassuring in isolation; only the before column shows how much work the matching actually did, and whether the treated and control groups were ever far apart to begin with. ## Why not test for balance The instinct to run a two-sample test per covariate and declare balance when nothing is significant is common and wrong, for two reasons. First, the p-value mixes the size of the imbalance with the size of the sample. Matching discards units, sometimes most of them, so `n` falls and p-values rise mechanically. A covariate can show exactly the same standardized gap before and after matching and go from "significant" to "not significant" purely because the matched sample is smaller. The test rewards you for throwing data away. Second, the test answers the wrong question. Balance is a property of the specific sample you are about to analyse, not a hypothesis about some larger population from which it was drawn. Once you hold the matched sample in your hand, there is nothing to infer: the means either are close or they are not. ## Beyond means Equal means do not imply equal distributions. A thorough check adds: - **Variance ratios** of treated to control for continuous covariates, with values near 1 desirable; ratios far from 1 mean one group is much more spread out even though the centres agree. - **Interactions and squared terms**, since a score can balance every covariate marginally while leaving the joint distribution skewed. - **Distributional comparisons** of the propensity score itself between groups, which shows both remaining imbalance and how much the two groups overlap at all. ## Iterating without cheating If balance is unacceptable you change something — add an interaction to the score, tighten the matching, alter the covariate set — and check again. This iteration is not p-hacking, and an interviewer may probe whether you know why. The design stage is conducted blind to outcomes: no outcome variable has been touched, so no amount of searching can be steered toward a preferred effect estimate. That protection evaporates the moment you start comparing candidate specifications by the treatment effect each one produces. Fix the design, freeze it, then look at the outcome once.

  • Why not just run a significance test per covariate to confirm balance?
    Because the p-value confounds imbalance with sample size. Matching discards units, so n falls and p-values rise even when the standardized gap is unchanged; the test can pass simply because the matched sample got smaller. Balance is also a property of the sample you are about to analyse, not a hypothesis about a wider population, so testing answers the wrong question altogether.
  • Every covariate now has an SMD of 0.04. Is the design credible?
    It is necessary, not sufficient. Balance speaks only to the covariates you measured and included, and mean-level balance can still hide differences in spread or in interactions. Report which covariates were balanced and at what values, and be explicit that the causal claim still rests on the untestable assumption that nothing important was left out.
  • Is it cheating to re-specify the score until balance looks acceptable?
    No, provided outcomes stay hidden. Iterating on the score and the matching design is the design stage of an observational study and it is judged on balance alone; since no outcome has been touched, the search cannot be steered toward a result. It stops being defensible the moment you start comparing specifications by the effect estimate each produces.

saying these in an interview costs you the question

  • Declares balance because no covariate test was significant
  • Reports only after-matching numbers with no before column
  • Recomputes the standardizing spread inside the matched sample
  • Checks means only, ignoring variance and interactions
  • Treats 0.1 as a hard law rather than a convention

context