In propensity score matching, when should you match with replacement rather than without?
answer
- who gets the good partners first
- greedy matching depends on order
- a thin pool changes the answer
- one control standing in many times
- effective sample size, not row count
basics
~20 sMatch with replacement when the control pool is thin or overlaps poorly, because every treated unit then gets its closest available control instead of the leftovers. The price is that heavily reused controls shrink the effective sample size and widen uncertainty.
solid answer
~50 sWithout replacement each control is used at most once, so a greedy pass through 200 treated units against 240 plausible controls gives the early treated units good partners and the last ones whatever remains. The estimate then depends on the order the treated units were processed and absorbs bias from those poor late pairs. With replacement, every treated unit takes its nearest control even if that control has already served, which removes the order dependence and cuts bias but concentrates the comparison on a few units: if one control is reused thirty times, the control side carries far less information than the row count suggests. That has to enter the analysis, by weighting each control by how often it was used and treating the matched structure rather than the raw pair count as the sample. Thin or poorly overlapping pools favour replacement; a large, comparable pool favours without.
go deeper
Know the distinction itself: without replacement a control can be matched only once, while with replacement the same control may serve several treated units.
Explain the tradeoff mechanically, that replacement gives every treated unit its closest control and removes order dependence, at the cost of leaning on a small number of reused controls.
Demonstrate that you handle the consequences in the analysis: reuse weights, an effective sample size well below the pair count, and a check on how many distinct controls actually carry the comparison.
Own the judgment of whether a control pool this thin can support a credible comparison at all, and be ready to argue for building a better comparison group or running an experiment instead.
## The two regimes Matching without replacement removes each control from the pool once it has been used. Every control appears in at most one matched set, so the matched data look like a clean collection of independent pairs. Matching with replacement leaves each control in the pool permanently. A single control can be the nearest neighbour of many treated units and appear in many matched sets. The choice looks administrative. It is not: it decides how much bias and how much variance the design carries, and it decides what the analysis has to do afterwards. ## What goes wrong without replacement when the pool is thin Consider 200 treated units and a comparison pool of 240 units that are plausible at all. That is a ratio of 1.2 controls per treated unit — thin. A greedy algorithm walks the treated units in some order and assigns each the closest remaining control. The first treated unit chooses from 240 candidates and gets an excellent partner. The two-hundredth chooses from 41 leftovers, most of which the earlier treated units already rejected, and gets whatever is left. Two consequences follow. **Order dependence.** Process the treated units in a different sequence — sorted by score, sorted by ID, shuffled — and you get a different matched sample and a different estimate. A result that depends on row order is not a property of the data. Some implementations mitigate this by matching hardest-to-match units first, or by solving for the assignment that minimises total distance across all pairs rather than greedily, which removes the order effect but not the underlying scarcity. **Systematic bias in the tail.** The badly matched pairs are not randomly scattered; they concentrate among the treated units whose scores are furthest from the bulk of the control pool. Those are exactly the units where a bad pair does the most damage. ## What replacement buys, and what it costs With replacement, each treated unit gets its genuinely nearest control. Order stops mattering entirely, because no unit is ever consumed. Bias from poor pairings drops, often sharply, in exactly the thin-pool case where without-replacement matching struggles. The cost is concentration. If one control at the top of the score range is the nearest neighbour of thirty treated units, thirty matched rows rest on one person's outcome. Those rows are not independent: they share the same control outcome, so a fluke in that single unit propagates into thirty comparisons. The information content of the control side is roughly the number of *distinct* controls used, not the number of rows. This has a concrete analytic consequence. Each control must enter the comparison with a weight equal to the number of times it was matched, and the uncertainty must be computed from the matched structure rather than from a naive pair count. An analysis that reports the standard error as if 200 independent pairs existed, when 200 pairs draw on 60 distinct controls, will overstate precision — sometimes badly. ## Diagnostics worth running - **Distinct controls used.** Report it alongside the number of matched pairs. A large gap is the headline fact about the design. - **Reuse distribution.** How many times is the most-used control matched? If the top few controls account for a large share of matched rows, the comparison is effectively resting on a handful of units. - **Match distance distribution.** Under either regime, look at how far apart the matched pairs actually are. Without replacement this exposes the leftovers problem; with replacement it confirms the pairs really are close. ## Ratios and other variants A related dial is how many controls to match per treated unit. Matching 3 controls to each treated unit uses more of the pool and lowers variance, but the second and third neighbours are by construction worse fits than the first, so bias rises with the ratio. With a thin pool the extra neighbours are usually poor, which is precisely when the variance saving is not worth the bias. Variable-ratio and full matching schemes attack this by letting the number of controls differ per treated unit according to how many good ones are locally available. ## How to answer the choice State a rule and its condition rather than a preference: - Control pool large relative to the treated group, good overlap, plenty of near neighbours everywhere: match **without** replacement. The design is simpler, sets are independent, no reuse weights are needed, and the leftovers problem never bites. - Thin pool, or treated units sitting in a region of the score distribution where controls are sparse: match **with** replacement, and carry the reuse into the analysis. And add the check that distinguishes a practitioner: run it both ways. If the estimate is similar under both regimes, the choice was not load-bearing and you can report the simpler design. If it differs materially, that difference is itself a finding about how fragile the comparison is.
- How does reusing one control for thirty treated units affect the uncertainty?It shrinks the effective sample size. Thirty pairs sharing one control carry roughly one control's worth of independent information, not thirty, so an analysis that treats the pairs as independent reports a standard error that is too small. The reuse must enter as a weight on each control, with uncertainty computed from the matched structure rather than the raw pair count.
- Would matching three controls per treated unit help when the pool is thin?Only if the extra controls are genuinely close. A higher ratio uses more of the pool and lowers variance, but each additional match is by construction a worse fit than the first, so bias grows with the ratio. In a thin pool the second and third neighbours are usually poor, which is exactly when the extra matches cost more than the variance saving is worth.
- Does matching without replacement have advantages worth keeping?Yes, simplicity and independence. Each control appears once, so matched sets share no units, no reuse weights are needed, and the design is easy to describe and audit. When the control pool is large relative to the treated group and overlap is good, the leftovers problem never bites and without replacement is the cleaner choice.
saying these in an interview costs you the question
- Treats matched rows as independent when controls are reused
- Ignores that greedy matching depends on processing order
- Assumes matching with replacement is always the safer default
- Reports the sample size as the number of pairs after heavy reuse
- Believes more controls per treated unit always improves the estimate