skip to content

Propensity matching dropped your enterprise accounts for lack of common support - what do you report?

level: principalimportance: nice to knowfreq 29%

answer

  1. the population quietly changed
  2. count and profile who was dropped
  3. no untreated counterpart exists
  4. the target shifts away from all treated
  5. extrapolation hides what trimming shows

basics

~20 s

Report that the estimate now covers only the treated units that had comparable controls, not all of them. Quantify how many were dropped and how they differ, and state plainly that this design cannot answer the enterprise question.

solid answer

~50 s

Dropping units outside common support is honest, but it silently changes what is being estimated: the effect on all treated accounts becomes the effect on the overlap subpopulation. Enterprise accounts with propensity near `0.99` have essentially no untreated counterpart, so no adjustment method can recover their effect from this data, because the information is not there. My write-up names the retained population explicitly, gives the count and the profile of the discarded treated units, and reports what share of the business they represent, since losing 8% of accounts that carry 60% of revenue is a different result from losing 8% of the tail. Then I say what would answer the original question, such as a staged or randomised rollout inside that segment, or a comparison pool that actually contains enterprise accounts. What I do not do is widen the match until everyone keeps a partner.

go deeper

for a junior

Know that matching can only compare units that have counterparts, so treated units with no similar control are dropped and the results no longer describe them.

for a middle

Explain what common support means and how you would show it, comparing the score distributions of the two groups, and be able to state how many treated units fell outside the overlapping region.

for a senior

Show that you quantify and communicate the loss: who was discarded, what share of the outcome they carry, and how the reported conclusion would have to be worded differently because of it.

for a principal

Own the decision of whether a trimmed study should run at all, and be ready to tell stakeholders that the segment they care about needs a different design rather than a weakened version of this one.

## What common support means Common support, or overlap, is the region of the propensity score range where both treated and untreated units actually exist. It is the practical face of the positivity condition that any adjustment method needs: for the comparison to mean anything, a unit like this one must have been able to end up in either arm. When enterprise accounts have estimated scores clustered near 0.99, that condition has failed for them empirically. Essentially every account of that kind was treated. There is no untreated enterprise account to serve as a counterfactual, and no amount of statistical machinery conjures one. Matching therefore leaves those treated units unmatched, and they fall out of the analysis. ## The estimand changes, and it changes quietly This is the part that gets missed. Before trimming, the target was the average effect on the treated: every treated account, enterprise included. After trimming, the number you can defend is the average effect on the treated units that lie inside the region of overlap. Those are different quantities about different populations, and only the second one is supported by the data. Nothing in the pipeline announces the switch. The estimate comes out with a tidy confidence interval, the balance table looks excellent — often *better* than before, because the hardest units were removed — and the write-up says "effect of the programme on treated accounts". That sentence is now false, and nobody reading the report can tell. So the first duty is naming. The results section should say which population the number describes, in the language the business uses: "among mid-market accounts with a comparable untreated counterpart" rather than "among treated accounts". ## Quantify the loss in the units the audience cares about Counting discarded units is necessary but not sufficient, because accounts are not interchangeable. Three numbers belong in the report: 1. **How many treated units were discarded**, as a count and a share of the treated group. 2. **How they differ** from the retained units on the covariates that drive the outcome — size, tenure, region, engagement. 3. **What share of the outcome they carry.** Losing 8% of accounts sounds negligible until you note that those accounts represent 60% of revenue. Then the study covers a minority of what the decision is about. That third number converts a methodological footnote into a business fact, and it is what separates a principal-level answer from a competent one. ## The judgment call With those numbers in hand, there are three honest paths. **Report the trimmed estimate, clearly labelled.** Correct when the retained population is one somebody will act on. If the rollout decision concerns mid-market accounts, an effect estimated on comparable mid-market accounts is exactly the right deliverable. **Declare the question unanswerable with this data, and say what would answer it.** Correct when the discarded segment *is* the question. A staged or randomised rollout within the enterprise segment creates the variation that history did not; so does finding a comparison pool that actually contains untreated enterprise accounts, such as another market or an earlier period before the programme reached that tier. Saying this early is far cheaper than saying it after three weeks of analysis. **Split the report.** Estimate where support exists, and state explicitly that the enterprise segment is out of scope with a named plan to cover it. This is usually the version stakeholders can act on. ## What not to do The tempting move is to relax the matching until nothing is discarded — widen the distance allowed, or fall back to a model that extrapolates a prediction into a region where no untreated units were ever observed. Both replace a visible, countable restriction with an invisible assumption. A trimmed sample tells you exactly what you do not know and how much of it there is. A forced match, or a model asked to predict outside the range of its data, produces a number with no warning label attached, and its error does not shrink as the sample grows. The second thing not to do is decide the trimming rule after seeing the outcomes. The overlap rule belongs in the analysis plan: how the region is defined, what happens to units outside it, and what fraction of treated units you are prepared to lose before the study is declared infeasible. Fixed in advance, it is a design decision. Chosen afterwards, it is a knob that gets turned until the estimate looks acceptable. ## The organisational point Teams that run many observational studies benefit from making this a standing requirement rather than a judgment each analyst re-makes: every causal write-up names its population, reports discarded units with their share of the outcome, and states the overlap rule that was fixed before outcomes were examined. That single convention prevents most of the quiet estimand drift that makes observational results untrustworthy over time.

  • Why not widen the caliper so the enterprise accounts stay in the analysis?
    Because their matches would not be comparable units. Pairing an account with a propensity near 0.99 to a control near 0.4 does not supply the missing counterfactual; it borrows one from a different kind of account and buries the assumption inside a pair nobody will inspect. Trimming makes the limitation countable and reportable, which is the version a decision-maker can actually act on.
  • How do you decide whether the trimmed population is still worth studying?
    Ask whether anyone would act on an effect estimated for that population. If the retained accounts are the ones the rollout decision concerns, the trimmed estimate is useful and should be labelled as theirs. If the decision is about the discarded segment, the study answers a question nobody asked, and the right move is to say so early rather than deliver a precise irrelevance.
  • What should the analysis plan say about common support before any outcomes are seen?
    It should fix the overlap rule in advance: how the region is defined, what happens to units outside it, and what fraction of treated units you are willing to lose before declaring the study infeasible. Setting that at the design stage keeps the trimming rule from becoming a knob turned until the estimate looks acceptable.

saying these in an interview costs you the question

  • Calls the trimmed result the effect on all treated units
  • Discards non-overlapping units without counting or describing them
  • Widens the matching rule until nobody is discarded
  • Treats absent overlap as a modelling problem to be solved
  • Chooses the trimming rule after seeing the outcome estimates

context