In an IPW analysis one user has propensity score 0.01 and weight 100 — what do you do?
answer
- one row speaking for a hundred
- look at the weight distribution first
- nominal n versus effective sample size
- drop units or cap the weights
- bias for variance, threshold set in advance
basics
~20 sTreat it as a diagnostic, not a nuisance. Check the largest weights and the effective sample size, see how far the estimate moves without that unit, then decide between trimming extreme-score units and capping weights.
solid answer
~50 sA weight of 100 means one observation is speaking for a hundred, so the estimate and its variance can hinge on a single person. First I quantify it: the largest weights, each unit's share of the arm's total weight, and the effective sample size `(sum of w)^2 / (sum of w^2)` — if 5,000 rows give an effective 300, the analysis is far weaker than it looks. Then I read it substantively: a score of 0.01 says there is a covariate region where almost nobody gets treated, so the data carry little information there. Remedies in order: re-examine the treatment model for a covariate that nearly determines treatment, restrict to the overlap region (trimming), or truncate weights at, say, the 1st and 99th percentiles — trading a little bias for a large variance reduction. I report the rule and the estimate with and without it.
go deeper
Know that a weight of 100 means one observation counts as a hundred, and that huge weights come from propensity scores close to zero or one. Be ready to say why that makes an estimate fragile.
Explain the diagnostics you would compute: maximum weight, each unit's share of its arm's total weight, and effective sample size, plus the mechanics of trimming versus capping and why capping trades bias for variance.
Demonstrate that you have handled this in real work: read the extreme weight as evidence about the data, investigate the treatment model before patching its output, run a sensitivity sweep over thresholds, and report the estimate with and without the rule.
Own the framing decision. Argue whether the honest response is to narrow the estimand to the region with real overlap and tell stakeholders which users the answer no longer covers, rather than producing a fragile number for the whole population.
## Why one weight of 100 matters An inverse probability weight of 100 comes from a treated unit with an estimated propensity of 0.01: the model says a person with those covariates almost never gets treated, and yet this one did. In the reweighted pseudo-population that single observation is cloned a hundred times, so its outcome enters the treated-arm average with a hundred times the influence of a unit with score 1.0. If that person happens to have an unusual outcome, the estimated effect moves with them. This is the single most common way an IPW analysis goes wrong in practice, and it is a favourite senior interview probe because the right answer is a diagnostic procedure, not a formula. ## Step 1: quantify the problem Several cheap diagnostics belong in every weighted analysis: - **The weight distribution.** Look at the maximum, the 99th percentile and the ratio of max to median, per arm. A max/median ratio in the hundreds is a red alert. - **Share of total weight.** Compute each unit's weight divided by its arm's total weight. If the top unit holds 10% of an arm's weight, the arm is effectively a handful of people. - **Effective sample size.** The standard measure is `ESS = (sum of w)^2 / (sum of w^2)`, computed per arm. With equal weights it equals the number of units; the more unequal the weights, the smaller it gets. Reporting 'n = 5,000 but ESS = 300' communicates the real precision immediately. - **Leave-one-out sensitivity.** Re-estimate the effect with the top-weighted unit removed. If the sign or the practical conclusion flips, you do not have an estimate, you have that person's outcome. ## Step 2: read it substantively A propensity of 0.01 is a statement about the design of your data: in that covariate region there is almost no treated experience to learn from. Two very different causes produce it. First, **the region genuinely has almost no treated units**. Then no estimator can rescue you there; any method that appears to produce a confident answer is extrapolating. The honest move is to narrow the question to the region where both arms exist and say so. Second, **the treatment model is at fault**. A covariate that nearly determines treatment (a variable recorded only for treated users, a proxy for the assignment rule itself, an over-flexible model that separates the arms) manufactures extreme scores. Also check for a variable that should not be in the model at all — one that predicts treatment strongly but has no path to the outcome adds enormous weight variance while removing little bias. Fixing the model is preferable to patching its output. ## Step 3: choose a remedy, and be explicit about it - **Trimming (dropping units).** Discard units whose scores fall outside a range such as [0.01, 0.99], or drop the tails of the score distribution. This is the most defensible option because it changes the *question*: you are now estimating the effect in the region where both treatments occur, and you can describe who was excluded. It reduces variance a lot and removes the extrapolation, at the cost of a target population you must characterise. - **Truncation / weight capping.** Keep every unit but cap weights at, say, the 1st and 99th percentiles of the weight distribution. This is a pure bias-variance trade: the capped estimator is biased for the original estimand even when the model is right, but its variance can fall by an order of magnitude, and total error is often lower. Sweep the cap (99th, 97.5th, 95th) and show the estimate as a function of it — a stable plateau is reassuring, a steadily marching estimate is not. - **Stabilization.** Rescaling weights by the marginal treatment probability helps the estimator's behaviour but does not compress the relative spread in a point-treatment setting, so it is not on its own a fix for one dominant unit. - **Change estimand.** Targeting the effect among the treated, or an overlap-focused estimand that deliberately downweights the extremes, sidesteps the problem by asking a question the data can answer. ## Step 4: report honestly Whatever you do, the write-up should contain the weight diagnostics, the rule applied (with its threshold chosen before seeing the effect, not after), the estimate before and after, and the effective sample size. A trimming rule tuned until the effect looks good is a researcher-degrees-of-freedom problem dressed as hygiene. ## The one-line summary Extreme weights are the data telling you that part of your comparison rests on almost no evidence. Measure the concentration, find out whether it is the world or the model, then either narrow the question or accept a controlled amount of bias — and say which you did.
- How do you compute and interpret the effective sample size of a weighted arm?Use `(sum of w)^2 / (sum of w^2)` within the arm. With equal weights it returns the number of units; as weights become unequal it shrinks toward the number of units that genuinely carry the estimate. It is the right number to quote alongside n, because it tells the reader how much independent information the weighted comparison actually has.
- What is the difference between trimming units and truncating weights, and when do you prefer each?Trimming removes units whose scores lie outside a range, which changes the target population to the overlap region and must be described. Truncation keeps every unit but caps the weight, keeping the nominal population while introducing bias toward the capped values. Prefer trimming when the extremes represent a genuine lack of comparable units, truncation when you want to keep the original population and can defend a modest bias.
- Why can a strong predictor of treatment that has no effect on the outcome make an IPW analysis worse?Such a variable pushes scores toward 0 and 1, inflating weights and variance, while removing essentially no confounding bias because it has no path to the outcome. The result is a noisier estimate with unstable weights for no bias gain, so covariate selection for the treatment model should be driven by plausible confounding, not by predictive power.
saying these in an interview costs you the question
- Deletes the biggest weights after seeing the effect move
- Reports nominal n while a few units carry the estimate
- Says stabilizing weights fixes one dominant observation
- Treats an extreme weight as a data-entry error to clean
- Caps weights without reporting the rule or the threshold