skip to content

Your logistic uptake model's scores will be reused as propensity scores - what changes in how you fit it?

level: seniorimportance: nice to knowfreq 31%

answer

  1. the label is treatment, not outcome
  2. keep confounders regardless of lift
  3. nothing measured after the send
  4. resampling moves the whole score level
  5. a near-perfect separation is bad news

basics

~20 s

The label becomes the treatment indicator and the probability level matters more than the ranking. Keep confounders even when they add no accuracy, exclude anything measured after treatment, and do not rebalance the classes or over-penalise the weights.

solid answer

~40 s

A propensity score is the fitted probability of having been treated - here, of having received the email campaign - given pre-treatment covariates, and it is exactly what a sigmoid classifier trained by cross-entropy produces. Three things change. First, the label is the treatment indicator, not the business outcome. Second, variable selection follows the causal story rather than predictive lift: keep covariates that plausibly drive both treatment and outcome even when their weight is small, and drop anything recorded after the campaign went out, since those are consequences of treatment. Third, the *level* of the scores matters, not just their ordering, so avoid oversampling, class weighting or heavy regularisation, all of which shift the fitted probabilities away from the observed treatment rate. And a near-perfect fit is bad news, not good.

go deeper

for a junior

Know what the score is: the fitted probability of having been treated given covariates measured beforehand, produced by the same sigmoid-plus-cross-entropy fit used for any binary target.

for a middle

Explain why the probability level matters rather than the ranking, and name what shifts that level: resampling, class weighting and a strong penalty on the weights all move the fitted probabilities away from the observed treatment rate.

for a senior

Show the judgment: choose covariates by their role in the assignment story rather than by predictive lift, exclude everything recorded after treatment, and read a near-perfect fit as thin overlap rather than as success.

for a principal

Own the standard your organisation holds this to. Decide what must be documented before scores from a model built for one purpose may be reused for another, and who signs off that the covariate list is defensible.

## What is being asked for A propensity score is the probability that a unit received the treatment, given its pre-treatment covariates. If you fit a sigmoid classifier whose label is "was sent the email campaign" and whose inputs are covariates measured before the send, the fitted probabilities *are* propensity scores. Cross-entropy is the right objective for this because its population minimiser is the true conditional event rate: fitting by negative log-likelihood pushes the predicted probability towards the actual share of similar units that were treated, which is precisely the quantity the downstream comparison needs. What changes is not the loss but everything around it, because the model's job has changed. A campaign-uptake model built for targeting is judged on whether it ranks people well. A propensity model is judged on whether its numbers describe assignment. ## 1. The label is the treatment, not the outcome The first mistake is fitting the wrong target. If the earlier model predicted who would *respond* to the campaign, that is an outcome model and its scores are not propensity scores. The propensity model's label is who *received* the campaign. If the two were built on the same feature set for different labels, the reuse is invalid and you refit. ## 2. Variable selection follows the causal story For pure prediction you keep whatever lifts accuracy. Here the selection rule is different: - **Keep confounders** - covariates that plausibly influence both who was emailed and whether they would have converted anyway - even if their fitted weight is small and dropping them barely changes discrimination. Leaving one out is the thing the whole exercise exists to prevent. - **Exclude post-treatment variables.** Anything recorded after the campaign was sent - opens, clicks, subsequent visits - is a consequence of treatment. Conditioning on it removes exactly the variation you are trying to measure. - **Be wary of pure instruments** - variables that drive treatment assignment but have no path to the outcome, such as a send-queue quirk or a mailing-list vendor batch. They add discrimination and can amplify bias from whatever confounding remains, because they push scores towards the extremes without removing confounding. This is why "my feature-selection step dropped it because it added nothing to accuracy" is not an acceptable justification on a propensity model. ## 3. The level of the probability matters, not just the order For a targeting model you can oversample, apply class weights, or lean on a strong penalty, and the ranking survives. A propensity score is consumed as a number, so all three are harmful: - **Resampling or class weighting** changes the base rate the model is fitting to. If 5% of your list was emailed and you resample to 50/50, the fitted probabilities describe the resampled world, not yours. - **Heavy regularisation** shrinks weights toward zero, which drags every score towards the overall treatment rate and can zero out a confounder entirely. Mild regularisation for numerical stability is defensible; aggressive selection-by-penalty is not. - **Rounding, clipping or bucketing the score** for convenience destroys the fine distinctions the downstream step depends on. ## 4. A near-perfect fit is a warning On a targeting model, separating the classes almost perfectly is a triumph. On a propensity model it means the treated and untreated groups barely overlap: there are regions of covariate space where essentially everyone was emailed and regions where essentially no one was. Scores piling up near 0 and 1 mean there are few comparable units to draw on, and any estimate that follows rests on extrapolation. Look at the two score distributions side by side; heavy separation is a finding to report, not a metric to celebrate. The converse is also informative: if the model can barely beat the base rate, assignment looks close to random on the covariates you have, which is a comfortable position - although it can also mean your covariates simply do not capture how assignment was actually made. ## 5. What you check, and what you hand on Before handing the scores on, confirm the mechanics: the label is treatment, every input predates the send, the average predicted probability is close to the observed treated share, and the score distributions for the two groups overlap over most of their range. Record which covariates went in and why, because the credibility of everything downstream rests on that list rather than on any accuracy number. How the scores are then used to compare treated and untreated units is a separate step with its own methodology and its own diagnostics. ## Common mistakes - Reporting discrimination as the model's quality bar, when the quantity being reused is the probability itself. - Rebalancing classes out of habit and shifting every score. - Including engagement features that only exist because the campaign was sent. - Treating a very strong fit as validation rather than as evidence of thin overlap.

  • Should a propensity model be regularised at all?
    Mild regularisation is fine and often helps numerically, especially with many correlated covariates or near-separation. What you cannot do is use a strong penalty as a variable-selection mechanism: it shrinks every score towards the overall treatment rate and can zero out a confounder whose inclusion is the entire point. If you regularise, keep it light and check that the average predicted probability still tracks the observed treated share.
  • Your propensity model achieves near-perfect discrimination on treatment assignment. What do you conclude?
    That treatment was close to deterministic given the covariates, so treated and untreated units barely overlap. There are few comparable pairs, and any comparison built on those scores leans on extrapolation into regions where one group is essentially absent. Report it as a limitation of the design, check whether a post-treatment variable leaked into the model, and consider narrowing the population to the region where both groups actually appear.
  • The campaign's engagement data - opens and clicks - is your strongest available feature. Can it go in?
    No. Opens and clicks only exist because the email was sent, so they are consequences of treatment rather than pre-treatment covariates. Including them conditions on the treatment's own effect and biases everything downstream. Only variables whose values were fixed before the send belong in the model, however much predictive power you give up.

saying these in an interview costs you the question

  • Judges the propensity model by its discrimination alone
  • Rebalances the treated and untreated classes before fitting
  • Includes features recorded after the treatment was delivered
  • Drops a confounder because it added no predictive lift
  • Treats near-perfect separation as a sign of a good model
  • Fits the outcome instead of the treatment indicator

context