skip to content

For an email campaign, do customer tenure, Black-Friday timing and clicking the email belong in the adjustment set for purchases?

level: seniorimportance: must knowfreq 58%

answer

  1. check each candidate against both conditions
  2. two shared causes, one consequence of treatment
  3. one candidate is campaign-level, not customer-level
  4. you cannot click an email you never received
  5. targeting inputs are the confounders you actually need

basics

~20 s

Tenure and Black-Friday timing each influence who was emailed and who buys, so both sit on backdoor paths and belong in the adjustment set. Clicking is caused by the campaign, so the criterion excludes it.

solid answer

~40 s

Decide node by node against the two conditions of the backdoor criterion. **Customer tenure** drives both targeting and baseline purchasing, so it sits on the fork `Send <- Tenure -> Purchase` — a backdoor path that tenure closes. **Black-Friday timing** has the same shape at campaign level: sends spike in that window and so does buying, giving `Send <- BlackFriday -> Purchase`, so it goes in too. **Clicked the email** fails condition (1) immediately: it is a descendant of the treatment, since nobody clicks an email that was never sent. Conditioning on it compares treated clickers with units that had no opportunity to click, and removes part of the effect being measured. The adjustment set is `{tenure, Black-Friday timing}`. Clicks are a fine engagement diagnostic — just not a control.

go deeper

for a junior

Recall the basic split: variables that came before the email and influence both who got it and who buys are controls, and anything the email itself caused is not.

for a middle

Walk each candidate against the two conditions out loud and name the path shape it creates. Explain why a variable that predicts the outcome strongly can still be inadmissible.

for a senior

Show that you start from the actual assignment mechanism: ask what the targeting rule used, name the unmeasured targeting score as the live threat, and quantify which direction it biases the estimate.

for a principal

Own the call between an observational estimate and a redesign. Decide when to spend a holdout on the next campaign rather than accept an unidentified effect, and set the standard for how the team documents adjustment choices.

## The setup Treatment T is "a campaign email was sent to this customer". Outcome Y is "the customer purchased within seven days". You did not randomise: an in-house targeting rule decided who got the email. You want the causal effect of sending on purchasing, and you have three candidate variables to control for. The right method is not intuition about "what predicts the outcome" — it is to draw the graph and check each node against the two conditions of the backdoor criterion: 1. No node in the adjustment set is a descendant of T. 2. The set blocks every backdoor path from T to Y — every path whose first edge points into T. ## Node 1: customer tenure How long someone has been a customer plausibly affects **whether they were emailed**: targeting rules favour established customers, or conversely exclude them from win-back campaigns. It also plausibly affects **whether they purchase**, independently of any email: established customers buy more often. That is the fork `Send <- Tenure -> Purchase`. It is a path from T to Y whose first edge points into T, so it is a backdoor path, and it is open because a fork transmits association unless its middle node is conditioned on. Leaving tenure out means part of the apparent "email effect" is really just the tendency of the sort of customer who gets emailed to buy anyway. Tenure is measured before the send and is not caused by it, so condition (1) is satisfied. **Tenure goes in.** ## Node 2: sent during Black Friday This is the same structure at a different level of the system. A shopping peak raises the volume of campaign sends *and* raises purchase propensity for reasons that have nothing to do with any particular email — the whole market is buying. So `Send <- BlackFriday -> Purchase` is a second open backdoor path with a shared cause in the middle. This one is easy to miss because it is a property of *when* rather than *who*, and it often lives in a different table from the customer attributes. It is nonetheless the single most dangerous confounder in campaign analysis: a campaign that fired during a peak will look spectacular until the timing is accounted for. It precedes the send in the causal order, so condition (1) is fine. **Black-Friday timing goes in.** ## Node 3: clicked the email Here the criterion answers before you have to think about mechanisms at all. A customer cannot click an email that was never sent, so `Send -> Clicked`. Clicking is a **descendant of the treatment**, and condition (1) bans descendants of T from the adjustment set. **Clicks stay out.** It is worth being able to say *why* the rule exists here, in the vocabulary of this criterion. Clicks lie on the route by which the treatment reaches the outcome — the email works partly by getting opened and clicked. The backdoor criterion is built to leave that route intact while closing the routes that carry no effect; conditioning on a node on the causal route closes what you are trying to measure. Practically, "clicked" is also undefined or structurally zero for every untreated unit, so comparing clickers to non-recipients is not a like-for-like contrast at all. The temptation is strong because clicks correlate beautifully with purchases. Predictive strength is not the criterion. A variable's admissibility is decided by where it sits in the graph relative to T, not by how much of Y it explains. ## The resulting analysis Adjustment set: `{tenure, Black-Friday timing}`. The identified quantity is `P(Purchase | do(Send)) = sum over z of P(Purchase | Send, Z=z) * P(Z=z)` with Z ranging over the joint levels of the two variables, averaged over the population distribution of those levels. ## What to say about the parts you cannot see A strong answer does not stop at the three variables handed to you. It states the remaining threat: if the targeting rule used a score built from browsing behaviour that also predicts purchasing, and that score is not in your data, there is an unmeasured node with arrows into both T and Y, and no set of measured variables closes that path. In campaign analysis this is the usual situation, because targeting is deliberately built to select likely buyers. The correct move is to recover the actual targeting inputs — they exist, they are logged somewhere, and they are exactly the variables the criterion asks for — rather than to add whatever features are convenient. If they cannot be recovered, say the effect is not identified and argue for a holdout group on the next campaign instead of an elaborate observational model. ## The interview signal This question separates people who apply a rule from people who reason about the mechanism. Weak answers control for clicks because it improves fit; medium answers get tenure but miss timing because it is not a customer-level column; strong answers walk each node against both conditions, name the unmeasured targeting score as the residual threat, and propose the holdout that would settle it.

  • The team argues that clicks should be controlled for because they strongly predict purchase. What is your response?
    Predictive strength is not the admissibility rule. A variable enters an adjustment set because of where it sits relative to the treatment in the graph, and clicking is caused by the send, so the criterion excludes it. Controlling for it also compares recipients who clicked against units that could not click at all, which is not a like-for-like contrast. Report clicks as a diagnostic instead.
  • The targeting model used a propensity score built from browsing data you cannot access. What does that do to identification?
    It introduces an unmeasured node with arrows into both the send decision and purchasing, so a backdoor path exists that no measured variable blocks and the effect is not identified by adjustment. Since targeting deliberately selects likely buyers, the bias is large and upward. Recover the targeting inputs from the campaign system, or run the next campaign with a randomised holdout.
  • Would you also adjust for the customer's number of purchases in the previous quarter?
    Yes if it precedes the send and plausibly influences both targeting and current purchasing, which is the usual case: it is then a common cause sitting on an open backdoor path. Check the timing carefully, though. If the window overlaps the campaign period, the variable is partly a consequence of the treatment and falls under the descendant ban instead.

saying these in an interview costs you the question

  • Controls for clicks because they predict purchases well
  • Misses campaign timing because it is not a customer column
  • Puts every available column into the adjustment set
  • Claims randomisation is unnecessary if enough controls exist
  • Ignores that the targeting rule itself defines the confounders

context