skip to content

How do you turn an exec's 'does the coupon work?' into a well-defined causal estimand?

level: principalimportance: should knowfreq 42%

answer

  1. the question is underspecified as asked
  2. name the intervention, not the behaviour
  3. say which population and which horizon
  4. let the decision pick the estimand
  5. watch for subgroups formed after treatment

basics

~10 s

Pin down four things before any method: the intervention (being offered the coupon versus redeeming it), the population, the outcome and its window, and the decision the number feeds. The estimand follows from those.

solid answer

~50 s

"Does it work?" is not yet a question. I would settle four things. The **intervention**: being *offered* the coupon, or redeeming it — those are different treatments, and only the offer is something we control. The **population**: every eligible customer, or only the slice we currently mail. The **outcome and window**: orders in 14 days, or margin in 90, since a coupon can lift one and sink the other. And the **decision**: launch to everyone, or keep or kill the existing targeted programme. Those map to different estimands — the ATE over the eligible base for a full launch, the ATT on the mailed slice for keep-or-kill. The tempting answer, "the effect on redeemers", is the trap: redemption happens after the coupon lands, so redeemers versus non-redeemers compares self-selected groups, not treatment arms. Naming the estimand also tells you which data or experiment you need.

go deeper

for a junior

Practise restating a vague business question as a specific one: which action, applied to which customers, changing which metric over which period. That habit alone lifts a screening answer.

for a middle

Be ready to map each phrasing of the ask onto a named estimand, and to explain why a subgroup defined by post-treatment behaviour cannot carry a causal claim.

for a senior

Show that you negotiate the estimand with stakeholders before analysing, and that you can say which design would identify it and what a fallback number is honestly worth.

for a principal

Own the norm that no causal readout ships without a stated intervention, population, outcome window and decision, and be able to defend the cost of holdouts as the price of decisions that survive scrutiny.

## Why the estimand comes first A causal question is only well posed once four things are fixed: **what the intervention is**, **who it would be applied to**, **what outcome is measured over what horizon**, and **what decision the number will drive**. Skip that and the analysis silently answers whichever question the available data happens to fit — usually the easiest one, rarely the one asked. Choosing the estimand first also disciplines the conversation with stakeholders: it converts a vague ask into something that can be right or wrong. ## Unpacking the ask ### 1. The intervention "The coupon" hides at least two distinct treatments. **Being offered** the coupon is something the business controls and can assign; **redeeming** it is a customer behaviour that follows the offer. The offer is the policy lever, so the effect of the *offer* is what a launch decision needs. It is also the only one of the two that can be assigned. This matters for a second reason. If you define the treatment loosely, different units get materially different things — a 10% code in an email versus a banner versus an automatic discount at checkout — and the average effect becomes an average over an ill-defined mixture. State the intervention concretely enough that two people would implement it the same way. ### 2. The population Ask who the number is supposed to describe. All eligible customers? Only the segment currently mailed? New customers only? The estimand changes with the answer: the average effect over the whole eligible base is the ATE; the average effect on the currently-mailed slice is the ATT for the programme as it runs. A pilot restricted to one segment cannot speak for another without an explicit extrapolation. ### 3. The outcome and the window "Work" hides a metric choice. Orders in the next 14 days, gross margin over 90 days, and retention at 6 months will not all move the same way for a discount: the coupon can pull demand forward and subsidise purchases that would have happened anyway. Fix the outcome and the horizon before estimating anything, so the answer cannot be shopped after the fact. ### 4. The decision Different decisions need different estimands: - **Launch to the whole eligible base versus run nothing** — the ATE over that base. - **Keep or kill the current targeted programme** — the ATT on the customers it currently reaches. - **Extend the existing programme to the customers not currently mailed** — the effect on that untreated group, the ATU, which the current programme's data never measured. ## The redeemer trap The most common wrong answer is "compare people who used the coupon with people who did not". Redemption occurs *after* the treatment is delivered, so conditioning on it splits the population by a behaviour the treatment itself influenced. Redeemers are people who were already inclined to buy; non-redeemers include people who never opened the email. The resulting gap is a comparison of self-selected groups and is not a causal effect of anything. The well-defined alternative is the effect of the *offer* on everyone offered, which stays a comparison of assigned arms. The same trap wears other costumes: effect "on engaged users", "among customers who completed onboarding", "for accounts that stayed". Any subgroup defined by post-treatment behaviour has the same defect. Subgroups defined by characteristics fixed *before* assignment — tenure at mailing, region, prior-year spend — are legitimate. ## Running the conversation A workable script: restate the ask as a sentence with all four pieces filled in, and get agreement on it before analysing anything. For example, "the average effect of mailing the 10% coupon to every eligible customer, on gross margin per customer over 90 days, to decide whether to launch it to the whole base." Then say what it would take to answer that — often a randomised holdout — and what the fallback estimate would be worth without it. Stakeholders rarely object to this once they see that different readings of their own question give different answers, and sometimes opposite recommendations. The payoff is not pedantic tidiness. It changes what you build: the launch estimand demands a holdout across the full base, the keep-or-kill estimand can be answered on the mailed slice alone, and the extension estimand demands treating people the programme currently ignores. Getting that wrong costs a quarter, not an afternoon.

  • Why is 'the effect on customers who redeemed the coupon' not a valid estimand?
    Because redemption happens after the coupon is delivered and is influenced by it, so the redeemer group is defined by a post-treatment behaviour. Comparing redeemers to non-redeemers compares self-selected people, not arms of a treatment. The defensible version is the effect of being offered the coupon on everyone offered.
  • The exec wants a number this week and no experiment was run. How do you respond?
    State the estimand anyway, then say clearly what the available data can support and under what assumption. Give the observational estimate with its selection story and the likely direction of bias, propose the smallest randomised holdout that would settle it, and avoid presenting the biased number as if it answered the launch question.
  • Which pre-treatment slices are fair to report alongside the headline estimand?
    Slices defined by characteristics fixed before assignment — tenure, region, prior spend band, acquisition channel. They are legitimate conditional estimands. Fix the list before looking at results, because scanning many slices after the fact turns a reporting exercise into a search for a favourable number.

saying these in an interview costs you the question

  • Picks an estimator before naming the population
  • Compares redeemers with non-redeemers
  • Assumes the exec meant the population-wide effect
  • Quotes a targeted programme's effect for a full rollout
  • Leaves the outcome and time window unfixed until after the analysis

context