A launch can only be evaluated with difference-in-differences — how do you decide whether to act on it?
answer
- grade the design before the number
- ask what really ruled out randomising
- match the bar to reversibility
- state the assumption in business language
- name the confirmatory readout
basics
~10 sGrade the design before the number: why randomisation was impossible, how the comparison group was chosen, whether the pre-adoption path is flat. Then match the evidence bar to how reversible the decision is.
solid answer
~50 sI start by asking whether randomisation was genuinely unavailable or merely inconvenient — a staged rollout with randomised order is often still possible and worth a delay. If it truly was not, I grade the design rather than the point estimate: was the comparison group chosen for a substantive reason before anyone saw the outcome, do the pre-adoption coefficients sit flat, and how large a differential trend would be needed to erase the result. Then I match the evidence bar to the decision. A reversible rollout with cheap rollback can proceed on a quasi-experimental readout; a permanent pricing commitment or a contract cannot. I report a range with the assumption stated in business language — this assumes the two markets would have moved together — not a clean number that reads like an experiment. And I name the confirmatory readout: what randomised or staged evidence we will collect next, and by when.
go deeper
Know that a quasi-experimental estimate rests on an assumption a randomised experiment does not need, and that the assumption belongs in any report of the number.
Be ready to say which specific checks you would run before trusting the readout, and to describe how the comparison group was chosen and why that choice matters.
Show that you can quantify fragility — how large a differential trend would erase the result — and that you can state what evidence would change your conclusion.
Own the decision framing: scale the evidence bar to reversibility, push back on soft reasons randomisation was skipped, and commit the organisation to a confirmatory readout with a date.
## The real question When the only causal evidence for a launch is a quasi-experimental readout, the question stops being 'what is the estimate' and becomes 'how much of a decision will this estimate carry'. That is a judgment call, and it has structure. ## Step one: interrogate the constraint Before accepting the design, push on why the experiment was impossible. The genuine reasons are real: the treatment was a regulatory change, a whole-market launch, a contractual commitment, a feature that cannot be withheld from half a country. The soft reasons are not: 'the team already shipped it', 'legal was slow', 'we did not want to hold anyone back'. The useful middle ground is a **staged rollout with randomised order**. If a feature is going to every market eventually, the order it reaches them in is often free to randomise, and that converts a quasi-experiment into something much stronger at almost no cost. A lead's most valuable contribution is often noticing this *before* the launch, not analysing around it afterwards. ## Step two: grade the design, not the number Ask these in order, and notice that none of them is about the point estimate: **Was the comparison group chosen before anyone saw the outcome, and for a substantive reason?** A comparison market picked because it shares the same product maturity, seasonality and competitive structure is evidence. A comparison market picked because it made the chart look good is not, and if several were tried, the reported result is the maximum of several draws. **What does the pre-adoption path look like, and how informative is it?** Flat pre-adoption coefficients help. Flat pre-adoption coefficients with intervals wide enough to hide the entire estimated effect help very little. Insist on the magnitude the check could not have ruled out, not on a pass or fail. **How large a violation would it take to erase the result?** If a differential trend of a size the two markets routinely exhibit would wipe out the effect, the readout is not decision-grade. If it would take a divergence larger than anything in three years of history, that is a much stronger position, and it is a statement executives can evaluate. **What is the estimate an estimate of?** The effect on the treated markets, over the observed window. Not the global effect, not the steady-state effect. Rolling out worldwide on a single-market readout extrapolates twice: to other markets, and to a longer horizon. ## Step three: match the bar to the decision This is where leads earn their title. Evidence requirements should scale with the cost of being wrong, not be a fixed significance level. - **Cheap and reversible** — a feature flag you can turn off, a rollout you can pause. A credible quasi-experimental readout with a plausible effect size is enough. The cost of a wrong 'yes' is a rollback. - **Expensive but reversible** — an engineering investment, a marketing spend. Require the design to survive scrutiny and set a decision review with better evidence at a stated date. - **Effectively irreversible** — a pricing change, a contract, a deprecation, anything with a public commitment. A quasi-experimental readout should not carry this alone. Either find a way to randomise part of it, or accept the decision is being made on judgment with the estimate as one input among several. ## Step four: report it as what it is Give the point estimate with a range, and one sentence naming the assumption in the language of the business: 'this assumes that, without the launch, usage in the two markets would have moved together — they did for the previous eight quarters'. Resist the pressure to compress it into the same format an experiment readout uses, because the identical format implies identical warrant. Say explicitly what would change your mind: a pre-adoption divergence appearing when you extend the history, a comparison market shock you had not known about, a second market's rollout disagreeing with the first. ## Step five: plan the confirmation The estimate is a bet, and bets should be settled. Name the follow-up: the next market rollout randomised in order, a holdback in a market where one is legal, a period of staged reactivation. Organisations that never revisit a quasi-experimental call accumulate a portfolio of decisions whose evidence nobody ever verified — and each one becomes precedent for the next. ## The failure modes to name The two symmetrical failures are treating a quasi-experimental estimate as interchangeable with a randomised one, and refusing to act on anything short of a randomised trial. The first ships bad decisions confidently. The second cedes every decision that cannot be randomised — which includes most of the important ones — to whoever is loudest in the room.
- How do you communicate a quasi-experimental result to executives who want one number?Give the estimate with a range, then one sentence naming the assumption in their language — 'this assumes the two markets would have moved together without the launch, as they did for eight quarters'. Attach the decision it supports and the one it does not. Do not format it identically to an experiment readout, because identical formatting implies identical warrant.
- What would make you refuse to act on a difference-in-differences readout at all?A sloping pre-adoption path, a comparison group plausibly touched by the launch itself, an estimate smaller than the routine divergence between the two markets, or a comparison group that was clearly chosen after seeing outcomes. Any of those means the number is descriptive, and I would label it that way rather than negotiate around it.
- When is running the randomised experiment worth the delay?When the decision is expensive to reverse, when the effect sits near the threshold that flips the call, or when the same decision will recur — then the experiment amortises across every future instance. A staged rollout with randomised order often gets you most of the way for almost no delay at all.
- How do you stop teams from shopping for the comparison group that gives the best answer?Require the comparison group and the analysis window to be registered before the outcome is looked at, and require any alternative comparison groups tried to be reported alongside the chosen one. Choosing on substance is defensible; choosing on the resulting chart is a garden of forking paths dressed up as causal inference.
saying these in an interview costs you the question
- Treats a quasi-experimental estimate as interchangeable with a randomised one
- Reports a point estimate with no identifying assumption stated
- Applies the same evidence bar regardless of how reversible the decision is
- Accepts that randomisation was impossible without probing why
- Never plans a confirmatory readout after acting