skip to content

A lift measured on US desktop power users is proposed for a mobile-first market abroad. How do you judge whether it transports?

level: principalimportance: should knowfreq 45%

answer

  1. internal validity does not imply external
  2. it turns on effect modification
  3. compare modifier distributions, then reweight
  4. overlap is required, not optional
  5. relative and absolute transport differently

basics

~20 s

Ask which variables modify the effect and how their distribution differs in the target population. An estimate transports only if the effect modifiers are measured, overlap between populations, and the mechanism and baseline rate carry over. Otherwise reweight or re-test.

solid answer

~50 s

Internal validity - the estimate is right for the population it was measured on - does not imply external validity. Transport turns on effect modification: if the effect varies with device type, tenure, bandwidth or baseline engagement, then a target population with a different mix of those will see a different average effect. The practical procedure is to list the plausible modifiers, compare their distributions between source and target, and reweight the source estimate toward the target's covariate mix. That requires every relevant modifier to be measured and to have support in both populations - you cannot reweight toward mobile-heavy users if the source had almost none. Separately, check the scale: a relative lift and an absolute lift transport differently when baseline rates differ. When the mechanism itself may differ, reweighting is not enough and a confirmatory test in the target is the honest answer.

go deeper

for a junior

Know that an effect measured on one group need not hold for another, and that the reason is the effect differing across people rather than the original measurement being wrong.

for a middle

Explain effect modification concretely: name the variables the effect plausibly depends on, and describe reweighting a source estimate toward the target population's mix of those variables.

for a senior

Demonstrate the checks - measured modifiers, overlap, baseline rates, scale of the effect - and say when reweighting is inadequate because the intervention or context itself differs.

for a principal

Own the triage. Decide which markets get a confirmatory test and which ship on a transported estimate, judged by whether plausible modification would change the decision rather than the number.

## Two different validities An estimate can be perfectly correct and still be the wrong number to act on. Internal validity asks whether the effect you measured is the true effect for the units you measured it on. External validity asks whether it is the true effect for the units you now care about. A clean randomised result on US desktop power users has strong internal validity for exactly that group and says nothing on its own about a mobile-first user base in a different country. ## The mechanism: effect modification Transport fails for one structural reason. If the treatment effect is constant across everyone, the average effect is the same in any population and transport is trivial. It fails when the effect varies with some characteristic - an effect modifier - and the two populations differ in the distribution of that characteristic. A redesign that helps users with slow connections a great deal and heavy desktop users not at all will produce a small average lift in a desktop-heavy sample and a large one in a mobile-heavy one, with no contradiction between the two measurements. So the first question is never `will it transport?` but `what does the effect depend on, and does the target differ on that?` ## The formal condition and the practical procedure Reweighting a source estimate to a target population requires two things. Every variable that modifies the effect must be measured in both populations - unmeasured modification is exactly as untestable as unmeasured confounding, and calls for the same kind of sensitivity reasoning. And there must be overlap: every covariate profile common in the target must also occur in the source. If the source contained almost no low-bandwidth mobile users, no weighting scheme conjures an estimate for them; the weights simply blow up on a handful of units and the reweighted result carries an enormous, often understated, variance. Given both, the procedure is to estimate how the effect varies across the modifiers in the source, then average that variation over the target's covariate distribution rather than the source's. In practice this looks like a weighted average of subgroup effects, weighted by the target's subgroup shares. ## The scale question A point regularly missed: the same effect can transport on one scale and not another. Suppose a change reduces checkout abandonment by 20 percent relative, and the baseline abandonment rate is 5 percent in the source but 30 percent in the target. On the relative scale, the effect is identical; on the absolute scale it is 1 percentage point versus 6 - a sixfold difference in the number that drives the business case. Neither scale is inherently the transportable one; which is more stable is an empirical question about the mechanism. What is not acceptable is silently transporting an absolute effect when the baseline rate differs by a factor of six. Baseline differences also change the ceiling. If the target already converts at 80 percent, an intervention that removes friction has less room to work regardless of how well the mechanism carries over. ## When reweighting is not enough Reweighting assumes the same treatment does the same thing to the same kind of person, only the mix of people changes. Several failure modes break that assumption outright: - **The intervention itself differs.** A version rendered on a small screen, translated, or routed through different payment infrastructure is not the intervention that was tested. - **The context differs.** Competitive alternatives, network conditions, regulatory constraints and price sensitivity all change what the same feature means to a user. - **Interference.** If the effect works through social or marketplace channels, the equilibrium response depends on how many others are treated, and a result from a small treated share does not scale. - **Time.** A result from eighteen months ago transports to today only if the product and the user base have not moved. Temporal transport is the version of this problem that teams forget entirely. ## The judgment call A principal-level answer is not `re-test everything`. Confirmatory tests cost time, and many markets are too small to power one at the effect size in question. The defensible framing is a triage. Ship on the transported estimate when the mechanism is plainly universal, the modifiers you can measure are similar between populations, and the downside of being wrong is bounded and reversible. Reweight and ship with monitoring when the populations differ on measured modifiers but the mechanism is intact - and instrument the guardrail metrics you would need to catch a reversal. Insist on a confirmatory test when the mechanism itself may differ, when the decision is expensive or hard to reverse, or when the transported estimate is close to the decision threshold so that modest modification would flip the sign of the call. The last of these is the real discriminator. What matters is not whether the number changes but whether the *decision* changes. A lift that would still clear the bar under any plausible reweighting does not need a new test; one that clears it only under the source population's mix does. Finally, treat every transported claim as a hypothesis with an owner and an expiry. Record which population the estimate came from, which modifiers were assumed stable, and what would falsify the assumption. Organisations accumulate confident effect sizes whose original population nobody remembers, and those become the hardest errors to unwind.

  • Why is reweighting impossible when the source population lacks the target's profiles?
    Weighting redistributes information you already have; it cannot create an estimate for a profile you never observed. With near-zero support, a handful of units carry enormous weight, and the reweighted estimate becomes both unstable and driven by extrapolation from a model rather than by data.
  • Which is more likely to transport - a relative or an absolute effect?
    It depends on the mechanism and is an empirical question, not a rule. Effects working by removing a proportional friction often carry better on the relative scale, while effects with a fixed ceiling may carry better in absolute terms. State which scale you assumed and why, because baseline differences can make the two answers differ severalfold.
  • What is the analogue of unmeasured confounding in a transport argument?
    Unmeasured effect modification. If a variable changes the size of the effect, differs between populations and is not recorded, no reweighting can correct for it and no data can detect it. The response is the same as for confounding: bound how much modification it would take to change the decision.
  • How do you handle a result measured eighteen months ago on today's users?
    Treat it as a transport problem across time. The population, the product surface and the competitive context have all shifted, and none of that is captured by the original covariates. If the decision is significant, re-measure; if not, state the assumption of temporal stability explicitly so it can be challenged later.

A drug dose validated on adults does not carry to children by arithmetic alone. The mechanism may hold, but the population differs on things that change the response, so you check those before you prescribe.

saying these in an interview costs you the question

  • Assumes a clean randomised result generalises everywhere
  • Ignores that the effect may vary across user segments
  • Reweights without checking overlap between populations
  • Carries an absolute lift across very different baseline rates
  • Forgets that results also fail to transport across time

context