A checkout-button test wins only when a concurrent shipping-price test is in treatment. How do you diagnose that?
answer
- four cells, a quarter of traffic each
- difference of differences is noisy
- roughly 4x sample to see it
- check the joint arm table first
- did the other arm change who was eligible
basics
~10 sAsk first whether the pattern is real: the four-cell contrast measuring an interaction carries about twice a main effect's standard error, so flips are usually noise. Then rule out correlated assignment and eligibility differences.
solid answer
~50 sWork through three explanations in order. First, **noise**: two concurrent 50/50 tests split the population into four cells, and the interaction contrast — the difference of two differences — carries roughly twice the standard error of either main effect, so it needs about four times the traffic to detect an effect of the same size. Second, **broken independence**: cross-tabulate the two layers' arms, because correlated assignment confounds the treatments and makes the split meaningless. Third, **eligibility**: if the shipping-price treatment changes who reaches checkout at all, conditioning on that arm compares different populations, which mimics an interaction. Only if all three survive do I call it genuine non-additivity — and then I either rerun the two mutually exclusive, or report the effect averaged over the other test's arms if that mix reflects the world we ship into.
go deeper
Know that splitting a result by another running experiment's arm cuts the sample into four groups, and that small-group differences are frequently just noise.
Explain the arithmetic: the interaction is a difference of two differences over four quarter-sized cells, so its standard error is about double a main effect's and needs roughly four times the traffic.
Demonstrate the full triage — precision, then joint independence of the layers, then whether the other treatment changed who was even eligible — before conceding a genuine interaction, and say what you would rerun.
Own the reporting stance: decide whether the organisation reports effects averaged over the concurrent landscape or conditioned on a specific arm, and defend that choice to stakeholders who want one clean number.
## The pattern A checkout-button experiment reads flat overall, but split by the concurrent shipping-price experiment it looks decisive: strongly positive among users in the shipping treatment, slightly negative among users in the shipping control. The temptation is to declare an interaction and ship conditionally. Resist it long enough to work through three explanations. ## Explanation 1: it is noise Two concurrent 50/50 experiments partition the population into four cells, each holding about a quarter of the traffic. Call them by their arm pair: (0,0), (1,0), (0,1), (1,1). A **main effect** compares two cells against two cells, so each side of that comparison holds half the traffic. With per-user standard deviation `sigma` and `n` users per cell, the standard error is `sigma * sqrt(1/(2n) + 1/(2n)) = sigma / sqrt(n)`. The **interaction** is a difference of differences: `(mean11 - mean01) - (mean10 - mean00)`. Every one of the four cell means enters with a coefficient of plus or minus one, so its variance is the sum of four cell-mean variances: `4 * sigma^2 / n`, giving a standard error of `2 * sigma / sqrt(n)`. So the interaction estimate has roughly **twice** the standard error of a main effect on the same data, and detecting an interaction of the same magnitude as a main effect needs about **four times** the sample. An experiment powered to detect its own main effect is badly underpowered for the interaction, and the noisiest subgroup split will reliably produce a flattering story. Before anything else, compute the standard error of that difference of differences and look at how wide it is. ## Explanation 2: the layers were not independent Cross-tabulate the arm held in the checkout layer against the arm held in the shipping layer. Under proper layering, each of the four cells should hold about a quarter of exposed users. Two possibilities go wrong here. If the cells collapse toward the diagonal, the two layers are sharing or correlating their assignment, the treatments are confounded, and the split you are looking at is not a split at all. If the cells are merely skewed, partial correlation is biasing both main effects, which needs fixing before any interaction claim is entertained. ## Explanation 3: post-treatment conditioning on eligibility This is the subtle one and the most common real cause. Suppose the shipping-price treatment reveals shipping cost earlier in the funnel. That changes *who reaches the checkout page at all* — perhaps the price-sensitive users now abandon earlier, so the checkout population in the shipping treatment is systematically more committed than in the shipping control. The checkout-button experiment is measured on people who reached checkout. Splitting its result by the shipping arm therefore compares two **different populations**, not the same population under two conditions. Any estimate conditioned on a variable that the other treatment itself affects loses the protection of randomization; the difference you see may be entirely composition. Diagnose it by comparing the exposed population sizes and their pre-experiment characteristics across the shipping arms: if the checkout-exposed counts or the pre-period behaviour differ by the shipping arm, you are looking at selection, not interaction. ## If it survives all three Only now is a genuine non-additive interaction the leading explanation — and it is plausible: both treatments push on the same moment of the funnel, and once shipping cost stops being a surprise the button's prominence may matter more. Options, in ascending cost: **Report the average.** A layered platform estimates each effect averaged over the other experiment's arms. If the shipping change is going to ship, the world your button will live in *is* the shipping-treatment world, and the relevant number may simply be the effect within that arm — acknowledged as an underpowered estimate. **Rerun deliberately.** Rerun the two experiments in a mutually exclusive group so each is read against a clean baseline, or rerun the button test after the shipping change has fully launched so the baseline is the new world. **Power for the interaction.** If the combination decision is genuinely valuable, run the 2x2 on purpose with enough traffic for the interaction contrast — roughly four times what the main effect needed. ## The judgment being tested An interviewer is checking whether you reach for "interaction!" immediately or whether you first ask how precise that four-cell contrast actually is, whether the randomization held, and whether the second treatment moved the population you are conditioning on.
- Why does an interaction need roughly four times the sample size of a main effect?The main effect compares two cells against two, so each side holds half the traffic. The interaction is a difference of two differences, and all four cell means enter with weight one, so its variance is four cell-mean variances — a standard error twice as large. Since required sample scales with the square of the standard error, doubling it means quadrupling the traffic.
- How would you tell a composition effect apart from a real interaction?Compare the exposed populations across the other test's arms: counts reaching the surface, and pre-experiment characteristics measured before either treatment applied. A true interaction leaves those identical and changes only the outcome; a composition effect shows up as different exposure volumes or different pre-period behaviour. Pre-period covariates cannot be affected by treatment, so any imbalance in them is diagnostic.
- If the interaction is real, which number do you report to the decision-maker?It depends on the world the change ships into. If the shipping treatment is launching, report the button effect within that arm and flag its wider interval. If the shipping test may be abandoned, the averaged effect across arms is the honest summary. Either way, state explicitly which arm mix the estimate is averaged over rather than quoting a bare lift.
saying these in an interview costs you the question
- Treats a sign flip across four cells as proof of interaction
- Ignores that the interaction contrast is far less precise
- Splits results by a variable the other treatment affects
- Immediately blames the hashing without checking the joint table
- Ships conditionally on an underpowered subgroup result