When should two experiments be mutually exclusive rather than overlapping in separate layers?
answer
- same layer means never both
- incoherent combination, not mere unease
- two redesigns of one page
- cost is shared traffic and longer runs
- one control arm, several variants
basics
~10 sMake two experiments mutually exclusive when their treatments cannot sensibly coexist for one user, such as competing redesigns of the same page. Everything else should overlap, because exclusivity costs traffic and time.
solid answer
~50 sMutual exclusivity means the two experiments live in the *same* layer and split that layer's traffic between them, so a user can enter one or the other but never both. Reach for it when the combination is incoherent rather than merely uncertain: two teams each testing a full redesign of the same checkout page cannot both render, so they belong in one exclusive group. It also helps when you expect a strong interaction and want a clean read of each variant on its own. The cost is real — each experiment now gets a slice of the layer instead of the whole population, so it needs longer to reach the sample size it needs, and you lose the ability to learn anything about the combination. Because of that, overlapping should be the default and exclusivity the deliberate exception, granted per shared surface rather than per team preference.
go deeper
Know the definition: mutually exclusive experiments sit in one layer and split its traffic, so a user is in at most one of them, while different layers let a user be in both.
Explain the criterion and the cost. Exclusivity is for treatments that cannot coexist on the same surface, and it lengthens each experiment because they now share one layer's traffic.
Show the tradeoff in numbers and the escape hatch: quantify the duration penalty of splitting a layer, and use a shared control arm to reclaim traffic while acknowledging the correlation it introduces.
Own the policy. Argue for exclusivity granted by surface ownership rather than by team request, and be able to defend that rule to a team that wants a private universe for every launch.
## The two placements A layered platform offers exactly two ways to relate a pair of experiments. **Different layers (overlapping).** Both experiments randomize the full eligible population independently. Every user gets an arm in each, so all four combinations occur. Both experiments enjoy the whole traffic, and each estimate is unbiased for that experiment's effect averaged over the other's arms. **Same layer (mutually exclusive).** The layer's arm space is carved up between the experiments. A user who falls into experiment A's range is in experiment A and is simply not eligible for experiment B. No user ever sees both treatments, and the combination is never observed. ## When exclusivity is the right call **The combination is incoherent.** This is the dominant reason. Two competing redesigns of the same page cannot both render — one of them wins the DOM and the other silently does nothing, or the page becomes a hybrid nobody signed off on. The same applies to two experiments that each change the same copy string, the same pricing rule, or the same ranking decision. Here exclusivity is not a statistical nicety; it is a correctness requirement. **You need a clean read on each variant.** When a strong interaction is expected — two changes both aimed at the same behaviour, pushing on the same friction — the overlapping design still gives you unbiased main effects, but each is an average over the other's arms, and stakeholders may find that hard to act on. If the decision you are about to make is "which of these two designs do we ship", an exclusive group gives you two clean contrasts against a common baseline. **Risk containment.** A user who is simultaneously in three aggressive treatments is the user most likely to have a terrible session. Some platforms keep a small set of high-risk experiments mutually exclusive purely to bound how much novelty any single person absorbs at once. ## When exclusivity is the wrong call It is the wrong call when the only argument is discomfort. "I don't want anything else running during my test" is not a reason — independent layers already give the estimate the cleanliness that argument is reaching for, and granting exclusivity on request turns a scarce, shared resource into a queue. It is also wrong when the combination is exactly what you want to learn about: if two changes are expected to ship together, overlapping them lets you observe all four cells and see whether the pair behaves as hoped. ## The price Suppose an exclusive layer holds three experiments. Each gets roughly a third of the traffic. Because the width of a confidence interval shrinks with the square root of the sample size, a third of the traffic makes each interval about 1.7 times wider at a fixed duration — or, equivalently, forces each experiment to run about three times as long to reach the same precision. Multiply that across an organisation and exclusivity-by-default is the single biggest brake on experiment throughput. A **shared control arm** recovers part of that cost. If three variant experiments in one exclusive group are all being compared against the current experience, they do not each need their own control: allocate one control arm and compare all three variants against it. Four arms of 25% become one 25% control plus three 25% variants instead of three separate 20/20 pairs, so each comparison gets more users on both sides. The caveat is that the three comparisons now share a denominator, so their estimates are positively correlated — a control group that happens to run high makes all three variants look worse together, which matters when you rank them against each other. ## Design guidance that survives contact with an org The workable rule is to define exclusivity by **surface ownership**, not by team request. Each surface that can only be rendered one way — the checkout page, the search results template, the pricing rule — gets a layer, and experiments touching it queue in that layer as an exclusive group. Everything that touches a different surface goes into a different layer and overlaps freely. That policy is mechanical, easy to arbitrate, and puts the cost where the genuine conflict is. ## What a strong answer sounds like Name the mechanism (same layer, split traffic, never both), give the incoherent-combination criterion, quantify the cost in duration or interval width, and mention shared controls as the way to claw some of it back.
- What do you lose analytically by making two experiments mutually exclusive?The combination cell. No user ever receives both treatments, so you cannot estimate what happens when they ship together — only each against the shared baseline. If the two changes are destined to launch as a pair, that is precisely the number you most wanted, and an overlapping design would have given it to you.
- How does a shared control arm help an exclusive group, and what does it complicate?Three variants compared against one common control need far less traffic than three separate variant-plus-control pairs, because the control users are reused across all three contrasts. The complication is correlation: every comparison shares the same control mean, so an unusually high or low control shifts all three estimates in the same direction, which distorts head-to-head ranking of the variants.
- A team asks for exclusivity so 'nothing else interferes' with their test. How do you respond?Explain that independent layers already balance every other experiment across their arms, so the estimate is unbiased without exclusivity. Exclusivity buys isolation from a *combined rendering* problem, not from statistical contamination. If their surface genuinely cannot host two changes at once, grant it on that basis; if not, the request costs traffic that other teams need.
saying these in an interview costs you the question
- Grants exclusivity on request instead of by surface conflict
- Thinks exclusivity is free because assignment is still random
- Believes overlapping tests statistically contaminate each other
- Forgets a shared control correlates the comparisons
- Makes tests exclusive when the combination is the thing being decided