How do you set an experimentation platform's policy on which tests may overlap?
answer
- overlap is the default, exclusivity the exception
- the requester never pays the cost
- grant it by surface, not by request
- monitor after, do not gate before
- state what the reported effect averages over
basics
~10 sDefault to overlapping everything and grant exclusivity per shared surface, never per team request. Monitor for interactions retrospectively rather than gating launches, and state openly that reported effects average over the concurrent landscape.
solid answer
~50 sThe policy question is really about who bears the cost of caution. Exclusivity feels safe to the requesting team and is paid for by everyone else in queue time, so it cannot be self-service. I set the default to full overlap, define layers around **surfaces that can only render one way** — the checkout page, the ranking template, the pricing rule — and grant exclusivity automatically when two experiments contend for the same surface and never otherwise. Shared control arms inside those exclusive groups claw back some of the traffic. Instead of gating launches on interaction review, I run periodic scans on the highest-traffic concurrent pairs and act on the rare real finding, because pre-restriction throttles hundreds of experiments to prevent a handful of problems. And I make the reporting convention explicit: every effect is an average over whatever else was running, which is usually the number the ship decision needs anyway.
go deeper
Know that a platform deliberately lets most experiments overlap, and that making them exclusive is a decision with a traffic cost rather than a free safety measure.
Be able to explain the mechanics behind the policy: overlapping preserves each experiment's full sample, while an exclusive group divides one layer's traffic and lengthens every run inside it.
Show you can operate the rule — classify experiments by the surface they render, use shared controls to reclaim traffic, and run retrospective interaction scans instead of blocking launches.
Own the allocation and the narrative: who may spend exclusivity, how many layers the surface taxonomy needs, and the written convention that reported effects are averages over the concurrent landscape.
## Why this is a policy question, not a statistics question The statistics are settled: independent layers give unbiased estimates of each experiment's effect averaged over everything else running. The open question is organisational — how much throughput to trade for how much interpretive comfort, and who gets to make that trade. Left ungoverned, exclusivity becomes a status good: whoever asks loudest gets a private universe, and the platform's capacity quietly collapses. ## Default: overlap The starting position should be that everything overlaps. Layering exists precisely so that traffic is not the binding constraint on how many decisions the organisation can make in a quarter. Every experiment moved into an exclusive group takes a share of a layer rather than the whole population, and because precision improves with the square root of sample size, splitting a layer three ways roughly triples the time each experiment needs for the same interval width. Across a large program, exclusivity-by-default is the difference between hundreds of decisions a year and dozens. ## The exception rule: surface contention A rule has to be mechanical or it becomes a negotiation. The cleanest criterion is **rendering contention**: if two experiments modify the same thing and only one of them can win, they cannot overlap. That covers competing page redesigns, two changes to the same copy string, two rules setting the same price, two rerankers on the same result list. Give each such surface its own layer; experiments touching it queue there as a mutually exclusive group. This rule has three virtues. It is checkable from the experiment's own configuration, so it can be automated. It puts the cost exactly where a genuine conflict exists. And it removes the argument entirely from the realm of preference — a team cannot request exclusivity, they can only declare which surface they touch. ## Reclaiming traffic inside exclusive groups When several variants of one surface compete, a **shared control arm** cuts the cost substantially: instead of each variant carrying its own control, one control serves all comparisons, so more of the group's traffic goes to the things being tested. The tradeoff to state up front is that the comparisons become positively correlated through the common control, so head-to-head ranking of the variants is less reliable than each variant-versus-control contrast is on its own. For a ship-or-not decision that is a good bargain; for a rank-these-five decision, less so. ## Monitor rather than pre-restrict The instinct to review every launch for possible interactions with everything else running does not scale and does not pay. Interactions between unrelated surfaces are usually small, the four-cell contrast that would detect them is inherently imprecise, and a review board that must clear every start becomes both a bottleneck and a thing teams route around. The better posture is retrospective: periodically scan the highest-traffic concurrent pairs for non-additivity, investigate the small number that look real, and feed genuine findings back as new surface definitions. That converts a permanent tax into an occasional, evidence-driven cost. ## Make the reporting convention explicit The most valuable thing a leader can do here is stop the recurring anxiety that concurrent experiments are corrupting results. Write down, once, that a reported effect is the effect **averaged over the experiment landscape in which it ran**, that this is unbiased for that quantity, and that it is typically the decision-relevant number because the product ships into a world that also contains other changes. Pair it with the honest caveat: for two changes aimed at the same user behaviour, the average may hide a real interaction, and that is when a deliberate 2x2 or an exclusive rerun is worth the traffic. ## Where to spend the scarce judgment A few calls genuinely deserve leadership attention rather than a rule. How many layers to define at all — too few and everything contends, too many and the surface taxonomy stops matching how the product is built. Whether to bound how many concurrent treatments a single user can absorb, trading a little throughput for a less erratic experience for the unlucky. And how to handle a strategically important launch whose read must be unimpeachable, where buying exclusivity is a reasonable one-off purchase rather than a precedent. ## What a strong answer conveys That exclusivity is a resource allocation problem with a mechanical rule attached, that the default must be overlap or throughput dies, and that the organisation is better served by an explicit reporting convention plus retrospective monitoring than by a pre-launch review board.
- A senior team insists every launch of theirs needs a private, non-overlapping universe. What do you do?Separate the two claims. If their change contends for a surface someone else is also testing, exclusivity is automatic and no argument is needed. If it does not, explain that independent layers already balance every other experiment across their arms, so the estimate is unbiased — and that the request costs other teams queue time. Reserve discretionary exclusivity for genuinely strategic launches, visibly and rarely.
- How would you detect that concurrent experiments are interacting at a rate worth worrying about?Run a retrospective scan across the highest-traffic concurrent pairs, estimating the four-cell non-additivity for each and comparing the distribution of those estimates against what pure noise would produce. If the observed spread matches the noise expectation, the program is fine. A visible excess, especially concentrated in pairs touching related surfaces, is the signal to redraw the surface definitions.
- How do you explain to an executive that many experiments run on the same users at once?Say that each experiment randomizes the population afresh, so everything else running is split evenly between its treatment and control and cancels out of the comparison. The measured lift is the lift in the real, busy world the product actually ships into. Then name the one exception you actively manage: two changes aimed at the same behaviour, which get run against each other deliberately.
saying these in an interview costs you the question
- Lets teams self-serve exclusivity on request
- Gates every launch behind an interaction review board
- Treats concurrent experiments as inherently invalidating
- Never states what the reported effect averages over
- Defines layers by team ownership instead of rendered surface