skip to content

In a social feed A/B test, why can users in the control arm be affected by the treatment?

level: juniorimportance: should knowfreq 58%

answer

  1. one user's outcome depends on others
  2. friends land in different arms
  3. control is partly treated
  4. the gap between arms shrinks
  5. more traffic does not fix it

basics

~20 s

Users influence each other. If treated users post or share more, their control-arm friends see that extra content, so the control arm is partly treated and the measured gap between arms understates the feature's true effect.

solid answer

~50 s

A standard A/B test assumes one user's assignment does not change another user's outcome. In a social product that assumption breaks: the feed is built out of other people's activity, so a control user whose friends are in treatment is exposed to the treatment through them. If the feature makes treated users post more, control users see more content and their engagement rises too — control is contaminated upward, the treated-minus-control difference shrinks, and the test understates what a full launch would deliver. Spillover can also run the other way: if treated users' extra posting pushes control users' own posts out of ranked feeds, control is depressed and the test overstates the effect. The direction is not automatic; it follows from the mechanism. The fix is a design that keeps interacting users in the same arm, not a bigger sample.

go deeper

for a junior

Be ready to say in plain words that an A/B test assumes one user's assignment does not affect anyone else, and that social products break this because content flows between friends across arms.

for a middle

Explain the mechanism and the direction: work out whether the leaked exposure raises or lowers control outcomes, and show how that shrinks or inflates the measured difference between the arms.

for a senior

An interviewer expects you to catch contaminated results before a launch decision, propose a containment design, and put a rough size on how much of the observed lift could be leakage.

for a principal

Own the standard for when a social product's results can be trusted at all: which launches require an interference-safe design, and how much precision the organisation should trade for an unbiased answer.

## The assumption that quietly breaks Every ordinary A/B test rests on an assumption so basic it usually goes unstated: **the outcome of a user depends only on that user's own assignment**. Randomization guarantees the two arms are comparable *before* treatment; this extra assumption is what lets you read the difference in average outcomes as the effect of the feature. When users can influence one another, the assumption fails and the difference between arms stops answering the question you asked. This failure is called **interference**, or **spillover**. Social products break it by construction. A feed is not a private experience assembled from a user's own actions — it is assembled from other people's actions. If your friend is in the treatment arm and the treatment makes them post more, your feed changes even though you were never assigned the feature. ## How the contamination flows Think of the product as a graph: users are nodes, and an edge means one user's activity can reach the other. Randomizing users independently means that, for a 50/50 split, roughly half of every user's edges cross the arm boundary. Almost nobody is a *pure* control: a control user with twenty friends is, in expectation, exposed to ten treated friends. The control arm is therefore not the world without the feature — it is a diluted, partially treated world. The consequence depends on the sign of the spillover: - **Positive spillover.** The feature makes treated users create or share more; control friends receive more content and engage more. Control rises, treatment rises, and the *difference* between them is smaller than the true effect of launching to everyone. The test is **biased toward zero** and you may ship nothing because the result looked flat. - **Negative spillover.** Treated users' extra activity crowds a fixed feed, displacing content control users would otherwise have seen and dragging their engagement down. Control falls while treatment rises, so the difference is **larger** than the true global effect, and the launch underdelivers. Both are common, and which one you get is a fact about the mechanism, not something you can assume. Reasoning through the mechanism before reading the result is the skill being tested. ## Why this is bias, not noise The most frequent mistake is to treat spillover as an imprecision problem and reach for more traffic or a longer run. Contamination does not average out with sample size. A larger experiment gives you a tighter confidence interval around the *contaminated* contrast: a more precise estimate of a quantity you do not want. Statistical significance is no protection either — a heavily spilled-over experiment can be significant, replicable and consistently wrong about launch impact. It is also not fixed by checking that the arms look balanced. The split can be perfectly 50/50 and every pre-treatment covariate balanced, and the estimate will still be biased, because the problem happens *after* assignment, in how treatment reaches users it was never assigned to. ## Detecting it A cheap first look is a **dose-response** check inside the control arm: among control users, compare those with many treated friends to those with few, holding the total number of friends roughly constant. If control users with more treated contacts behave systematically differently, exposure is leaking. This is suggestive rather than conclusive — the number of treated friends is random given the number of friends, which helps, but activity levels of the friends themselves are not controlled. Stronger evidence comes from a design that deliberately creates both a *pure* control and an *exposed* control: randomize groups of connected users, and within groups assigned to treatment, randomize individuals. Comparing untreated individuals inside treated groups against individuals in fully control groups measures the spillover directly. ## Fixing it The fix is always a **design** change, never an analysis patch. The general principle is to make the randomized unit large enough to contain the interaction: assign whole clusters of connected users to the same arm so that most edges stay inside one condition, or switch an entire market or time window between conditions so that everyone in it experiences the same world. Both come at a real cost — the independent replicates become clusters or time windows rather than users, so precision drops sharply — and that trade is the substance of the design conversation. ## What good sounds like in an interview Name the assumption, name the mechanism by which it breaks in this specific product, state which direction the bias runs and why, say explicitly that sample size does not help, and propose a containment design rather than a statistical correction. Candidates who stop at "there might be some spillover" have identified the word without demonstrating they can reason about the consequence.

  • Does increasing the sample size reduce spillover bias?
    No. Spillover is a bias, not noise. More users tighten the confidence interval around the contaminated contrast, giving a more precise estimate of the wrong quantity. Only a design change removes it — keeping connected users in the same arm, or switching whole markets or time windows between conditions.
  • How would you get evidence that spillover is actually happening, rather than just suspecting it?
    Look for dose-response inside the control arm: compare control users with many treated friends against those with few, holding total connections roughly constant. A gradient is suggestive. Stronger evidence comes from a two-level design that randomizes groups and then randomizes individuals inside treated groups, so both pure and exposed controls exist.

saying these in an interview costs you the question

  • Claims a randomized test is automatically unbiased
  • Assumes spillover always inflates the measured effect
  • Suggests a larger sample fixes contamination
  • Treats connected users in opposite arms as independent
  • Proposes an analysis correction instead of a design change

context