Why must an A/B test give the same user the same variant on every visit?
answer
- same person, same experience
- re-rolling blends the two arms
- returning users become partly treated
- estimate pulled toward zero, not noisier
- deterministic hash, no stored coin flip
basics
~20 sSticky assignment gives each user one consistent experience instead of a flickering mix. If the variant is re-drawn per visit, returning users receive both treatments, which blends the two arms and shrinks the measured difference toward zero.
solid answer
~50 sAssignment has to be sticky for two reasons. The product reason is obvious: a user who sees the new checkout on Monday and the old one on Tuesday gets a confusing experience, and that inconsistency itself changes behaviour. The statistical reason is worse. If the variant is re-drawn on each visit, every returning user is part treated and part control, so the two arms stop being two different populations and become the same blended population measured twice. The estimated effect is diluted toward zero, and you ship a verdict of `no difference` for a feature that actually works. In practice stickiness comes from deriving the assignment from a hash of a stable identifier rather than storing a coin flip, so the same identifier always maps to the same arm, in every service, with no lookup.
go deeper
Be ready to state the rule plainly: one user, one variant, for the whole experiment. Know that the arm is normally computed from the user's identifier rather than drawn fresh at request time.
Explain the failure mechanically: users who see both variants are partly treated, so the arms blend and the estimated effect shrinks toward zero rather than merely getting noisier. Describe how a deterministic hash gives stickiness with no stored state.
Expect to be asked how you would catch broken stickiness in production - logging the assigned arm on every exposure and alerting when one identifier appears under two arms in the same experiment.
Own the platform rule: assignment is computed once from a documented identifier and is never re-drawn during a running test. Decide which exceptions exist, such as a deliberate re-salt between experiments, and who may authorise them.
## What sticky assignment means An online experiment splits users into arms and shows each arm a different version of the product. **Sticky assignment** means that once a user is placed in an arm, they stay in that arm for the whole experiment: every visit, every device session in which they are recognisable, every request. The opposite is a fresh random draw each time the user shows up, which sounds harmless and is not. ## The product argument A product that changes shape between visits is a bad product. A user who found a feature yesterday and cannot find it today files a bug, contacts support, or simply stops trying. Worse, the inconsistency is itself an intervention: you are no longer measuring `new design vs old design`, you are measuring `stable experience vs unstable experience`, which is not a shipping decision anyone wants to make. ## The statistical argument This is the one interviewers are usually probing for. An experiment estimates a difference between two groups that differ in exactly one thing: which version they received. If assignment is re-drawn per visit, a user with ten visits receives roughly five of each. Averaged over their behaviour, that user is now half treated. Do this to everyone and both arms contain the same mixture, so the difference between arm means collapses. The practical consequence is **attenuation**: the estimated effect is pulled toward zero, roughly in proportion to how much cross-contamination there is. The confidence interval does not widen to warn you; it narrows around the wrong value, because you still have plenty of observations. So the failure mode is not noisy results, it is confidently wrong results. A real 3% lift is measured as 0.4%, declared non-significant, and the feature is killed. Note which direction this goes. Contamination almost always biases toward `no effect`, so it produces false negatives rather than false positives. That is why it survives so long undetected: nothing looks broken, experiments just quietly stop finding anything. ## How stickiness is implemented The standard mechanism is deterministic: take a stable identifier for the randomisation unit, combine it with something specific to the experiment, hash the result, and map the hash into an arm. Because the function is pure, any service can recompute the same answer from the identifier alone. There is no assignment table to read on the request path, nothing to replicate between regions, nothing to lose, and the analysis job can rebuild every assignment from logs after the fact. The alternative — flip a coin the first time you see a user and store the result — also gives stickiness in principle, but it buys it with state. That state has to be written on first exposure, read on every subsequent one, kept consistent across services, and it becomes a source of outages and of silent divergence when two systems disagree about who is in which arm. ## What legitimately changes an assignment Stickiness does not mean an assignment is frozen forever. It means it never changes *spontaneously*. Assignments legitimately move when you deliberately re-salt an experiment (which reshuffles everyone and must therefore happen between experiments, never during one), or when the exposed fraction of the population is deliberately widened so that users who were previously unexposed become eligible. Widening exposure is safe precisely because it adds users without moving the ones already assigned. ## How you detect broken stickiness Log the arm on every exposure, keyed by the identifier. Then a single query answers the question: does any identifier appear under more than one arm within one experiment? In a healthy platform that count is zero or explainable. A nonzero and growing count means either the assignment function is not deterministic across services (different hash implementations, or a salt that differs by deployment), or the identifier being hashed is not as stable as assumed. ## What to say in an interview State the rule, then the mechanism, then the consequence of breaking it. `One user, one arm, for the life of the experiment; we get it from a deterministic hash rather than stored state; if it breaks, users are partly in both arms and the effect estimate is biased toward zero rather than merely noisy.` That last clause is what separates a memorised answer from an understood one.
- Why derive the assignment from a hash instead of storing each user's coin flip in a table?A hash is stateless and instant: any service recomputes the same arm from the identifier alone, with no database read on the request path and nothing to replicate or lose. A stored flip needs a write on first exposure and a read on every later one, and two services can silently disagree about who is where. The hash is also reproducible offline, so analysis can rebuild assignments from logs.
- Does sticky assignment mean a user's variant can never change?It means it never changes on its own. Assignment moves when you deliberately re-salt the experiment, or when you widen the exposed fraction so previously unexposed users become eligible - and widening is safe because it adds users without moving assigned ones. What stickiness rules out is an independent draw on each request, which is what destroys the comparison.
- How would you detect that stickiness is broken in production?Log the assigned arm on every exposure alongside the identifier, then count identifiers that appear under more than one arm inside a single experiment. A nonzero, growing count means the assignment function is not deterministic everywhere - different hash implementations across services, or a salt that differs by deployment - and the experiment's data is already contaminated.
It is like a drug trial in which each patient takes the real pill on some days and the placebo on others. At the end nobody was purely in either group, so the difference between the groups has been averaged away.
saying these in an interview costs you the question
- Says a fresh draw per request is fine because it averages out
- Treats stickiness as a UX nicety rather than a statistical requirement
- Believes contaminated users only add noise, never bias
- Claims re-randomising mid-test makes the sample more representative
- Assumes contamination inflates the effect rather than shrinking it