Why is randomizing an A/B test per request instead of per user usually wrong?
answer
- who gets one coin flip?
- same person, two page loads
- both arms contain the same blended experience
- effect pulled toward zero, attribution impossible
- acceptable only if invisible, stateless, request-scoped metric
basics
~20 sPer-request randomization re-rolls the assignment on every page load, so one person sees both variants. That breaks the experience and mixes the arms, so the comparison is between two blends rather than treatment versus control.
solid answer
~50 sThe unit of randomization should be the thing whose whole experience you want to change. If a pricing test is randomized per request, the same shopper can see one price on the product page and a different price after a refresh — visibly inconsistent, and in a pricing context a real trust and compliance problem. Statistically it is worse: nearly every shopper receives a mix of both variants, so the "treatment" group and the "control" group contain almost the same people having almost the same experience. The measured difference is pulled toward zero, and an outcome like a purchase cannot be attributed to one arm at all. Randomize on a stable identifier so the assignment sticks. Per-request randomization is only defensible when the change is invisible and stateless to the person and the outcome itself is per-request — for example a backend routing change measured by response latency.
go deeper
Be ready to say plainly what gets one coin flip in your experiment, and to notice when a design would let one person see both variants on consecutive page loads.
Explain the mechanics of dilution: if everyone receives roughly half of each variant, both arms hold nearly the same experience and the estimated difference shrinks toward zero.
Show judgment about when the finer unit is safe — invisible, stateless changes with request-scoped metrics — and be able to reject a proposed design on user-experience grounds, not only statistical ones.
Own the guardrail: decide which unit is the platform default, and what review a team must pass before shipping anything finer than a person-level split.
## What the unit of randomization is An experiment splits a population into arms and compares outcomes. The *unit of randomization* is the entity that gets a single coin flip: whatever that entity is, everything belonging to it goes into one arm. It can be a request, a session, a person, a device, or a whole cluster such as a store or a city. Choosing it is the first design decision in an experiment, because it determines what a "replicate" is, what experience each person actually receives, and which metrics are even definable. ## Why per-request assignment breaks the experience With per-request randomization, each HTTP request draws its own coin. A shopper loading a product page gets variant A, refreshes, and gets variant B. Concretely: a pricing test randomized per request can show the same shopper 19.00 on one load and 24.00 on the next. Beyond the statistics, that is a product failure — it looks broken, it destroys trust, and for price it may be legally fraught. The same problem appears for any stateful change: a redesigned checkout that appears and disappears between steps, a tutorial that starts and vanishes, a recommendation strip that flickers. ## Why it breaks the measurement The deeper problem is contamination. Suppose a person makes 20 requests during a visit. Under fair per-request assignment, that person sees roughly 10 requests of each variant. Everyone in the study therefore receives a mixture close to 50/50, and the two arms are populated by essentially the same blended experience. What you estimate is no longer "the effect of B versus A" but something closer to the difference between two nearly identical mixtures, which is close to zero regardless of how good B is. This is dilution, or attenuation: a real effect is systematically shrunk toward the null, so the experiment can conclude "no difference" for a change that actually works. Attribution collapses as well. If a purchase happens at the end of a 20-request journey that included both variants, which arm gets credit? Any rule you invent — last request wins, first request wins — is arbitrary and introduces its own bias. Finally, there is the analysis problem. Requests from the same person are not independent draws: a single heavy user contributes hundreds of correlated requests. Treating them as independent observations understates variability and produces standard errors that are too small. ## When per-request randomization is legitimate It is fine — often the right choice — when three conditions hold together: 1. **The change is invisible to the person.** A different backend index, a different cache tier, a different retrieval implementation that returns equivalent results. 2. **The change is stateless.** Nothing carries over from one request to the next; seeing variant B once does not alter how the person responds to variant A later. 3. **The outcome is measured per request.** Latency, error rate, cost per call, server CPU. These are properties of the request, not of the person. Under those conditions per-request assignment is attractive precisely because it produces an enormous number of units, which gives high power for small latency differences, and because it balances traffic mix continuously over time. ## The rule to state in an interview Randomize at the coarsest unit that the treatment can affect, and never finer than the unit whose behaviour the metric describes. If the metric is "did this person convert", the unit must be the person. If the metric is "how long did this request take", the request can be the unit. When in doubt, randomize by a stable user or device identifier: it is the safe default because it keeps the experience coherent, keeps person-level metrics definable, and avoids self-contamination. ## A quick diagnostic Ask: can one individual end up in both arms? If yes, and the treatment is something they can perceive or remember, the unit is too fine. That single question catches most unit-of-randomization mistakes before an experiment ships.
- Give a case where randomizing per request is actually the right design.A backend change that is invisible to the person — a new search index, a different cache tier — measured by per-request latency or error rate. Nothing carries over between requests, the person cannot perceive which path served them, and the outcome is a property of the request itself. The huge number of units gives excellent power for small latency effects.
- If per-request assignment dilutes the effect, is the resulting p-value at least conservative?Not reliably. The point estimate is attenuated, which pushes toward false negatives, but the standard error is usually understated too because requests from one person are correlated and are being counted as independent. Two errors in opposite directions do not cancel in any controlled way, so neither the estimate nor its uncertainty can be trusted.
It is like testing two recipes by swapping the chef mid-meal on every bite: everyone tastes both dishes, so the two menus end up indistinguishable.
saying these in an interview costs you the question
- Thinks more requests always means more statistical power
- Claims mixing the arms averages out and is harmless
- Cannot say which entity received a single assignment
- Assumes any randomization is valid because it is random
- Ignores that the same person now sees two different prices