skip to content

When would you randomize an A/B test by session rather than by user?

level: middleimportance: must knowfreq 70%

answer

  1. does the change leave a trace?
  2. returning visitor gets a fresh coin flip
  3. learning and habit contaminate the later arm
  4. retention is undefined if the person was in both arms
  5. more units, but correlated within person

basics

~20 s

Randomize by session only when the change leaves no trace between visits and the outcome is finished within one visit — a layout tweak judged by same-visit conversion. If the treatment can be learned, remembered or accumulated, randomize by user.

solid answer

~50 s

Session randomization draws a fresh assignment each visit, so a returning person can be in treatment on Monday and control on Tuesday. That is acceptable when the change has no memory: a same-visit layout or copy tweak whose effect is over when the visit ends. It buys more units and therefore more power, and it can balance a metric that varies strongly across visits. It is wrong whenever carryover exists. A session-randomized onboarding tutorial is the classic failure: someone who completed the tutorial on their first visit is already taught on their second, so their control-arm behaviour is contaminated by treatment, and the estimated effect is diluted. It also makes person-level metrics meaningless — you cannot ask whether a user retained for a week if that user was in both arms during the week. And sessions from one person are correlated, so they cannot be treated as independent observations without care.

go deeper

for a junior

Know that session randomization re-rolls the assignment on each visit, so the same person can land in different arms on different days.

for a middle

Explain the two mechanics that decide the choice: carryover contaminating the later arm, and person-level metrics being undefined when a person spans both arms.

for a senior

Diagnose a shrinking effect curve as accumulating contamination rather than novelty, and know that sessions from one person are correlated and cannot be counted as independent.

for a principal

Frame the two designs as different estimands — effect per visit versus effect per person — and decide which one the shipping decision actually needs.

## The choice User randomization assigns a person once and keeps them in that arm for the life of the experiment. Session randomization assigns each visit independently, so the same person can move between arms across visits. Both are legitimate designs; they answer different questions and carry different risks. ## The case for sessions **More units.** A population of 100,000 people may produce 400,000 visits in the same period. Because the precision of a comparison improves roughly with the square root of the number of independent units, a valid session-level design can detect a smaller difference in the same calendar time — but only if the sessions really are close to independent, which they are not when a small group of people supplies most of the visits. **Tighter scoping.** If the change only exists inside a visit — a banner, a sort order, a page layout — and the outcome is completed inside the same visit, the session is the natural experimental frame. **Traffic mix.** Session-level splits balance day-of-week and campaign mix continuously, so both arms see similar traffic composition even if the population shifts. ## Carryover: the reason to prefer users Carryover means an earlier exposure changes later behaviour. It quietly destroys session designs. Take a session-randomized onboarding tutorial. On visit 1, a person is assigned treatment and completes the tutorial. On visit 2 they are re-randomized into control, and they navigate the product fluently because they were taught last time. Their control-arm session now looks good for reasons that belong to the treatment. Aggregated over many returning people, the control arm absorbs a share of the treatment's benefit and the measured gap shrinks toward zero. Worse, the contamination grows over the life of the experiment, so the effect appears to fade — a pattern that is easy to misread as novelty wearing off. Carryover comes in many forms: learning (the tutorial), habit (a new default the person now expects), state (a saved setting or a completed profile), and annoyance (a nagging prompt they already dismissed). Any of these makes session randomization invalid. ## What metrics you can no longer define A person-level metric requires the person to belong to exactly one arm. Retention ("did this user return within 7 days"), lifetime value, subscription conversion and complaint rate are all person-level. If a user visited three times with two treatment sessions and one control session, there is no honest way to place them in an arm, and any rule you pick induces bias. Session randomization therefore restricts you to session-scoped outcomes: same-visit conversion, visit length, task completion within the visit. ## Independence, and the analysis unit Even when carryover is genuinely absent, sessions from the same person are correlated: heavy users contribute many sessions with similar behaviour. Treating each session as an independent observation understates the true variability of the comparison and produces confidence intervals that are too narrow. If the same person can appear in both arms, that correlation also links the two arms, which is not something a simple two-sample comparison expects. The safe habit is to make the analysis unit match the randomization unit; when it cannot, the correlation has to be handled explicitly rather than ignored. ## The inconsistency people can see Session designs also create visible inconsistency: someone who liked yesterday's checkout finds today's different. For small cosmetic changes this is tolerable; for anything a person invests effort in, it is a product defect, and support tickets about "the site keeps changing" are a real cost. ## A decision procedure 1. Can an earlier exposure change later behaviour? If yes, randomize by user. 2. Is the primary metric defined over a person rather than a visit? If yes, randomize by user. 3. Would a returning person notice the switch and be bothered? If yes, randomize by user. 4. Otherwise session randomization is available, and its extra units are a genuine gain — provided the analysis respects that one person's sessions are correlated. ## What good candidates add Strong answers point out that the choice is not just statistical hygiene: user and session randomization estimate different things. A user-level design estimates the effect of *being the kind of person who always gets this experience*, which is what shipping actually does. A session-level design estimates the effect of *this visit having the feature*, which is a narrower and sometimes less decision-relevant quantity.

  • A session-randomized test shows the effect shrinking week by week. What is your first hypothesis?
    Accumulating carryover. As the experiment runs, more of the control-arm sessions come from people who already experienced treatment on an earlier visit, so control drifts toward treatment and the gap closes. Check whether the shrinkage is concentrated among returning visitors; if first-time visitors show a stable effect, contamination rather than a fading novelty effect is the explanation.
  • Why can a session-randomized test not report seven-day retention?
    Retention is a property of a person over a week, but a session design can place that person in treatment on one visit and control on another. There is no non-arbitrary arm to attribute the person to, and any tie-break rule — first session, most sessions — correlates with how often they visit, which is itself an outcome. The metric is simply not identified under that design.
  • Do session and user randomization estimate the same quantity when carryover is genuinely absent?
    They are close but not identical. Session randomization weights people by how often they visit, so heavy users influence the result more; user randomization weights each person equally. With no carryover both are unbiased for their own estimand, but the session-weighted number answers "what happens to a typical visit" while the user-weighted number answers "what happens to a typical person".

saying these in an interview costs you the question

  • Picks sessions purely because it yields more rows
  • Ignores that a returning visitor switches arms
  • Reports retention from a session-randomized experiment
  • Treats each session as an independent observation
  • Assumes carryover is only about novelty effects

context