How do logged-out cookies and cross-device logins corrupt user-level A/B assignment?
answer
- you randomize an id, not a human
- phone and laptop are two units
- arm can flip at the sign-in step
- contamination attenuates toward zero
- one identity per experiment, carried across login
basics
~20 sYou never randomize a person, only an identifier. A device cookie splits one person across a phone and a laptop into two units that can land in opposite arms, and signing in mid-visit can switch the arm entirely.
solid answer
~50 sUser-level randomization is really identifier-level randomization, and identifiers are leakier than people. A logged-out visitor is bucketed by a device cookie, so one person on a phone and a laptop is two units and can be in treatment on one and control on the other; experiencing both variants dilutes the effect. Cookies also churn — cleared, expired, absent in private browsing — creating fresh units mid-experiment whose history is lost. The sharpest problem is the login boundary: if the cookie bucket and the signed-in bucket disagree, the person changes arm in the middle of a funnel and the conversion is attributed to whichever identity was active. Practical answers: pick one identifier for the whole experiment and never switch mid-flow; carry the logged-out identifier forward at login; and if the test is genuinely about logged-out flows, accept the device as the unit, keep the horizon short, and report device-level metrics honestly.
go deeper
Know that the experiment buckets an identifier, not a human, and that one person with two devices can end up in both arms.
Explain the three concrete failures — multi-device duplication, identifier churn, and the arm changing at the login boundary — and what each does to the estimate.
Demonstrate the operational fix: choose one identity for the whole run, carry it across login, keep logged-out horizons short, and quantify leakage from the signed-in population.
Own the identity policy across the platform, including which identifier is canonical, and set the standard for reporting the randomization unit alongside every result.
## The identifier is not the person Every user-level experiment assigns arms by hashing some identifier. The design is only as good as the correspondence between that identifier and a human being, and in practice the correspondence is loose in both directions: one person can hold many identifiers, and one identifier can be used by several people. **One person, many identifiers.** Phone, laptop, work machine, a second browser, private browsing. Each carries its own cookie, each is randomized independently, and with a 50/50 split a person with two devices lands in both arms roughly half the time. **One identifier, many people.** A shared family tablet or a kiosk pools several people under a single assignment. Their behaviour is attributed to one unit. **Identifiers that expire.** Cookies get cleared, expire, or are dropped by privacy settings. When that happens, the person reappears as a brand-new unit with a fresh coin flip and no history, so the population under study quietly turns over during the run. ## What each failure does to the numbers **Cross-device contamination dilutes.** Someone who sees the new checkout on their laptop and the old one on their phone has partly received both treatments. The two arms become more alike, so the estimated difference shrinks toward zero. The direction is predictable: contamination attenuates. An experiment can look like a null result when the underlying effect is real. **Churn adds noise and shortens exposure.** New identifiers appear mid-run with no accumulated exposure, so the average dose of treatment across the arm is lower than the design assumed. It also breaks any metric that needs history, such as "return within 7 days", because the returning visitor is not recognisable as the same unit. **The login boundary can flip the arm.** A visitor browses logged out under cookie C, assigned to treatment. They sign in as account U, whose independent assignment is control. If the system re-evaluates the arm on login, the interface changes underneath them mid-funnel and the eventual purchase belongs to neither arm cleanly. This is the most damaging of the three because it happens exactly where conversions are recorded. ## Designs that handle it **Pick one identity per experiment and hold it.** Decide up front whether the unit is the device identifier or the account identifier, and evaluate the assignment from that identifier for the entire experiment. Never let the effective unit change while a person is mid-flow. **Carry the identifier across login.** If the experiment must span the login boundary, keep using the logged-out identifier after sign-in for the duration of the run, so the person stays in the arm they started in. The pre-login and post-login experience then stays consistent, which is usually what the product needs anyway. **Scope the experiment to one side of the boundary.** Signed-in-only experiments randomize on the account identifier, which is stable across devices and immune to cookie churn. Logged-out experiments accept the device as the unit. Mixing the two in one test is where most of the pain comes from. **Keep logged-out horizons short.** Because device identifiers decay, a logged-out experiment measuring a two-week outcome is measuring a population that partially reset. Same-visit or few-day outcomes are far more trustworthy under a cookie unit. ## Interpreting the result honestly Under a device-level design, the comparison is still a fair comparison — it is between devices assigned to each arm, and randomization guarantees the arms are comparable. What it is *not* is an unbiased estimate of the per-person effect: multi-device people are partly treated in both arms, so the per-person effect is larger in magnitude than the number you measured. Saying this out loud is a mark of a strong candidate: the test remains valid for the unit you actually randomized, and the attenuation is a known, one-directional bias when you translate the result to people. ## The diagnostic questions When reviewing someone's design, ask: what exactly is hashed to choose the arm? Can that value change during the person's journey? What fraction of the traffic crosses the login boundary during the funnel? How fast do the identifiers turn over? Answers to those four questions predict almost every identity-related surprise before the experiment starts.
- Which direction does cross-device contamination push the estimated effect?Toward zero. People who experience both variants respond to a blend, so the arms become more similar than the design intended and the measured gap shrinks. The practical implication is asymmetric: a significant result survives the bias, because the true per-person effect is at least as large, while a null result is genuinely ambiguous between no effect and a contaminated one.
- How would you estimate how badly cross-device leakage affects a given experiment?Use the signed-in portion of the traffic, where the same account can be linked across devices, to measure what share of accounts appear on more than one device within the experiment window and how often those devices sit in opposite arms. That share, applied to the split, bounds how much of the population received a mixed experience and therefore how much attenuation to expect.
- Is a device-randomized experiment invalid because devices are not people?No — it is a valid randomized comparison of devices, and the arms are comparable by construction. The care is in interpretation: person-level claims and person-level metrics like retention are not directly supported, and the per-person effect is understated. State the unit alongside the result rather than silently relabelling devices as users.
It is like mailing a coupon to addresses rather than people: someone with a home and an office address may get both versions, and someone who moves gets counted twice.
saying these in an interview costs you the question
- Assumes one cookie equals one person
- Re-evaluates the arm at the sign-in step
- Reports device counts as user counts
- Mixes cookie and account units in one test
- Treats a null result as clean despite heavy leakage