skip to content

Randomization & Assignment

Choosing what actually gets randomized — user, session, or cluster — and making the split trustworthy with sticky hash bucketing, holdouts and A/A runs. Every design question starts here.

on this pageshow

explore

questions

22

Why must an A/B test give the same user the same variant on every visit?

level: juniorimportance: must knowfreq 72%

answer

  1. same person, same experience
  2. re-rolling blends the two arms
  3. returning users become partly treated
  4. estimate pulled toward zero, not noisier
  5. deterministic hash, no stored coin flip

basics

~20 s

Sticky assignment gives each user one consistent experience instead of a flickering mix. If the variant is re-drawn per visit, returning users receive both treatments, which blends the two arms and shrinks the measured difference toward zero.

solid answer

~50 s

Assignment has to be sticky for two reasons. The product reason is obvious: a user who sees the new checkout on Monday and the old one on Tuesday gets a confusing experience, and that inconsistency itself changes behaviour. The statistical reason is worse. If the variant is re-drawn on each visit, every returning user is part treated and part control, so the two arms stop being two different populations and become the same blended population measured twice. The estimated effect is diluted toward zero, and you ship a verdict of `no difference` for a feature that actually works. In practice stickiness comes from deriving the assignment from a hash of a stable identifier rather than storing a coin flip, so the same identifier always maps to the same arm, in every service, with no lookup.

go deeper

for a junior

Be ready to state the rule plainly: one user, one variant, for the whole experiment. Know that the arm is normally computed from the user's identifier rather than drawn fresh at request time.

for a middle

Explain the failure mechanically: users who see both variants are partly treated, so the arms blend and the estimated effect shrinks toward zero rather than merely getting noisier. Describe how a deterministic hash gives stickiness with no stored state.

for a senior

Expect to be asked how you would catch broken stickiness in production - logging the assigned arm on every exposure and alerting when one identifier appears under two arms in the same experiment.

for a principal

Own the platform rule: assignment is computed once from a documented identifier and is never re-drawn during a running test. Decide which exceptions exist, such as a deliberate re-salt between experiments, and who may authorise them.

## What sticky assignment means An online experiment splits users into arms and shows each arm a different version of the product. **Sticky assignment** means that once a user is placed in an arm, they stay in that arm for the whole experiment: every visit, every device session in which they are recognisable, every request. The opposite is a fresh random draw each time the user shows up, which sounds harmless and is not. ## The product argument A product that changes shape between visits is a bad product. A user who found a feature yesterday and cannot find it today files a bug, contacts support, or simply stops trying. Worse, the inconsistency is itself an intervention: you are no longer measuring `new design vs old design`, you are measuring `stable experience vs unstable experience`, which is not a shipping decision anyone wants to make. ## The statistical argument This is the one interviewers are usually probing for. An experiment estimates a difference between two groups that differ in exactly one thing: which version they received. If assignment is re-drawn per visit, a user with ten visits receives roughly five of each. Averaged over their behaviour, that user is now half treated. Do this to everyone and both arms contain the same mixture, so the difference between arm means collapses. The practical consequence is **attenuation**: the estimated effect is pulled toward zero, roughly in proportion to how much cross-contamination there is. The confidence interval does not widen to warn you; it narrows around the wrong value, because you still have plenty of observations. So the failure mode is not noisy results, it is confidently wrong results. A real 3% lift is measured as 0.4%, declared non-significant, and the feature is killed. Note which direction this goes. Contamination almost always biases toward `no effect`, so it produces false negatives rather than false positives. That is why it survives so long undetected: nothing looks broken, experiments just quietly stop finding anything. ## How stickiness is implemented The standard mechanism is deterministic: take a stable identifier for the randomisation unit, combine it with something specific to the experiment, hash the result, and map the hash into an arm. Because the function is pure, any service can recompute the same answer from the identifier alone. There is no assignment table to read on the request path, nothing to replicate between regions, nothing to lose, and the analysis job can rebuild every assignment from logs after the fact. The alternative — flip a coin the first time you see a user and store the result — also gives stickiness in principle, but it buys it with state. That state has to be written on first exposure, read on every subsequent one, kept consistent across services, and it becomes a source of outages and of silent divergence when two systems disagree about who is in which arm. ## What legitimately changes an assignment Stickiness does not mean an assignment is frozen forever. It means it never changes *spontaneously*. Assignments legitimately move when you deliberately re-salt an experiment (which reshuffles everyone and must therefore happen between experiments, never during one), or when the exposed fraction of the population is deliberately widened so that users who were previously unexposed become eligible. Widening exposure is safe precisely because it adds users without moving the ones already assigned. ## How you detect broken stickiness Log the arm on every exposure, keyed by the identifier. Then a single query answers the question: does any identifier appear under more than one arm within one experiment? In a healthy platform that count is zero or explainable. A nonzero and growing count means either the assignment function is not deterministic across services (different hash implementations, or a salt that differs by deployment), or the identifier being hashed is not as stable as assumed. ## What to say in an interview State the rule, then the mechanism, then the consequence of breaking it. `One user, one arm, for the life of the experiment; we get it from a deterministic hash rather than stored state; if it breaks, users are partly in both arms and the effect estimate is biased toward zero rather than merely noisy.` That last clause is what separates a memorised answer from an understood one.

  • Why derive the assignment from a hash instead of storing each user's coin flip in a table?
    A hash is stateless and instant: any service recomputes the same arm from the identifier alone, with no database read on the request path and nothing to replicate or lose. A stored flip needs a write on first exposure and a read on every later one, and two services can silently disagree about who is where. The hash is also reproducible offline, so analysis can rebuild assignments from logs.
  • Does sticky assignment mean a user's variant can never change?
    It means it never changes on its own. Assignment moves when you deliberately re-salt the experiment, or when you widen the exposed fraction so previously unexposed users become eligible - and widening is safe because it adds users without moving assigned ones. What stickiness rules out is an independent draw on each request, which is what destroys the comparison.
  • How would you detect that stickiness is broken in production?
    Log the assigned arm on every exposure alongside the identifier, then count identifiers that appear under more than one arm inside a single experiment. A nonzero, growing count means the assignment function is not deterministic everywhere - different hash implementations across services, or a salt that differs by deployment - and the experiment's data is already contaminated.

It is like a drug trial in which each patient takes the real pill on some days and the placebo on others. At the end nobody was purely in either group, so the difference between the groups has been averaged away.

saying these in an interview costs you the question

  • Says a fresh draw per request is fine because it averages out
  • Treats stickiness as a UX nicety rather than a statistical requirement
  • Believes contaminated users only add noise, never bias
  • Claims re-randomising mid-test makes the sample more representative
  • Assumes contamination inflates the effect rather than shrinking it

context

open as a page

What do experiment layers do in an A/B platform running many tests at once?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Layers let many experiments share the same traffic at the same time. Each layer randomizes users independently, so a user's arm in one layer says nothing about their arm in another, and concurrent tests do not confound one another.

open as a page

Why ramp an A/B test through 1%, 5%, 25% and 50% instead of starting at full traffic?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A staged ramp limits blast radius: the first small stage exposes few users, so crashes, latency spikes or a collapsing metric are caught cheaply. Later, larger stages exist to measure the effect once the change has been shown to be safe.

open as a page

How does hash-based bucketing turn a user ID into an A/B test variant?

level: middleimportance: must knowfreq 68%

basics

~20 s

Concatenate the user ID with an experiment-specific salt, hash it to an integer, take it modulo the bucket count, and map bucket ranges to arms - 0-49 control, 50-99 treatment. Deterministic, so nothing is stored.

open as a page

When would you randomize an A/B test by session rather than by user?

level: middleimportance: must knowfreq 70%

basics

~20 s

Randomize by session only when the change leaves no trace between visits and the outcome is finished within one visit — a layout tweak judged by same-visit conversion. If the treatment can be learned, remembered or accumulated, randomize by user.

open as a page

An A/B variant read +12% lift at a 1% ramp and +2% at 50% — what explains the drop?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Most likely the winner's curse. The 1% estimate was extremely noisy, and the feature was promoted because that noisy number looked big. Selecting on a large observed effect inflates it, so the precise later read of +2% is the better estimate, not a decay.

open as a page

You randomize by user but run the t-test over page-views — why is that p-value wrong?

level: seniorimportance: must knowfreq 60%

basics

~20 s

The test counts page-views as independent, but page-views from one person share that person's single assignment. The effective sample size is the number of users, so the standard error is understated and false positives run far above the nominal rate.

open as a page

Why is randomizing an A/B test per request instead of per user usually wrong?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Per-request randomization re-rolls the assignment on every page load, so one person sees both variants. That breaks the experience and mixes the arms, so the comparison is between two blends rather than treatment versus control.

open as a page

What does an A/A test validate before you launch a real experiment?

level: middleimportance: should knowfreq 58%

basics

~20 s

An A/A test splits traffic normally but serves both arms the identical experience. It validates the plumbing - a split that produces comparable groups, a metric pipeline that agrees on identically treated users, and an analysis that does not flag phantom differences.

open as a page

Why does each experiment layer hash users with its own independent seed?

level: middleimportance: should knowfreq 54%

basics

~10 s

Independent seeds make assignment in one layer statistically independent of every other layer. Reusing one seed sends the same users to the same arm position in both layers, perfectly confounding the two experiments.

open as a page

When should two experiments be mutually exclusive rather than overlapping in separate layers?

level: middleimportance: should knowfreq 44%

basics

~10 s

Make two experiments mutually exclusive when their treatments cannot sensibly coexist for one user, such as competing redesigns of the same page. Everything else should overlap, because exclusivity costs traffic and time.

open as a page

How do logged-out cookies and cross-device logins corrupt user-level A/B assignment?

level: middleimportance: should knowfreq 48%

basics

~20 s

You never randomize a person, only an identifier. A device cookie splits one person across a phone and a laptop into two units that can land in opposite arms, and signing in mid-visit can switch the arm entirely.

open as a page

An A/A run flags 1 of 20 metrics at alpha 0.05 - do you block the launch?

level: seniorimportance: should knowfreq 47%

basics

~20 s

No. Scanning 20 metrics at a 0.05 threshold when nothing truly differs gives one flag on average, and at least one flag about 64% of the time. Investigate the flagged metric, but a single borderline result is not evidence of a broken split.

open as a page

Your new test reuses buckets 0-49, which held last quarter's shipped treatment - what breaks?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Those users carry a residual effect from the old treatment - a quarter of exposure changed their habits and their composition. Concentrating them in one arm makes the arms differ before the test starts. Re-salt to spread them.

open as a page

A checkout-button test wins only when a concurrent shipping-price test is in treatment. How do you diagnose that?

level: seniorimportance: should knowfreq 46%

basics

~10 s

Ask first whether the pattern is real: the four-cell contrast measuring an interaction carries about twice a main effect's standard error, so flips are usually noise. Then rule out correlated assignment and eligibility differences.

open as a page

Should data from earlier ramp stages be pooled into the final effect estimate of an experiment?

level: seniorimportance: should knowfreq 41%

basics

~10 s

Only when every stage estimated the same thing: unchanged treatment, unchanged eligible population, unchanged metric. Even then, pool by combining per-stage effect estimates, never by summing raw counts across stages with different traffic splits.

open as a page

In a store-level A/B test, why is your sample size the number of stores, not shoppers?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Every shopper in a store receives the same assignment and the same local conditions, so their outcomes move together. The independent replicates are the stores, and adding shoppers inside a store adds far less information than adding another store.

open as a page

After shipping a feature, is a permanent 1% holdback kept for a quarter worth its cost?

level: principalimportance: should knowfreq 30%

basics

~20 s

Sometimes, and rarely per feature. A holdback buys continued measurement after the experiment ends, but 1% is insensitive and every holdback forks the product. A single company-level holdback covering many launches usually beats one per feature.

open as a page

When ramping an experiment from 1% to 25%, should the original 1% users keep their assignment?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

Yes, by default. Widen who is eligible while leaving already-assigned users exactly where they are. Re-drawing assignment mid-ramp flips some users between arms, and their later behaviour still carries what they already saw, which contaminates both arms.

open as a page

How do you decide whether to keep a permanent 10% global holdout excluded from every experiment?

level: principalimportance: nice to knowfreq 32%

basics

~10 s

Weigh one clean never-treated reference population against three costs: every experiment now draws from 90% of traffic, every shipped feature must keep its old code path alive, and held-out users get a stale product.

open as a page

How do you set an experimentation platform's policy on which tests may overlap?

level: principalimportance: nice to knowfreq 28%

basics

~10 s

Default to overlapping everything and grant exclusivity per shared surface, never per team request. Monitor for interactions retrospectively rather than gating launches, and state openly that reported effects average over the concurrent landscape.

open as a page

As an experimentation lead, how do you set the default randomization unit for every team?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Make the most stable identity the default, record the chosen unit as a first-class property of every experiment, derive the analysis grain from it automatically, and require review before any team randomizes finer than a person.

open as a page