skip to content

When ramping an experiment from 1% to 25%, should the original 1% users keep their assignment?

level: middleimportance: nice to knowfreq 33%

answer

  1. expand, do not redraw
  2. a flipped user carries the old experience
  3. contamination pulls the estimate toward zero
  4. arm counts look perfectly fine
  5. changed treatment means a new experiment

basics

~20 s

Yes, by default. Widen who is eligible while leaving already-assigned users exactly where they are. Re-drawing assignment mid-ramp flips some users between arms, and their later behaviour still carries what they already saw, which contaminates both arms.

solid answer

~50 s

The default is to expand the exposed range without disturbing existing assignments: users already in treatment stay in treatment, users already in control stay in control, and the new traffic is randomized into both arms. A fresh randomization at each stage moves some users across arms, and those users bring the experience they already had with them — a user who used the new flow for a week and then lands in control is not a control user any more. That contamination pulls the measured difference toward zero and is invisible in the arm counts. It also breaks per-user metrics that span stages. The one legitimate reason to draw fresh assignments is that you want the later stage to be an independent replication, usually because the treatment itself changed materially; even then the clean move is to exclude previously exposed users rather than to reassign them.

go deeper

for a junior

Know the default: users already in the experiment keep their arm, and ramping means adding new users, not reshuffling old ones.

for a middle

Be ready to explain the mechanism — a reassigned user carries the experience they already had — and to state that the resulting bias pulls the measured effect toward zero.

for a senior

Show that you would catch this in review: arm counts and pipeline checks look clean under reassignment, so the only defence is knowing how the ramp step was implemented before you read the numbers.

for a principal

Own the platform-level rule: assignment expansion should be the only supported ramp operation, with a deliberate new-experiment path for redesigned treatments, so individual teams cannot silently reassign users.

## The question behind the question When a ramp moves from 1% to 25%, somebody has to decide what happens to the users who were already in the experiment. There are two options: **expand** (keep every existing assignment and randomize only the newly eligible traffic) or **re-randomize** (throw away the old assignment and draw the whole 25% afresh). The default is expand, and knowing why is the point of the question. ## Why re-randomizing mid-ramp is damaging Randomization gives you comparable groups; it does not give you comparable *experiences* once people have already lived in an arm. Suppose a user spent a week in treatment during the 1% stage, learned the new flow, and is then reassigned to control. Their subsequent behaviour reflects a week of treatment: they may keep using a habit the new design taught them, or they may be annoyed that something they got used to has vanished. That user is now a control-labelled unit carrying part of the treatment effect. The statistical consequence is straightforward. The estimate you compute is the difference between what the treatment-labelled group did and what the control-labelled group did. If a slice of the control group has already received treatment, the two labelled groups are less different than the underlying conditions are, so the estimated effect shrinks toward zero — an attenuation bias. The reverse flow (control users moved into treatment) does not cancel it; it adds a second group whose behaviour is a mixture. Worse, nothing in the arm counts or in a routine data check reveals the problem: both arms are exactly the size you asked for, and the metric pipeline is healthy. The damage is entirely in the interpretation. There are practical harms too. Any user-level metric that spans stages — retention over 14 days, cumulative spend, sessions per user — becomes ambiguous, because the user belonged to two arms during the measurement window. Anyone comparing a stage-2 report against a stage-1 report sees numbers that moved for reasons that have nothing to do with the feature. ## What expanding costs you Expanding is the right default, but it is not free of consequences, and a strong answer names them. The first is **exposure heterogeneity**. Users who joined at 1% have been in treatment far longer than users who joined at 25%. The treatment arm at the end of the ramp is a mixture of long-exposed and freshly exposed users, and the reported average is weighted by that mixture. It is a real property of the estimate, not a bug, but it means the number answers "what is the average effect over this mix of exposure lengths", not "what is the effect on a steady-state user". The second is **stage dependence**. Because the same users persist across stages, the stage-2 estimate is not an independent replication of the stage-1 estimate — the earlier sample is a subset of the later one. Treating consecutive stage readings as separate confirmations of the same result double-counts the overlapping users. ## When a fresh randomization is legitimate There is one case that genuinely justifies breaking assignments: the intervention changed. If the treatment was materially redesigned between stages — a different layout, a fixed algorithm, a different price — then the earlier stage measured a different thing, and pooling it with the later stage measures neither cleanly. The correct move is still not reassignment: it is to **start a new experiment** with a fresh randomization over users who were never exposed to the earlier version, or at minimum to exclude previously exposed users from the new estimate and note the reduced generalisability. A weaker but real case is a ramp that stayed at a tiny fraction for a very long time with a narrow eligible population; when eligibility itself is redefined (a new country, a new platform), you are effectively running on a different population and a clean restart is easier to reason about than a patched-together pooled estimate. ## How to answer in an interview Say the default (expand, keep assignments sticky), explain the contamination mechanism and its direction (attenuation toward zero), note that arm counts will not reveal it, and then name the exception (the treatment changed) with the correct remedy (a new experiment on unexposed users, not a reassignment). That sequence — default, mechanism, direction of bias, exception — is what separates someone who has run ramps from someone who has read about them.

  • Which direction does contamination from reassignment bias the estimated effect?
    Toward zero. Control-labelled users who already experienced the treatment behave partly like treated users, so the gap between the two labelled groups understates the gap between the underlying conditions. The attenuation is not symmetric-cancelling: moving users the other way simply adds a second mixed group. A real effect can end up looking small or inconclusive purely from the reassignment.
  • The treatment was materially redesigned between the 5% and 25% stages. What now?
    The 5% stage measured a different intervention, so it is not evidence about the current one. Treat the redesign as the start of a new experiment: estimate from the post-change stage only, and prefer users who were never exposed to the old version, since previously exposed users can carry habits from it. Say so explicitly in the write-up rather than quietly pooling.
  • If assignments stay sticky, why are consecutive stage results not independent confirmations?
    The earlier stage's users are a subset of the later stage's users, so the two estimates share data. Seeing the same direction twice is largely the same observation reported twice, not replication. If you want an independent check, it has to come from users who were not in the earlier stage, analysed as their own group.

saying these in an interview costs you the question

  • Re-randomizes at every stage to keep it fresh
  • Thinks equal arm sizes prove no contamination
  • Assumes flipped users cancel each other out
  • Pools stage readings as independent replications
  • Keeps old assignments after the treatment was rebuilt

context