skip to content

Your new test reuses buckets 0-49, which held last quarter's shipped treatment - what breaks?

level: seniorimportance: should knowfreq 44%

answer

  1. buckets are people, not fresh draws
  2. same salt means the same grouping
  3. old exposure changed habits and who stayed
  4. swapping the labels changes only the sign
  5. new salt balances rather than erases

basics

~20 s

Those users carry a residual effect from the old treatment - a quarter of exposure changed their habits and their composition. Concentrating them in one arm makes the arms differ before the test starts. Re-salt to spread them.

solid answer

~50 s

Bucket membership is not fresh randomness. Because the hash is deterministic, reusing the range 0-49 reuses the same people, and those people spent a quarter inside the previous treatment. Two things followed from that. Their **behaviour** adapted - they learned the new flow, changed habits, moved their baseline. And their **composition** changed, because whoever disliked the old treatment churned out of that group and not out of the other. So the new experiment starts with two arms that already differ, and any measured effect is that inherited difference plus whatever the new feature does. The fix is to re-salt: hash with a fresh experiment salt so bucket membership is unrelated to the old split. That does not erase carryover, it balances it - previously exposed users land in both arms in similar proportions, so their residual effect largely cancels in the difference.

go deeper

for a junior

Know that a bucket number is not a coin flip: with the same salt the same users sit in the same buckets, so reusing a range means reusing a specific group of people.

for a middle

Explain both inherited differences - the earlier treatment changed how those users behave and also which of them stayed - and that a new salt reshuffles who lands in each arm.

for a senior

Show the triage: confirm whether the salt is new, compare the proposed arms on a pre-period, and be explicit that re-salting balances carryover rather than erasing it.

for a principal

Own the platform default that every experiment receives a fresh salt, and decide the narrow cases where a deliberately reused split is the design - such as a follow-up measurement on a previously exposed cohort.

## Why bucket reuse is not a fresh randomisation Hash bucketing is deterministic on purpose: the same identifier plus the same salt always yields the same bucket. That is what gives sticky assignment. The corollary is that a bucket range is a **named group of people**, not a fresh draw. If last quarter's experiment used the same salt and mapped buckets 0-49 to its treatment, then buckets 0-49 today contain exactly the users who received that treatment - minus whoever left. So pointing a new experiment at the same range does not randomise anything. It reuses a cohort that has already been intervened on. ## The two inherited differences **Behavioural carryover.** Users who lived with a change for a quarter are not the same subjects they were before it. They learned the new navigation, formed a habit around the new default, or adjusted their expectations of what the product does. Their baseline metrics are permanently shifted relative to users who never saw it. When those users are all in one arm of the new test, that shift is confounded with the new treatment. **Compositional carryover (differential attrition).** If the old treatment pushed some users away - or attracted and retained others - the two former arms no longer contain comparable populations. Selection has already happened. Even if behaviour had fully reverted, the *set of people* differs, and comparing them measures the survivors of one experience against the survivors of another. Both effects point the same way: the arms are not exchangeable at the moment the new test starts, which is precisely the assumption randomisation exists to guarantee. ## Why relabelling does not help A tempting shortcut is to swap which range is called control: make 50-99 the control this time. This does nothing. The same concentration of previously exposed users sits in one arm; only the sign of the inherited difference changes. The imbalance is a property of *who is grouped together*, not of what you call the group. ## What actually fixes it **Re-salt.** Give the new experiment its own salt so the key being hashed is different. Because a good hash has avalanche behaviour, the new bucket for a user is unrelated to the old one, and previously exposed users are scattered roughly evenly across the new arms. Their residual behaviour still exists, but it now appears on both sides of the comparison and largely cancels in the difference. This is the standard reason platforms issue a fresh salt per experiment by default rather than letting teams pick bucket ranges by hand. Be precise about what re-salting buys you: **balance, not erasure**. If the previous change moved the whole population's baseline, the new experiment is being run on a different population than last quarter's was, and its result generalises to today's users only. That is a legitimate limitation to state, not a bug to fix. **Verify before launching.** Two cheap checks: - Compare the two proposed arms on historical metrics from a period *before* either experiment ran. If they already differ on engagement, spend or retention with no treatment applied, the split is inherited rather than fresh. - Run a short A/A on the exact split you intend to use. A null experiment on an inherited split will often surface the imbalance directly. ## When reuse is deliberate and fine Reusing a split is not always a mistake. If the explicit goal is a follow-up measurement on the same cohort - for example, checking whether last quarter's change still behaves the same way for the people who got it - then holding membership fixed is the design, not an accident. What makes it legitimate is that it is stated, and that the analysis is framed as a comparison of previously exposed versus never exposed rather than as a clean test of the new feature. ## What interviewers want to hear The chain in order: deterministic hash means bucket ranges are stable populations; a reused range therefore inherits a treated cohort; the inheritance is both behavioural and compositional; relabelling arms is not a fix; re-salting redistributes rather than removes; and a pre-period comparison or a short A/A catches it before launch. Candidates who only say `carryover effect` without explaining why the hash makes it inevitable have half the answer.

  • Does re-salting remove the carryover effect?
    No - it removes the imbalance. Previously exposed users still behave differently, but a fresh salt scatters them evenly across the new arms, so their residual effect appears on both sides and largely cancels in the difference. What re-salting cannot fix is a population-wide baseline shift; that changes the reference for everyone and limits how far the new result generalises.
  • Would swapping which bucket range is the control arm fix the problem?
    No. Swapping labels changes which arm carries the previously exposed users, not whether they are concentrated in one arm. The inherited difference stays the same size and merely flips sign. Only rehashing with a new salt actually redistributes membership.
  • How would you detect an inherited split before launching?
    Compare the two proposed arms on historical metrics from a period before either experiment ran, or run a short A/A on the exact split you intend to use. If the arms already differ on engagement or spend with no treatment applied, the split is inherited rather than fresh, and you re-salt before launching.

saying these in an interview costs you the question

  • Assumes each experiment gets fresh randomness automatically
  • Thinks swapping control and treatment labels cancels carryover
  • Claims residual behavioural effects always fade within days
  • Accepts a reused split because the arms are still equal in size
  • Believes re-salting erases the earlier treatment's effect

context