skip to content

How does hash-based bucketing turn a user ID into an A/B test variant?

level: middleimportance: must knowfreq 68%

answer

  1. deterministic, not drawn at runtime
  2. ID plus something experiment-specific
  3. modulo gives a bucket number
  4. bucket ranges are mapped to arms
  5. uniform hash so bucket sizes match

basics

~20 s

Concatenate the user ID with an experiment-specific salt, hash it to an integer, take it modulo the bucket count, and map bucket ranges to arms - 0-49 control, 50-99 treatment. Deterministic, so nothing is stored.

solid answer

~50 s

The assignment is a pure function of the identifier. You build a key such as `user_id + ':' + experiment_salt`, hash it to an integer, reduce it modulo the number of buckets - 100 buckets gives 1% resolution - and then map contiguous bucket ranges to arms: buckets 0-49 control, 50-99 treatment. Three properties make it work. It is **deterministic**, so every service and every offline analysis job derives the same arm with no stored state. The hash must be **well distributed**, so buckets end up close to equal in size and adjacent identifiers scatter rather than clump. And the **salt** makes each experiment a different partition of the same users, instead of every experiment repeating one split. Because arms are bucket ranges, you can widen an arm's range to expose more users without moving anyone already assigned.

go deeper

for a junior

Know the shape of the pipeline: identifier plus salt, hash, modulo, bucket, arm. Be able to say why calling the function twice for the same user gives the same answer both times.

for a middle

Explain the properties the hash must have - even distribution, small input changes scattering across the range, identical results in every service - and why arithmetic on the raw identifier fails them.

for a senior

Show how you would diagnose a skewed split: confirm one hash implementation everywhere, confirm the salt is identical across clients and servers, and check bucket sizes over real traffic before blaming chance.

for a principal

Own the bucketing contract as a platform decision: one pinned hash algorithm, salts issued per experiment, a bucket count sized for the finest exposure you will ever need, and a rule that a running experiment's salt is immutable.

## The pipeline, step by step 1. **Build the key.** Take the identifier of the randomisation unit and concatenate it with a value unique to this experiment - its salt. A key looks like `u-8831902:checkout-redesign-2026q1`. 2. **Hash the key** to a fixed-width integer with a good general-purpose hash. 3. **Reduce to a bucket** with the modulo operator: `bucket = hash(key) % 100`, giving an integer in 0..99. 4. **Map buckets to arms** by contiguous range: 0-49 control, 50-99 treatment. A 10% test uses 0-4 control and 5-9 treatment and leaves 10-99 unexposed. The whole thing is a pure function. There is no assignment table, no write on first exposure, no cross-region replication problem, and the analysis job can reconstruct every assignment from the identifier in the logs. ## What the hash function must give you **Determinism across implementations.** The same key must produce the same integer in every service, every language runtime, and in the offline pipeline. A hash whose result depends on process start-up, on runtime version, or on a language's built-in string hashing can differ between two services and split a single user across arms. Pick one algorithm, pin it, and treat it as a contract. **Uniformity.** Bucket sizes should be equal in expectation. A weak hash - a checksum, a sum of character codes, a truncation - leaves structure in the output and produces buckets of visibly different sizes. **Avalanche.** Keys that differ in one character should land in unrelated buckets. This is what makes consecutive identifiers scatter instead of marching through buckets in order. Cryptographic strength is not required; bucketing is not a security boundary. What is required is good distribution and stability. ## Why not just take the identifier modulo 100 Because raw identifiers carry structure, and modulo preserves it. - Sequentially issued IDs make the remainder a near-perfect function of signup order only in blocks, but any block-allocation scheme - IDs handed out in ranges per shard, per region, or per registration channel - makes buckets correlate with infrastructure or geography. - Some ID schemes embed a timestamp or a machine number in their low bits, so the remainder encodes when or where the account was created. The symptom is nasty because bucket **sizes** still look fine: each remainder gets roughly the same count of users. What differs is *who* is in them - one arm quietly ends up with older accounts, or with one data centre's traffic. Hashing destroys that structure so the bucket number carries no information about the user. ## Choosing the bucket count Pick enough buckets that the smallest exposure you will ever want is a whole number of buckets. With 100 buckets you cannot expose 0.5% of users. Using 1000 or 10000 buckets costs nothing, keeps arm sizes exactly on target, and reduces how much leverage any single unlucky bucket has on the overall split. The mapping from bucket range to arm stays just as simple. ## The salt, and when it may change The salt makes each experiment hash the same population differently, so that being in the first half of one experiment tells you nothing about where you land in the next. It must be identical everywhere the experiment is evaluated - a salt that differs between a mobile client and a server is one of the classic causes of a user appearing in two arms. Changing the salt of a **running** experiment re-buckets everybody: the hash of the same identifier under a new salt is unrelated to the old one, so roughly half of each arm swaps sides. Stickiness is destroyed mid-flight and the data is not analysable. Salt changes belong between experiments. ## Verifying the scheme Before trusting it, check the boring things: bucket sizes are close to equal over a real population; the same identifier evaluates to the same arm in every service; the modulo is applied to an unsigned or absolute value so negative hash outputs do not collapse into a subset of buckets; and the mapping from bucket to arm is stored as configuration you can read, rather than being computed differently in two places. ## What an interviewer is listening for They want the four-step pipeline, the reason a raw identifier is not a hash, and the fact that the salt is what stops every experiment from reusing one fixed split. Candidates who describe generating a random number at assignment time and storing it have missed the entire point of the design.

  • Why not simply take a numeric user ID modulo 100 and skip the hash?
    Because identifiers carry structure that modulo preserves. IDs allocated in blocks per shard, per region or per signup channel make bucket membership correlate with infrastructure or tenure, and some schemes embed a timestamp in their low bits. Bucket sizes still look even, so the imbalance hides in who is in each bucket rather than how many. A good hash destroys that structure.
  • What happens if you change an experiment's salt while it is running?
    Everyone is re-bucketed. The hash of the same identifier under a new salt is unrelated to the old one, so roughly half of each arm swaps sides mid-flight. Stickiness is gone, both arms contain partly treated users, and the run is not analysable. Salt changes belong between experiments, never inside one.
  • How many buckets should you divide the hash space into?
    Enough that the smallest exposure you will ever want is a whole number of buckets. 100 buckets caps you at 1% resolution; 1000 or 10000 cost nothing, keep arm sizes exactly on target, and reduce how much a single unlucky bucket can skew the split.

saying these in an interview costs you the question

  • Uses the raw numeric user ID modulo 100 as the bucket
  • Draws a random number at assignment time and stores it
  • Reuses a single salt across every experiment
  • Assumes any hash distributes uniformly, including a checksum
  • Changes the salt mid-experiment to refresh the split

context