skip to content

As an experimentation lead, how do you set the default randomization unit for every team?

level: principalimportance: nice to knowfreq 32%

answer

  1. make the safe choice the default
  2. record the unit as structured metadata
  3. analysis grain derived, not hand-chosen
  4. escape hatch behind a short review
  5. never switch units mid-run

basics

~20 s

Make the most stable identity the default, record the chosen unit as a first-class property of every experiment, derive the analysis grain from it automatically, and require review before any team randomizes finer than a person.

solid answer

~50 s

The lead's job here is to remove a decision most teams get wrong. I would set the default to the most stable identity available — the account identifier where the surface is signed in, a long-lived device identifier otherwise — because that keeps experiences coherent and keeps person-level metrics such as retention definable. The unit must then be a recorded, structured property of the experiment, not a comment in a doc, so the analysis pipeline can force the analysis grain to match it and refuse a mismatch instead of trusting an analyst. Finer units (session, request) and coarser ones (store, city, geo) stay available but behind a short review that asks about carryover, the metric's scope and the true replicate count. I would also forbid changing the unit mid-experiment, since it invalidates the run, and require every readout to state its unit, because a session-randomized result and a user-randomized result answer different questions and should not be compared as though they were interchangeable.

go deeper

for a junior

You will not set this policy, but know that your organisation has a default unit and that following it is what keeps your result comparable with everyone else's.

for a middle

Be able to explain why a platform would default to the most stable identity even though finer units would give more observations.

for a senior

Argue the specific cases for deviating — stateless infrastructure changes, geo-level tests — and say what evidence you would bring to justify one.

for a principal

Own the whole design: the default, the metadata that makes analysis follow it, the review gate for deviations, the ban on mid-flight changes, and the reporting convention that keeps results comparable.

## The problem a platform lead is actually solving Individual teams choose randomization units under time pressure, usually by copying the last experiment. Left to itself, an organisation accumulates a mixture of user-, session- and request-randomized tests whose results are compared to each other in a review meeting as though they measured the same thing. The lead's leverage is not in reviewing each design; it is in making the safe choice automatic and the unsafe choice deliberate. ## Choosing the default The default should be the **most stable identity available on the surface**: - **Signed-in surfaces:** the account identifier. Stable across devices, immune to cookie churn, and it makes person-level metrics well-defined. - **Logged-out surfaces:** a long-lived device identifier, accepting that it is a device and not a person. This default is not chosen because it is statistically optimal — finer units often give more power — but because it fails safely. A user-level design is almost never *invalid*; it is sometimes merely underpowered. A session- or request-level design can be silently invalid, and the failure does not announce itself. ## Making the unit a first-class object The single highest-value piece of infrastructure is recording the randomization unit as structured metadata on every experiment, then wiring the analysis to it: - The readout aggregates to the recorded unit automatically, so the number of independent observations matches the number of assignments by construction. - A metric declared at a different grain than the experiment's unit either gets aggregated or is refused, rather than being quietly compared row by row. - Power calculations take the unit from the same field, so the sizing and the analysis cannot disagree. This converts the most common statistical error in experimentation from a judgement call into an impossibility. ## Governing deviations Deviations should be allowed — some experiments genuinely need a finer or coarser unit — but gated by a small number of questions: 1. **Carryover.** Can exposure change later behaviour? If yes, a finer unit than the person is out. 2. **Metric scope.** Is the primary metric defined over a person or a longer horizon? If yes, the unit must be at least a person. 3. **Replicate count.** For coarse units, how many stores or cities are there really? A geo experiment with 20 units needs its expectations set before it starts, not after. 4. **Visibility.** Would a person notice the experience changing between visits or requests? A lightweight checklist answered in the experiment's configuration is enough; the point is that someone had to say the words, not that a committee met. ## Rules that prevent whole classes of failure - **Never change the unit mid-experiment.** Switching from cookie to account, or from session to user, after the run starts scrambles the assignment history and invalidates the comparison. The correct move is to stop and restart. - **Every readout states its unit.** A one-line label on the result page prevents a session-randomized 3% lift being stacked against a user-randomized 3% lift in a portfolio review. - **Comparability discipline.** Aggregate reporting of "experiment wins this quarter" should not mix estimands. If the organisation cares about cumulative impact, it needs one unit convention for the headline metric. ## The tradeoffs to own out loud **Power versus validity.** Standardising on the person costs power relative to session or request designs. The lead should be explicit that the platform is paying that price deliberately, and should invest in variance reduction to buy the power back rather than in finer units. **Consistency versus autonomy.** A rigid rule slows teams with legitimate needs — latency experiments, infrastructure tests, geo-level marketing measurement. The gate must be cheap enough that those teams route through it instead of around it. **Identity investment.** Better identity resolution across devices improves every experiment at once. That is a platform-level investment with platform-level returns, and it competes against features; the lead is the person who has to make that argument with numbers about attenuation and lost sensitivity. ## What a strong answer sounds like A convincing candidate does not just say "use user-level randomization". They describe the default, the mechanism that makes the analysis follow the default automatically, the escape hatch and its gate, the prohibition on mid-flight changes, and the reporting convention that stops incomparable results being summed. They also name what the standard costs, because a policy whose downsides you cannot state has not really been thought through.

  • A team wants request-level randomization for an infrastructure change. What do you require before approving it?
    Evidence on three points: the change is imperceptible to the person, nothing carries over between requests, and the primary metric is a property of the request such as latency or error rate. I would also require that no person-level metric be reported from that experiment, and that the readout label the unit so nobody later compares its numbers with a user-level result.
  • How do you stop a session-randomized 3% lift being compared with a user-randomized 3% lift in a portfolio review?
    Label the unit on every result, and set a convention that headline impact accounting uses one unit only. The two numbers answer different questions — effect per visit versus effect per person — and they weight the population differently, so summing them produces a portfolio figure that corresponds to nothing. The discipline has to live in the reporting template, not in each reviewer's memory.
  • What is the strongest argument against standardising on a single default unit?
    It costs sensitivity. Finer units give far more observations, and for genuinely stateless changes that extra power translates into faster, cheaper decisions. The counter is that the failure modes are asymmetric: an underpowered user-level test yields an inconclusive result you can act on cautiously, while an invalid session-level test yields a confident wrong answer. Buy the power back through variance reduction instead.

saying these in an interview costs you the question

  • Mandates one unit with no escape hatch
  • Leaves the unit as free-text documentation
  • Allows the unit to change mid-experiment
  • Compares results across units without labelling them
  • Cannot state what the standard costs in power

context