skip to content

What is a sample ratio mismatch in an A/B test?

level: juniorimportance: must knowfreq 72%

answer

  1. look at the denominators first
  2. planned split versus observed split
  3. the gap is bigger than chance allows
  4. users were lost non-randomly
  5. exchangeability between arms is gone

basics

~20 s

A sample ratio mismatch is when the observed split of users across experiment arms differs from the planned split by more than chance can explain. It signals a broken assignment or logging pipeline, so the comparison itself is untrustworthy.

solid answer

~50 s

A sample ratio mismatch (SRM) is a statistically implausible gap between the traffic split you configured and the split you actually observe in the data. If you planned 50/50 and logged 100,000 users in control against 98,500 in treatment, a goodness-of-fit test on those two counts tells you whether chance could produce that gap; here it could not. The reason this voids the experiment is that the whole causal claim rests on the two arms being exchangeable populations. Missing users are never missing at random — they were lost by a mechanism, and that mechanism usually removes a distinctive slice, such as slow devices or users who abandoned an extra redirect. Once that happens, any metric difference mixes the treatment effect with a selection effect and cannot be separated. The correct response is to stop reading the metrics, find the cause, fix it, and rerun.

go deeper

for a junior

Be ready to define it in one sentence: observed traffic split versus planned traffic split, checked on user counts. Know that the standard response is to stop and investigate, not to report the numbers.

for a middle

Explain why it invalidates rather than weakens a result: randomization guarantees exchangeable arms, and non-random loss of users breaks that guarantee so a metric gap and a selection effect become inseparable.

for a senior

Show you treat it as a hard gate in practice. Interviewers want to hear that you refuse to read a flagged experiment, that you hunt the mechanism, and that you never rebalance the arms to make the check pass.

for a principal

Own the organisational side: whether the check blocks result readouts by default, who may override it and on what evidence, and how you keep the alert credible enough that teams do not learn to route around it.

## The definition An online experiment is configured with an intended traffic allocation — say 50% of users to control and 50% to treatment. **Sample ratio mismatch (SRM)** is the condition where the counts you actually observe in each arm deviate from that intended allocation by an amount that random assignment could not plausibly produce. Note what is being compared. SRM is a statement about **denominators**, not about the outcome metric. It asks "did the right number of users end up in each bucket?", not "did conversion go up?". You can have a beautiful, tightly-estimated lift and still have an SRM — and if you do, the lift is not interpretable. ## Why counts are never exactly equal Random assignment does not guarantee equal counts; it guarantees that each unit has the planned probability of landing in each arm. With 200,000 users and a fair coin, the two arms differ by hundreds of users routinely. So the question is never "are the counts equal?" but "is the observed gap larger than sampling noise?" That is a hypothesis test on the arm counts against the configured proportions, and it is the reason a formal check exists instead of eyeballing the dashboard. ## Why an SRM voids the result Randomization buys one thing: the two arms are, in expectation, identical on every characteristic — measured, unmeasured, and unimaginable — before treatment is applied. That is what licenses attributing a metric difference to the treatment. An SRM is evidence that units disappeared from (or were added to) one arm through a non-random mechanism. Three properties make this fatal: 1. **The loss is selective.** A redirect that only the treatment arm passes through loses exactly the users with the least patience and the worst connections. A client crash loses exactly the users on old hardware. The survivors in that arm are systematically different from the survivors in the other arm. 2. **The bias is unbounded in direction and size.** If the lost users were low-converters, treatment looks better than it is; if they were high-converters, worse. You cannot sign the bias without knowing the mechanism, and you usually do not. 3. **It cannot be repaired analytically.** Dropping random users from the larger arm restores the counts and does nothing about the bias, because the users missing from the smaller arm were removed by the mechanism, not by chance. Reweighting requires knowing exactly who was lost and why — which, if you knew, you would have fixed instead. The practical rule used by mature experimentation programs is blunt: an SRM-flagged experiment does not get read. Not "read with a caveat", not "read the secondary metrics" — not read. ## What SRM is not - **It is not merely a power problem.** A common wrong answer is "unequal arms cost you a bit of statistical power". Deliberately unequal allocations (say 90/10) are perfectly valid and only cost power. SRM is different in kind: the split you got is not the split you asked for, which means something is broken. - **It is not a small-numbers curiosity.** Whether a given percentage gap is alarming depends entirely on sample size. A 0.5% gap is unremarkable noise at ten thousand users and overwhelming evidence at ten million. - **It is not a treatment effect.** "Treatment made fewer people sign up, that is why there are fewer of them" is only coherent if arm membership is defined by an outcome — which is itself a design bug. Assignment must be independent of anything that happens afterwards. ## Where it comes from Briefly, mismatches enter wherever units can be lost between being bucketed and being counted: an extra redirect hop in one arm, an instrumentation event that fails to fire when one arm's client crashes, or a post-hoc filter (bots, deduplication, outlier trimming) that removes different numbers from each arm. The common shape is a step that happens **after** assignment and behaves differently by arm. ## How to talk about it in an interview A strong answer covers three beats: SRM compares observed to planned allocation and is tested formally on the arm counts; it invalidates rather than degrades the result, because it breaks exchangeability; and the response is diagnose-and-rerun, never rebalance-and-report. Candidates who describe it as "an imbalance that slightly hurts precision" have missed the point entirely.

  • Why can't you fix an SRM by randomly downsampling the larger arm?
    Because the problem is bias, not arithmetic. The users missing from the smaller arm were removed by a mechanism that selects a particular kind of user; discarding random users from the other arm restores the counts while leaving the two populations different. You would end up with a balanced, still-biased comparison, which is worse than an obviously broken one.
  • Does a mismatch under one percent still matter?
    It can be decisive. Whether a percentage gap is alarming depends on sample size: the same 0.5% gap is noise at ten thousand users and overwhelming evidence at ten million. What matters is not the size of the gap but that something other than chance produced it, and that mechanism usually removed a distinctive slice of users.
  • Can the treatment itself legitimately cause fewer users in its arm?
    No, not if assignment is done correctly. Users are bucketed before they experience anything, so arm membership cannot depend on what the treatment does to them. If your counts move because of user behaviour, the exposure or logging step is downstream of the treatment experience, and that is the defect to fix.

It is like a taste test where you promised a hundred cups of each drink but only ninety-eight of one were handed out. Before comparing scores you have to ask which cups went missing, and to whom.

saying these in an interview costs you the question

  • Says an unequal split only costs a little statistical power
  • Calls a sub-one-percent gap too small to matter regardless of sample size
  • Proposes discarding random users to rebalance the arms
  • Reports the metric lift anyway with a written caveat
  • Claims correct randomization should produce exactly equal counts

context