skip to content

A 50/50 A/B test logged 100,000 control and 98,500 treatment users — is that a sample ratio mismatch?

level: middleimportance: must knowfreq 62%

answer

  1. eyeballing is not a test
  2. total first, then expected per arm
  3. squared deviation over expected count
  4. two arms means one degree of freedom
  5. p near 0.0008, far below any threshold

basics

~10 s

Yes. Against a 50/50 plan those counts give a chi-square goodness-of-fit statistic near 11.3 on one degree of freedom, a p-value around 0.0008. Far too extreme for chance, so treat the experiment as invalid.

solid answer

~40 s

Run the check rather than eyeballing it. The total is 198,500 users, so a 50/50 plan expects 99,250 per arm and each arm deviates by 750. The chi-square goodness-of-fit statistic on the two counts is `2 * 750^2 / 99250 ≈ 11.3` on one degree of freedom, which corresponds to a two-sided p-value of roughly 0.0008; equivalently the observed proportion is about 3.4 standard errors from 0.5. Many experimentation platforms alert well below the usual 0.05 — 0.001 or stricter — because the check runs on every experiment, and this result clears even that bar. So yes, this is a sample ratio mismatch. Note how ordinary it looks by eye: treatment is short by 1.5%, the kind of gap someone would wave off on a dashboard. That is exactly why the test exists.

go deeper

for a junior

Know that the verdict comes from a test on the two user counts against the planned split, not from looking at how close the numbers seem. Remember that a percentage gap means nothing without the sample size.

for a middle

Be able to do the arithmetic out loud: total, expected per arm, squared deviation over expected, one degree of freedom for two arms, then read the p-value and state the conclusion.

for a senior

Show you know where the check is applied — on the assignment population before eligibility filters — and why the alerting threshold sits far below 0.05 when it runs on every experiment in the platform.

for a principal

Be ready to defend the threshold as a policy choice: the tradeoff between undetected small mismatches and alert fatigue across a large experiment portfolio, and what evidence lets a team clear a flag.

## Set the check up properly The question the check answers is: given the configured allocation, could random assignment plausibly have produced these counts? Two ingredients: the **planned proportions** (here 0.5 and 0.5) and the **observed counts** (100,000 and 98,500). **Total:** 100,000 + 98,500 = 198,500. **Expected per arm:** 0.5 * 198,500 = 99,250. **Deviation:** control is +750, treatment is -750. The deviations are equal and opposite by construction, because the total is fixed. ## The arithmetic The goodness-of-fit statistic sums the squared deviation divided by the expected count over the arms: ``` (100000 - 99250)^2 / 99250 + (98500 - 99250)^2 / 99250 = 562500/99250 + 562500/99250 ≈ 5.67 + 5.67 ≈ 11.34 ``` With two arms this has one degree of freedom, and the corresponding two-sided p-value is about 0.0008. The identical check can be written as a test of a single proportion. The observed share in control is 100,000/198,500 ≈ 0.50378. Under the plan, the standard error of that share is `sqrt(0.5 * 0.5 / 198500) ≈ 0.00112`, so the deviation of 0.00378 is about 3.37 standard errors out. Squaring 3.37 gives back 11.34 — the two formulations are the same test, and either is acceptable to present. ## Reading the result A p-value near 0.0008 means: **if** assignment and logging were behaving, a gap this large or larger would show up in roughly eight experiments in ten thousand. It is not a claim that the arms are 99.92% different, and it is not a probability that the experiment is broken — it is the improbability of the data under a healthy pipeline. The conclusion is that a healthy pipeline is an implausible explanation, so something removed users from treatment (or added them to control) by a mechanism. ## Why the threshold is usually stricter than 0.05 Run this check on every experiment and you get a false alarm on 5% of healthy tests at the conventional cutoff. On a platform with hundreds of live tests that is dozens of spurious investigations, and teams quickly learn to ignore the alert. Common practice is to fire at 0.001 or lower, accepting that very small real mismatches go undetected in exchange for alerts that are almost always real. This example clears 0.001 comfortably, so the verdict does not depend on where exactly the line is drawn. ## The trap: eyeballing 1,500 users out of 100,000 is a 1.5% shortfall. Written as "98,500 vs 100,000", most people's instinct is that it is close enough. The check says otherwise, and the check is right: at this sample size the standard error on the split is about a tenth of a percent, so a 1.5% relative shortfall is enormous in units of noise. The lesson generalises — whether a percentage gap is alarming depends entirely on N. The same 1.5% gap on 2,000 total users would be unremarkable. ## What changes with an unequal planned split Nothing about the method: you compare observed counts to whatever proportions were configured, so a 90/10 plan expects 0.9N and 0.1N. What does change is sensitivity. The same **absolute** deviation contributes far more to the statistic when it lands against a small expected count, because the deviation is divided by that expectation. Unequal splits therefore tend to surface mismatches with less traffic than a balanced split does. ## Practical notes - Test the **assignment population**, not a metric-eligible subset that a filter has already touched — filtering before counting can create or hide a mismatch. - The check needs the planned proportions, so record them with the experiment configuration; comparing against 50/50 when the plan was 60/40 manufactures alarms. - Because the check looks at counts and not at the outcome metric, running it continuously from day one is normal and lets a broken experiment be killed early. Bear in mind that anything tested repeatedly will eventually breach a fixed threshold by chance, which is another reason the threshold sits well below 0.05. - Report the statistic and the p-value, not just "looks off" — the number is what makes the case for stopping the test.

  • Would the same check work if the planned split were 90/10 instead of 50/50?
    Yes — you compare the observed counts to whatever proportions were configured, so the expected counts become 0.9N and 0.1N. Sensitivity changes though: the same absolute shortfall is divided by a much smaller expected count in the small arm, so it contributes far more to the statistic. Unequal splits surface mismatches with less traffic than balanced ones.
  • Why do platforms alert at 0.001 rather than the conventional 0.05?
    Because the check runs on every experiment. At 0.05 you would investigate roughly one healthy test in twenty, and with hundreds of concurrent experiments the alert stops being believed. A stricter cutoff trades away detection of very small mismatches for alerts that are nearly always real defects.
  • Is it safe to run this check before the experiment reaches its planned sample size?
    Yes, and it is standard. The check reads arm counts, not the outcome metric, so monitoring it daily lets you kill a broken test early rather than wasting weeks of traffic. Just remember that repeatedly testing anything raises the chance of an eventual false alarm, which is part of why the threshold is set so low.

saying these in an interview costs you the question

  • Dismisses a 1.5% shortfall as obviously within normal variation
  • Compares the arms' conversion rates instead of their user counts
  • Interprets the p-value as the probability the pipeline is healthy
  • Concludes the split is fine because the config said 50/50
  • Applies the check after metric-eligibility filters have removed users

context