skip to content

How do you find the root cause of a sample ratio mismatch in a running A/B test?

level: seniorimportance: should knowfreq 50%

answer

  1. count at every stage, not just the end
  2. where does the ratio first break
  3. assignment, exposure, then analysis filters
  4. slice by device, browser, country, day
  5. did the gap start at a deploy timestamp

basics

~20 s

Localise the loss. Recount users at each pipeline stage — assignment, exposure, then any post-hoc filters — and slice by platform, browser, device, country and day. The first stage and segment where the ratio breaks names the bug.

solid answer

~50 s

Treat it as a funnel problem on counts. Start at assignment: if the arms are already skewed at the moment users are bucketed, the problem is upstream in the assignment service. If assignment is balanced but exposure is skewed, users were lost between bucketing and logging — the classic causes are an extra redirect hop that only the treatment arm passes through, so impatient users drop off before the page loads, and a treatment-only crash on older devices whose exposure events never fire. If both are balanced and the gap appears only in the analysis dataset, look for a filter applied after assignment that hits one arm harder, such as bot removal or a join that silently drops rows. Then slice: a mismatch concentrated in one browser, OS version or country usually points straight at the defect, and a gap that starts at a deploy timestamp names the release.

go deeper

for a junior

Know that the investigation is about counting users at more than one point in the pipeline, and that slicing by device, browser and date is how the affected population is identified.

for a middle

Explain the assignment-exposure-analysis funnel and name the mechanisms at each stage: an extra redirect hop, a client crash that suppresses exposure events, and filters applied after assignment.

for a senior

Demonstrate a systematic triage you have actually run: per-stage counts, dimension slicing, deploy-timestamp overlay, and a bar for accepting a cause that requires it to explain the full magnitude of the gap.

for a principal

Speak to prevention: instrumentation standards that log assignment and exposure separately, a rule that filters may only use pre-assignment information, and automated diagnostics so triage takes hours instead of days.

## The mental model: counts as a funnel Every user in an experiment passes through stages, and users can be lost at each one: 1. **Assigned** — the platform decides which arm this user belongs to. 2. **Exposed** — the user actually receives the experience and an event records it. 3. **Analysed** — the user survives whatever cleaning the analysis pipeline applies. A sample ratio mismatch is a claim about the ratio at whichever stage you measured. The diagnostic move is to compute the ratio **at every stage** and find the first one where it breaks. That single step turns an open-ended hunt into a bounded one. ## Stage 1: skewed at assignment If the counts are already off at the moment of bucketing, nothing downstream is to blame. This is the rarest and most serious case, and it belongs to the assignment service itself. Confirm it on raw assignment records that no filter has touched, because measuring "assignment" through a table that has already been joined or deduplicated will mislead you. ## Stage 2: balanced at assignment, skewed at exposure This is where most real mismatches live, and the mechanisms are physical rather than statistical. **The extra redirect.** The treatment variant is served behind a redirect — an additional network round trip before the experiment page renders. Every hop loses users: slow connections time out, impatient users close the tab, some clients mishandle the hop. The users who survive the redirect are systematically faster and more patient than the control arm's, so the arm is both smaller and different. The tell is that the shortfall concentrates in mobile traffic and slow-network regions, and that it appears immediately at launch rather than drifting in. **The one-arm crash.** The treatment code path crashes on a subset of clients — an old OS version, an unusual screen size, a device without some capability. Those users never fire the exposure event, so they vanish from the treatment count entirely. The tell is razor-sharp segmentation: the ratio is fine everywhere except one OS version band, where treatment is missing almost completely. **Asymmetric event timing.** If the exposure event fires at a different point in the render path in each arm — early in control, after an extra component mounts in treatment — then anything that interrupts loading costs treatment more. This is a design bug in the instrumentation, not in the feature. ## Stage 3: balanced at exposure, skewed in analysis Any filter applied **after** assignment can create a mismatch if its hit rate correlates with the arm. Bot and crawler filtering is the standard example: automated traffic is assigned like any other traffic, and if the treatment experience changes what a crawler can reach or how it behaves, the filter removes a different number from each arm. Deduplication, session stitching, outlier trimming and inner joins against another table all behave the same way. The defensive rule is that filters should use only information determined before assignment; a filter that depends on post-assignment behaviour is a mismatch generator. ## Slicing: the second axis Once the stage is known, cut the counts by dimension: - **Platform / OS version / browser / app version** — isolates crashes and rendering defects. - **Country and network type** — isolates latency-driven losses such as the redirect. - **Day and hour** — a ratio that is clean for five days and breaks on the sixth points at a deploy, a config change, or an incident. Overlay release timestamps. - **New versus returning traffic** — an imbalance confined to one of these hints at caching or an identity-resolution problem. A mismatch that is uniform across every slice is unusual and points at the assignment or configuration layer; a mismatch concentrated in one slice hands you the bug. ## Confirming the fix A cause is only confirmed when it explains the **magnitude and the shape**, not just the direction. If a crash on one OS version accounts for 200 missing users but 1,500 are missing, keep looking — partial explanations are how teams talk themselves into re-reading a broken experiment. After the fix, rerun the experiment; do not patch the historical data and reinterpret it. ## What a weak answer looks like "It is probably randomness, let us restart it" — no diagnosis, so the same defect recurs. "The assignment must be buggy" — jumping to the least likely stage. Or reporting the metrics because the cause was found and seemed small: knowing the mechanism does not undo the selective loss it caused.

  • The counts are balanced at assignment but skewed at exposure — what does that narrow it to?
    The loss happens between bucketing and logging, so the assignment service is exonerated. Look at the client and the delivery path: an extra redirect hop in one arm, a crash that stops the exposure event firing, or an exposure event that fires later in one arm's render sequence. Segment by device and network to find which users are disappearing.
  • Why is bot filtering such a common cause?
    Bots are assigned like any other traffic, and the filter removes them after assignment. If the treatment experience changes what automated clients can reach or how they behave, the filter deletes a different number from each arm and manufactures a mismatch. The general rule is that filters must depend only on information available before assignment.
  • When do you consider the cause confirmed?
    When it accounts for the magnitude and the shape of the gap, not just its direction. If a crash explains 200 of 1,500 missing users, something else is still operating. Partial explanations are how teams convince themselves to read a broken experiment; hold out for a mechanism that reproduces the observed deficit.

saying these in an interview costs you the question

  • Blames chance and restarts the test without diagnosing
  • Only inspects the final aggregate counts, never per-stage counts
  • Assumes the assignment service is at fault by default
  • Applies bot or outlier filters after assignment and never rechecks the ratio
  • Accepts a cause that explains only a fraction of the missing users

context