skip to content

A sample ratio mismatch alert fired mid-test but arm counts matched by the end — do you trust the result?

level: seniorimportance: nice to knowfreq 28%

answer

  1. an alert clearing is not evidence of health
  2. two stories: late records or a fixed defect
  3. recompute an old window after ingestion settles
  4. look at the daily ratio, not the total
  5. one arm short early, the other short later

basics

~20 s

Not automatically. Late-arriving logs from one arm can create a mismatch that resolves on its own. But a defect that started and stopped leaves corrupted data behind. Establish which you saw before reading any metric.

solid answer

~50 s

The benign explanation is ingestion lag: one arm's events batch or flush more slowly, so a snapshot taken mid-test undercounts it and the gap closes once late records land. Confirm that by recomputing the ratio for the **same historical window** after ingestion has settled — if yesterday's counts look balanced today, you saw latency, not loss. The dangerous explanation is a real defect that was live for part of the run and then stopped. Then the days it was active carry non-random user loss, and a healthy grand total only means a later period diluted it. Check per-day ratios rather than the aggregate, and watch for a shortfall in one arm being offset by a shortfall in the other later — two bugs cancelling is not health. If the affected days cannot be cleanly excluded, rerun.

go deeper

for a junior

Know that a mismatch alert clearing on its own does not mean the experiment is fine, and that late-arriving log data is one common but not automatic explanation.

for a middle

Explain the recompute test: measure the same historical window again after ingestion has settled, and see whether the earlier gap persists. That distinguishes delayed records from users who were never there.

for a senior

Show the operational habits — daily ratio time series, closed-window evaluation, correlating the alert period with deploys and rollbacks — and the discipline to rerun when a contaminated period cannot be cleanly excluded.

for a principal

Own the monitoring design: which windows the check evaluates, whether alerts require a sustained breach, what history is retained so an old window can be recomputed, and how exclusion decisions are governed so they are not made after seeing the result.

## Two very different stories, one symptom An alert that fires and then clears has exactly two families of explanation, and they call for opposite actions. **Story A: measurement lag.** Events from the two arms do not arrive at the same speed. One client batches events and flushes on a timer or on app background; another posts immediately. A regional pipeline backs up for a few hours. In all of these, the users existed and were exposed — their records simply had not arrived when the check ran. Once ingestion completes, the historical counts are correct and the experiment is unharmed. **Story B: a defect with a lifetime.** A bad build shipped on day three and was rolled back on day five. During those two days, treatment lost a slice of users for real. From day six the ratio is healthy again, and the cumulative counts creep back toward balance because the healthy days outnumber the broken ones. The alert clearing means nothing here: the broken days are still in the dataset, still biased. ## The test that separates them Recompute the ratio for a **fixed historical window** at two different times. If the counts for "day three" looked skewed when measured on day three but look balanced when recomputed on day ten, the missing records simply arrived late — story A. If day three still looks skewed when recomputed after ingestion has fully settled, the users were never there — story B. This is why the SRM check should run on **closed windows**: periods whose ingestion delay has fully elapsed. Evaluating on partial, still-filling data produces exactly the transient alerts that teach teams to ignore the alarm. ## Why the grand total is a bad summary The cumulative ratio is a sum, and sums hide structure. Two failure patterns survive a clean-looking total: - **Dilution.** Treatment is short 1,000 users over two broken days; the remaining twelve healthy days add so much balanced traffic that the aggregate test no longer reaches the threshold. The bias in those two days does not disappear because it was averaged with clean data — it is still in the numbers you will analyse. - **Cancellation.** Control is short early, treatment is short later, and the totals reconcile. This is strictly worse than a persistent mismatch, because it means two distinct defects rather than one, and the aggregate check is blind to both. The defence is to look at the ratio **per day** (and per major segment) over the life of the experiment, not just at the end state. A time series of the daily split makes both patterns obvious at a glance. ## Deciding what to do If you confirm ingestion lag: nothing to fix in the experiment; fix the monitoring so it evaluates settled windows and alerts on a sustained breach rather than a single snapshot. If you confirm a defect with a lifetime: the affected period is contaminated. Dropping those days is tempting but is only defensible when the exclusion rule is decided on grounds independent of the outcome — excluding days because they were broken is reasonable; excluding them after seeing which choice makes the result significant is not. When the contaminated period is a meaningful share of the run, or when the affected users are still in the dataset on later days having already had a broken experience, the honest answer is to rerun. If you cannot distinguish the two stories — no per-day history retained, no way to recompute an old window — you do not have grounds to trust the experiment. That gap is itself a finding worth fixing in the platform. ## The interview point What is being tested here is whether you treat an alert clearing as evidence. A weak answer says "it resolved, so it was a glitch". A strong answer names both stories, gives the recompute-a-fixed-window test that tells them apart, and refuses to substitute the grand total for the daily series.

  • How do you make the mismatch check immune to logging lag?
    Evaluate on closed windows only — days whose ingestion delay has fully elapsed — rather than on partially filled recent data. Pair that with alerting on a sustained breach across consecutive windows instead of a single snapshot. Both changes remove the transient alarms that otherwise train teams to dismiss the alert.
  • Can the totals balance while the experiment is still corrupted?
    Yes, in two ways. Dilution: a short broken period is swamped by many healthy days, so the aggregate test no longer fires while the biased records remain in the analysis. Cancellation: one arm is short early and the other short later, so the sums reconcile while two separate defects were active. Per-day ratios expose both.
  • Is dropping the affected days a legitimate fix?
    Only if the exclusion rule is chosen on grounds independent of the outcome and applied to both arms identically — excluding a period because a rollback record shows the defect was live is defensible. Choosing the exclusion after seeing which version makes the result significant is not. When the contaminated share is large, rerun instead.

saying these in an interview costs you the question

  • Treats a self-resolving alert as proof the experiment is healthy
  • Checks only the cumulative ratio and never the daily series
  • Assumes every transient alert is ingestion lag
  • Excludes the flagged days after seeing which choice helps the result
  • Evaluates the check on a still-filling recent window

context