skip to content

What does an A/A test validate before you launch a real experiment?

level: middleimportance: should knowfreq 58%

answer

  1. both arms get the identical experience
  2. tests the plumbing, not the feature
  3. biggest yield is instrumentation defects
  4. double-counted events show as a huge gap
  5. a clean run is not proof

basics

~20 s

An A/A test splits traffic normally but serves both arms the identical experience. It validates the plumbing - a split that produces comparable groups, a metric pipeline that agrees on identically treated users, and an analysis that does not flag phantom differences.

solid answer

~50 s

An A/A test is a null experiment: users are bucketed and analysed exactly as in a real test, but both arms get the same product. Any difference you then observe is noise or a defect. It can catch three distinct classes of problem. A **broken split**, where the assignment mechanism does not produce the ratio or the comparable groups you configured. A **broken metric pipeline**, which is the most common real finding - one arm's events double-counted by a second instrumentation path, or a join dropping sessions for one arm, showing up as a large apparent effect between two identical experiences. And a **mis-calibrated analysis**, where repeated A/A runs flag far more often than the alpha you set, meaning the variance model does not fit the data. What an A/A cannot do is validate the treatment itself, and one clean run is a smoke test rather than proof.

go deeper

for a junior

Be able to define it: both arms receive the same experience, so any measured difference is noise or a defect. Know that it is run to check the experiment platform, not a feature.

for a middle

Explain the three things it can catch - a split that does not match the configuration, duplicated or dropped events in one arm, and an analysis that flags more often than its stated error rate - and name what it cannot catch.

for a senior

Demonstrate the triage: separate assignment counts from downstream metric counts, follow one user through both, and decide whether the finding blocks a launch or is credible as noise.

for a principal

Set the policy: whether A/A runs are a gate before each launch or a permanent background check, which findings block a release, and how much live traffic the organisation will spend on validating its own platform.

## The idea An A/A test runs the entire experiment machinery with the treatment removed. Users are hashed into buckets, buckets are mapped to two arms, exposures are logged, metrics are computed, and the comparison is run - but both arms see the identical current product. Under a correct platform there is nothing to find, so anything the A/A reports is telling you about the platform rather than about the product. That is the whole point: it turns a silent system into an observable one. ## What it can catch **A broken split.** The assignment mechanism may not be producing what you configured. Symptoms include arm memberships that do not match the intended ratio, or two arms that differ systematically in who is in them - one arm carrying older accounts, one region, or one client version. The first of those is an imbalance in the split itself and is treated as a blocking defect diagnosed on its own terms; the second is what a comparison of pre-existing user characteristics across the arms surfaces. **A broken metric pipeline.** This is the finding that pays for A/A testing. Between two identical experiences, a 15% or 20% gap in an event count cannot be sampling variation at experiment scale - it is instrumentation. The classic case is **double counting**: one arm's events are emitted by both a legacy and a new logging path, so its counts are inflated. Other variants are a join that silently drops sessions for one arm, a deduplication rule applied asymmetrically, or a metric defined over a table that is populated for one code path only. None of these are visible in a real A/B test, because there they masquerade as a treatment effect. **A mis-calibrated analysis.** Even with a clean split and clean data, the statistical procedure can be wrong for the data it is given. If the variance is understated, p-values come out too small and the platform flags differences that do not exist far more often than its stated error rate. You do not see this in a single run; you see it across many A/A runs, where the share that flag any given metric should sit near your alpha and the p-values should be scattered across the whole range rather than piling up near zero. That is the acceptance criterion for the platform. ## What it cannot catch An A/A test says nothing about anything that only happens when the treatment is on: a slower code path in the new variant, an error that fires only on the new screen, a metric that is only emitted by the new UI. It also cannot certify the split as unbiased in general. Like any test, it only detects imbalances that are large relative to its noise, so a small systematic bias passes unnoticed. Treat a single clean A/A as ruling out gross breakage, not as a certificate. ## How teams actually run them Three patterns are common, and they are not exclusive: - **Pre-launch gate.** Run an A/A on the exact split the upcoming experiment will use, particularly when the metric or the population is new. - **Continuous background A/A.** Keep a permanent null experiment running so the platform's false-positive behaviour is measured continuously rather than remembered from a one-off exercise. - **Retrospective A/A.** Re-analyse historical data by splitting a past period with a fresh salt and running the analysis as if it were an experiment. This is cheap because it spends no live traffic, but it only exercises the metric and analysis layers, not the live assignment path. ## Reading a flagged A/A The triage question is always *where does the asymmetry first appear*. Compare the number of assigned users per arm; if that already deviates from the configured ratio, the problem is upstream in assignment and nothing downstream can be trusted. If assignment counts match but a downstream count is inflated, follow a single user through both tallies and find where one of their events is counted twice or dropped. Only when both of those are clean does `expected noise` become a reasonable explanation - and then the size of the discrepancy and how extreme the p-value is decide whether it is credible as noise. ## Interview framing A strong answer names all three failure classes, insists that the biggest practical yield is instrumentation rather than randomisation, and volunteers the limitation: an A/A is underpowered against small biases, so its real value comes from running it repeatedly and checking calibration, not from one green tick.

  • An A/A run shows one arm with 8% more sessions than the other - split bug or pipeline bug?
    Find where the asymmetry starts. If the count of assigned users already deviates from the configured ratio, the assignment mechanism is at fault and everything downstream is untrustworthy. If assignment counts match but the session tally is inflated, the pipeline is duplicating or dropping events for one arm - a second logging path firing on one code branch is the classic cause. Trace one user through both counts before deciding.
  • Is a single clean A/A run evidence that the split is unbiased?
    Only weakly. An A/A has the same power limits as any test: it detects only imbalances that are large relative to its noise, so modest systematic bias passes unnoticed. Treat one clean run as ruling out gross breakage, and get real confidence from repeated runs whose flag rate matches the alpha you set.
  • What can an A/A test never catch?
    Anything that only occurs when the treatment is on - a slow path in the new variant, an error that fires only on the new screen, or a metric emitted only by new code. An A/A exercises assignment, logging and analysis with the feature absent, so treatment-specific defects need their own monitoring once the real experiment starts.

saying these in an interview costs you the question

  • Calls an A/A a waste of traffic because nothing changes
  • Treats one clean A/A run as proof the split is unbiased
  • Expects zero flagged metrics in a healthy A/A run
  • Blames randomness for a 20% gap between identical arms
  • Believes an A/A also validates the treatment code path

context