A/B Testing & Experiment Design
You will learn to design a trustworthy online experiment end to end: pick the randomization unit, size the sample, choose metrics and guardrails, and avoid peeking, SRM, and novelty traps. This is the signature interview loop for data scientist and analyst roles — 'design an A/B test for feature X' is near-guaranteed.
on this pageshowhide
explore
- Randomization & Assignment22 questions
- Unit of Randomization6 questions
- Bucketing & A/A Tests6 questions
- Overlapping Experiment Layers5 questions
- Ramp Plans & Rollout Decisions5 questions
- Sample Size Planning10 questions
- Minimum Detectable Effect5 questions
- Runtime & Seasonality5 questions
- Metrics & Guardrails30 questions
- Overall Evaluation Criterion5 questions
- Guardrail & Counter Metrics5 questions
- Delta Method for Ratios4 questions
- Heavy-Tailed Metrics6 questions
- Metric Definition & Denominators5 questions
- Retention, Funnels & Activity5 questions
- Readout & Decision Rules15 questions
- Interval-Based Decisions5 questions
- Segmentation Without P-Hacking5 questions
- Triggered Analysis & Dilution5 questions
- Validity Threats20 questions
- Peeking & Optional Stopping5 questions
- Sample Ratio Mismatch5 questions
- Novelty & Primacy Effects5 questions
- Network Interference5 questions
- Sensitivity & Adaptive Designs15 questions
- CUPED Variance Reduction5 questions
- Always-Valid Sequential Tests5 questions
- Bandits vs Fixed Horizon5 questions
questions
112 · 6 sectionsWhy must an A/B test give the same user the same variant on every visit?
basics
~20 sSticky assignment gives each user one consistent experience instead of a flickering mix. If the variant is re-drawn per visit, returning users receive both treatments, which blends the two arms and shrinks the measured difference toward zero.
What do experiment layers do in an A/B platform running many tests at once?
basics
~20 sLayers let many experiments share the same traffic at the same time. Each layer randomizes users independently, so a user's arm in one layer says nothing about their arm in another, and concurrent tests do not confound one another.
Why ramp an A/B test through 1%, 5%, 25% and 50% instead of starting at full traffic?
basics
~20 sA staged ramp limits blast radius: the first small stage exposes few users, so crashes, latency spikes or a collapsing metric are caught cheaply. Later, larger stages exist to measure the effect once the change has been shown to be safe.
How does hash-based bucketing turn a user ID into an A/B test variant?
basics
~20 sConcatenate the user ID with an experiment-specific salt, hash it to an integer, take it modulo the bucket count, and map bucket ranges to arms - 0-49 control, 50-99 treatment. Deterministic, so nothing is stored.
When would you randomize an A/B test by session rather than by user?
basics
~20 sRandomize by session only when the change leaves no trace between visits and the outcome is finished within one visit — a layout tweak judged by same-visit conversion. If the treatment can be learned, remembered or accumulated, randomize by user.
Why do A/B tests usually run for whole weeks rather than an arbitrary number of days?
basics
~20 sUser behaviour differs by day of week, so a partial week samples an unrepresentative mix. Whole multiples of seven days weight every weekday equally, so the measured lift describes a typical week, not whichever days happened to be included.
What is the minimum detectable effect (MDE) in an A/B test sample-size calculation?
basics
~20 sThe minimum detectable effect is the smallest true difference between variants a test is designed to catch. Given the sample size, significance level and target power, a true effect smaller than the MDE will usually go undetected.
A test spec says 'detect a 5% lift' on a 3% conversion rate. Why is that underspecified?
basics
~10 sIt never says whether 5% means 5 percentage points, 3% to 8%, or a 5% relative lift, 3% to 3.15%. Those targets differ by hundreds of times in required traffic, so sizing cannot start.
Your experiment hits its planned sample size in three days. Why keep it running to the planned end date?
basics
~20 sA sample count and a calendar window are different requirements. Three days covers only three weekdays and over-represents your most frequent visitors, so the sample is large but not representative. Read the result when both are met.
A planned two-week test window overlaps Black Friday. How do you handle the runtime?
basics
~20 sDecide before launch. The cleanest option is usually to shift the window so it sits entirely before or after the holiday. If the test must span it, keep whole weeks and treat the result as an estimate for holiday traffic.
What is a guardrail metric in an A/B test, and how does it differ from a driver metric?
basics
~20 sA guardrail is a metric the experiment must not damage — latency, crash-free sessions, unsubscribes — and a big enough degradation can veto the launch. A driver metric only explains why the primary metric moved; it never vetoes.
Why does revenue per user make an A/B test hard to call when a few users spend far more than the rest?
basics
~10 sAlmost all the variance in revenue per user comes from a handful of very large spenders, so the interval around the treatment-control difference stays wide. Ordinary experiment sizes then cannot resolve realistic revenue effects.
What is an Overall Evaluation Criterion (OEC) in an A/B test?
basics
~20 sThe Overall Evaluation Criterion is the single metric, agreed before launch, that an experiment's ship-or-not decision is defined against. It expresses what the team means by the product getting better, so the result cannot be reinterpreted afterwards.
What does a DAU/MAU stickiness ratio of 0.2 tell you about how a product is used?
basics
~10 sA DAU/MAU of 0.2 means the typical monthly active user opens the product on about 6 days out of 30. It measures return frequency, not audience size, and averages over very different user types.
Why report a confidence interval on A/B lift instead of just a p-value?
basics
~20 sA p-value only says how compatible the data are with zero effect. A confidence interval reports the lift itself with its precision, so you can compare the plausible range against the minimum detectable effect and the launch bar.
Why does slicing a flat A/B test result into many segments so often produce a false winner?
basics
~20 sEvery segment cut is another hypothesis test. At a 5% false-positive rate, about twenty cuts of a genuinely flat experiment throw up roughly one significant segment by chance alone, so the winner you find is usually noise.
What is dilution in an A/B test where only 3% of assigned users ever see the change?
basics
~20 sDilution is the shrinking of a measured effect when most assigned users never encounter the change. At a 3% trigger rate an effect among exposed users appears about thirty times smaller in the all-up read, while the noise stays.
In an A/B readout, how does an intent-to-treat analysis differ from a triggered-only analysis?
basics
~20 sIntent-to-treat counts every assigned user and answers what shipping does to the whole population. A triggered analysis keeps only users who met the trigger condition in both arms, answering what the change does to those who encounter it.
An A/B test reads +0.4% lift, 95% CI [-2.0%, +2.8%], planned MDE 2% — is that flat?
basics
~20 sNo — it is inconclusive, not flat. The interval reaches +2.8%, above the 2% minimum detectable effect the test was planned around, so a lift worth shipping has not been ruled out, and neither has a 2% loss.
What are novelty and primacy effects in an A/B test?
basics
~20 sBoth are temporary reactions to the fact that something changed, not to what it changed into. Novelty inflates an early result as users explore something new; primacy, or change aversion, depresses it as users unlearn a habit.
On day 2 of a planned 14-day A/B test the dashboard shows p = 0.04 — do you ship?
basics
~20 sNo. The 0.05 threshold is calibrated for a single analysis at the planned sample size, so a day-2 reading is not the test you designed. Run to the agreed horizon unless the plan already allowed an early call.
What is a sample ratio mismatch in an A/B test?
basics
~20 sA sample ratio mismatch is when the observed split of users across experiment arms differs from the planned split by more than chance can explain. It signals a broken assignment or logging pipeline, so the comparison itself is untrustworthy.
In a marketplace test where treated sellers win demand from control sellers, which way is the effect biased?
basics
~20 sUpward. Treated sellers capture bookings that control sellers would otherwise have received, so treatment is pushed up while control is pushed down. The measured lift counts one transfer twice and mixes real growth with redistribution that vanishes at full rollout.
How do you check whether an A/B test effect is decaying over exposure time?
basics
~20 sPut every user on their own clock: days since that user first saw the change. Plot the treatment-minus-control gap against exposure age, grouped by first-exposure date. A gap that shrinks with exposure age, repeating per group, indicates decay.
Would you use a bandit allocator or a fixed-horizon A/B test for twelve news headlines whose value expires within hours?
basics
~20 sUse adaptive allocation. With twelve options, clicks arriving in seconds and value gone in hours, the cost of serving losers dominates and no precise per-headline lift is needed. A fixed-horizon test would still be running when the headlines are worthless.
What is CUPED, and how does a pre-experiment covariate reduce the variance of an A/B test?
basics
~20 sCUPED replaces outcome Y with Y - theta*(X - mean X), where X is a pre-experiment covariate. With theta = cov(Y, X) / var(X), variance falls to (1 - rho^2) of the original, rho being the correlation.
In a group-sequential A/B test, what does an alpha-spending function do?
basics
~20 sAn alpha-spending function states in advance how much of the total 0.05 false-positive budget each interim analysis may consume, as a function of how much of the planned data has arrived. The pieces sum to 0.05.
Why is the sample mean of an arm abandoned early by a bandit allocator biased downward?
basics
~20 sBecause allocation depended on outcomes: an arm gets starved precisely after a run of unlucky results, so the few observations it holds are the unlucky ones. Its sample mean understates its true mean, and ordinary confidence intervals under-cover.
What does cumulative regret measure in a multi-armed bandit experiment?
basics
~10 sCumulative regret is the total reward given up by not serving the best arm every round: summed over rounds, the best arm's true mean reward minus the mean reward of the arm actually served.