skip to content

A/B Testing & Experiment Design

You will learn to design a trustworthy online experiment end to end: pick the randomization unit, size the sample, choose metrics and guardrails, and avoid peeking, SRM, and novelty traps. This is the signature interview loop for data scientist and analyst roles — 'design an A/B test for feature X' is near-guaranteed.

on this pageshow

explore

questions

112 · 6 sections

Why must an A/B test give the same user the same variant on every visit?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Sticky assignment gives each user one consistent experience instead of a flickering mix. If the variant is re-drawn per visit, returning users receive both treatments, which blends the two arms and shrinks the measured difference toward zero.

open as a page

What do experiment layers do in an A/B platform running many tests at once?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Layers let many experiments share the same traffic at the same time. Each layer randomizes users independently, so a user's arm in one layer says nothing about their arm in another, and concurrent tests do not confound one another.

open as a page

Why ramp an A/B test through 1%, 5%, 25% and 50% instead of starting at full traffic?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A staged ramp limits blast radius: the first small stage exposes few users, so crashes, latency spikes or a collapsing metric are caught cheaply. Later, larger stages exist to measure the effect once the change has been shown to be safe.

open as a page

How does hash-based bucketing turn a user ID into an A/B test variant?

level: middleimportance: must knowfreq 68%
basics
~20 s

Concatenate the user ID with an experiment-specific salt, hash it to an integer, take it modulo the bucket count, and map bucket ranges to arms - 0-49 control, 50-99 treatment. Deterministic, so nothing is stored.

open as a page

When would you randomize an A/B test by session rather than by user?

level: middleimportance: must knowfreq 70%
basics
~20 s

Randomize by session only when the change leaves no trace between visits and the outcome is finished within one visit — a layout tweak judged by same-visit conversion. If the treatment can be learned, remembered or accumulated, randomize by user.

open as a page

Why do A/B tests usually run for whole weeks rather than an arbitrary number of days?

level: juniorimportance: must knowfreq 72%
basics
~20 s

User behaviour differs by day of week, so a partial week samples an unrepresentative mix. Whole multiples of seven days weight every weekday equally, so the measured lift describes a typical week, not whichever days happened to be included.

open as a page

What is the minimum detectable effect (MDE) in an A/B test sample-size calculation?

level: juniorimportance: must knowfreq 82%
basics
~20 s

The minimum detectable effect is the smallest true difference between variants a test is designed to catch. Given the sample size, significance level and target power, a true effect smaller than the MDE will usually go undetected.

open as a page

A test spec says 'detect a 5% lift' on a 3% conversion rate. Why is that underspecified?

level: middleimportance: must knowfreq 68%
basics
~10 s

It never says whether 5% means 5 percentage points, 3% to 8%, or a 5% relative lift, 3% to 3.15%. Those targets differ by hundreds of times in required traffic, so sizing cannot start.

open as a page

Your experiment hits its planned sample size in three days. Why keep it running to the planned end date?

level: middleimportance: should knowfreq 58%
basics
~20 s

A sample count and a calendar window are different requirements. Three days covers only three weekdays and over-represents your most frequent visitors, so the sample is large but not representative. Read the result when both are met.

open as a page

A planned two-week test window overlaps Black Friday. How do you handle the runtime?

level: seniorimportance: should knowfreq 44%
basics
~20 s

Decide before launch. The cleanest option is usually to shift the window so it sits entirely before or after the holiday. If the test must span it, keep whole weeks and treat the result as an estimate for holiday traffic.

open as a page

In an A/B test, what is the difference between a count, a rate, and a share metric?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A count metric sums events, such as total orders. A rate metric divides a count by the units that were exposed, such as orders per visitor. A share metric divides one count by a related total, such as the percentage of orders placed on mobile.

open as a page

What is a guardrail metric in an A/B test, and how does it differ from a driver metric?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A guardrail is a metric the experiment must not damage — latency, crash-free sessions, unsubscribes — and a big enough degradation can veto the launch. A driver metric only explains why the primary metric moved; it never vetoes.

open as a page

Why does revenue per user make an A/B test hard to call when a few users spend far more than the rest?

level: juniorimportance: must knowfreq 74%
basics
~10 s

Almost all the variance in revenue per user comes from a handful of very large spenders, so the interval around the treatment-control difference stays wide. Ordinary experiment sizes then cannot resolve realistic revenue effects.

open as a page

What is an Overall Evaluation Criterion (OEC) in an A/B test?

level: juniorimportance: must knowfreq 74%
basics
~20 s

The Overall Evaluation Criterion is the single metric, agreed before launch, that an experiment's ship-or-not decision is defined against. It expresses what the team means by the product getting better, so the result cannot be reinterpreted afterwards.

open as a page

What does a DAU/MAU stickiness ratio of 0.2 tell you about how a product is used?

level: juniorimportance: must knowfreq 68%
basics
~10 s

A DAU/MAU of 0.2 means the typical monthly active user opens the product on about 6 days out of 30. It measures return frequency, not audience size, and averages over very different user types.

open as a page

Why report a confidence interval on A/B lift instead of just a p-value?

level: middleimportance: must knowfreq 72%
basics
~20 s

A p-value only says how compatible the data are with zero effect. A confidence interval reports the lift itself with its precision, so you can compare the plausible range against the minimum detectable effect and the launch bar.

open as a page

Why does slicing a flat A/B test result into many segments so often produce a false winner?

level: middleimportance: must knowfreq 72%
basics
~20 s

Every segment cut is another hypothesis test. At a 5% false-positive rate, about twenty cuts of a genuinely flat experiment throw up roughly one significant segment by chance alone, so the winner you find is usually noise.

open as a page

What is dilution in an A/B test where only 3% of assigned users ever see the change?

level: middleimportance: must knowfreq 72%
basics
~20 s

Dilution is the shrinking of a measured effect when most assigned users never encounter the change. At a 3% trigger rate an effect among exposed users appears about thirty times smaller in the all-up read, while the noise stays.

open as a page

In an A/B readout, how does an intent-to-treat analysis differ from a triggered-only analysis?

level: middleimportance: must knowfreq 60%
basics
~20 s

Intent-to-treat counts every assigned user and answers what shipping does to the whole population. A triggered analysis keeps only users who met the trigger condition in both arms, answering what the change does to those who encounter it.

open as a page

An A/B test reads +0.4% lift, 95% CI [-2.0%, +2.8%], planned MDE 2% — is that flat?

level: seniorimportance: must knowfreq 58%
basics
~20 s

No — it is inconclusive, not flat. The interval reaches +2.8%, above the 2% minimum detectable effect the test was planned around, so a lift worth shipping has not been ruled out, and neither has a 2% loss.

open as a page

What are novelty and primacy effects in an A/B test?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Both are temporary reactions to the fact that something changed, not to what it changed into. Novelty inflates an early result as users explore something new; primacy, or change aversion, depresses it as users unlearn a habit.

open as a page

On day 2 of a planned 14-day A/B test the dashboard shows p = 0.04 — do you ship?

level: juniorimportance: must knowfreq 66%
basics
~20 s

No. The 0.05 threshold is calibrated for a single analysis at the planned sample size, so a day-2 reading is not the test you designed. Run to the agreed horizon unless the plan already allowed an early call.

open as a page

What is a sample ratio mismatch in an A/B test?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A sample ratio mismatch is when the observed split of users across experiment arms differs from the planned split by more than chance can explain. It signals a broken assignment or logging pipeline, so the comparison itself is untrustworthy.

open as a page

In a marketplace test where treated sellers win demand from control sellers, which way is the effect biased?

level: middleimportance: must knowfreq 66%
basics
~20 s

Upward. Treated sellers capture bookings that control sellers would otherwise have received, so treatment is pushed up while control is pushed down. The measured lift counts one transfer twice and mixes real growth with redistribution that vanishes at full rollout.

open as a page

How do you check whether an A/B test effect is decaying over exposure time?

level: middleimportance: must knowfreq 56%
basics
~20 s

Put every user on their own clock: days since that user first saw the change. Plot the treatment-minus-control gap against exposure age, grouped by first-exposure date. A gap that shrinks with exposure age, repeating per group, indicates decay.

open as a page

Would you use a bandit allocator or a fixed-horizon A/B test for twelve news headlines whose value expires within hours?

level: middleimportance: must knowfreq 66%
basics
~20 s

Use adaptive allocation. With twelve options, clicks arriving in seconds and value gone in hours, the cost of serving losers dominates and no precise per-headline lift is needed. A fixed-horizon test would still be running when the headlines are worthless.

open as a page

What is CUPED, and how does a pre-experiment covariate reduce the variance of an A/B test?

level: middleimportance: must knowfreq 62%
basics
~20 s

CUPED replaces outcome Y with Y - theta*(X - mean X), where X is a pre-experiment covariate. With theta = cov(Y, X) / var(X), variance falls to (1 - rho^2) of the original, rho being the correlation.

open as a page

In a group-sequential A/B test, what does an alpha-spending function do?

level: middleimportance: must knowfreq 55%
basics
~20 s

An alpha-spending function states in advance how much of the total 0.05 false-positive budget each interim analysis may consume, as a function of how much of the planned data has arrived. The pieces sum to 0.05.

open as a page

Why is the sample mean of an arm abandoned early by a bandit allocator biased downward?

level: seniorimportance: must knowfreq 50%
basics
~20 s

Because allocation depended on outcomes: an arm gets starved precisely after a run of unlucky results, so the few observations it holds are the unlucky ones. Its sample mean understates its true mean, and ordinary confidence intervals under-cover.

open as a page

What does cumulative regret measure in a multi-armed bandit experiment?

level: juniorimportance: should knowfreq 42%
basics
~10 s

Cumulative regret is the total reward given up by not serving the best arm every round: summed over rounds, the best arm's true mean reward minus the mean reward of the arm actually served.

open as a page