skip to content

Minimum Detectable Effect

Working backwards from the smallest lift worth shipping to the sample each arm needs, given baseline conversion, metric variance, alpha and power. Interviewers push on the tradeoffs between them.

on this pageshow

questions

5

What is the minimum detectable effect (MDE) in an A/B test sample-size calculation?

level: juniorimportance: must knowfreq 82%

answer

  1. a planning input, not an outcome
  2. smallest lift the design can catch
  3. needs baseline, alpha, power fixed
  4. traffic and sensitivity trade off
  5. detection probability equals target power

basics

~20 s

The minimum detectable effect is the smallest true difference between variants a test is designed to catch. Given the sample size, significance level and target power, a true effect smaller than the MDE will usually go undetected.

solid answer

~40 s

MDE is a planning input, not a result. Before the test runs you fix the significance level `alpha`, the target power (commonly 80%), the baseline metric level and its variance, and the traffic split. MDE and sample size are then two views of one equation: name the effect you must be able to see and it tells you how much traffic you need; name the traffic you have and it tells you the smallest effect you can see. Formally, the MDE is the true effect at which the test's probability of returning a significant result equals the target power. It is not a significance threshold and not a forecast of the real lift. In practice you anchor it on business value: the smallest lift that would actually justify shipping the change.

go deeper

for a junior

Be ready to define it in one sentence and say which quantities you must fix first: baseline, variance, significance level, power and the split.

for a middle

An interviewer expects the mechanics: the sample-size formula, the rearrangement that turns it into an MDE, and the quadratic relationship between effect and traffic.

for a senior

Show that you set the MDE before the test, from what the business would act on, and that you catch design details that quietly worsen it, such as counting sessions while randomising visitors.

for a principal

Own the framing that MDE is the sensitivity your experimentation programme can afford, and be able to argue when a decision simply should not be routed through a test at all.

## The idea in one line The **minimum detectable effect (MDE)** is the smallest *true* difference between control and treatment that an experiment is built to detect reliably. "Reliably" is quantified: at the chosen significance level, with the planned sample, a true effect of exactly the MDE would produce a statistically significant result with probability equal to the target power (80% is the usual convention). It is a **design parameter**, decided before a single user is bucketed, and it is the number that turns a vague ambition ("let's see if the new checkout is better") into a concrete traffic and calendar commitment. ## The equation it comes from For a two-arm test with an equal split, comparing a metric with standard deviation `sigma` between arms, the classical per-arm sample size is `n = 2 * (z_alpha/2 + z_beta)^2 * sigma^2 / delta^2` where `delta` is the effect you want to detect, `z_alpha/2` is the critical value for the two-sided significance level, and `z_beta` is the value corresponding to the target power. For a conversion rate `p`, the variance term is `sigma^2 = p * (1 - p)`. Rearranged for `delta`, the same relationship gives the MDE directly: `delta = (z_alpha/2 + z_beta) * sigma * sqrt(2 / n)` That single rearrangement is the whole concept. Sample size and MDE are not two topics; they are the same equation solved for different unknowns. Two consequences follow immediately: - **MDE shrinks only as `1 / sqrt(n)`.** Four times the traffic halves the MDE. Doubling traffic improves it by about 29%, not 50%. - **Sample size grows as `1 / delta^2`.** Insisting on detecting an effect half as large costs four times the users. ## What you must fix before an MDE means anything An MDE quoted on its own is uninterpretable. To state it you need: 1. **Baseline metric level** — the control conversion rate or mean. It sets the variance and it is the reference for any relative statement. 2. **Variance of the metric** — for a proportion it follows from the baseline; for revenue-style metrics it must be measured from history, and heavy tails make it much larger than people guess. 3. **Significance level `alpha`** — commonly 0.05 two-sided. 4. **Target power** — commonly 0.80. 5. **Traffic split and number of arms** — the per-arm sample, not the total, drives sensitivity. 6. **Unit of analysis** — visitors, sessions or accounts. Randomising by visitor but counting sessions inflates the effective variance and quietly makes the real MDE worse than the plan says. 7. **Direction** — one-sided or two-sided, since the critical value differs. ## What the MDE is not - **Not a prediction.** Saying "our MDE is a 5% lift" says nothing about how large the true lift is. It describes the instrument, not the world. - **Not a significance threshold.** A measured difference smaller than the MDE can still be significant on a given run, and a measured difference larger than the MDE can fail to reach significance. The MDE is a statement about long-run detection probability for a fixed *true* effect, not a cutoff applied to the observed number. - **Not a floor on what the product can achieve.** A change with a true 1% lift is still worth having; a test sized for 9% simply cannot see it. - **Not free to shrink.** The only ways down are more traffic, less variance, or accepting weaker guarantees. ## Why interviewers ask it Because it separates people who have planned an experiment from people who have only read one out. The follow-up is almost always numerical ("what happens to sample size if you halve it?") and the expected answer is the quadratic relationship. The second-order signal is whether you know MDE is chosen from business value, not reverse-engineered from whatever traffic happens to exist and then written into the plan as if it were a decision. ## A worked intuition Suppose you plan a test and the arithmetic says the MDE is a 9% relative lift. That means: if the treatment truly moves the metric by 9%, roughly four runs in five would come back significant; if it truly moves it by 3%, most runs would not. Before spending two weeks of traffic, the team should ask whether a change that only delivers 3% is worth knowing about. If it is, the design is wrong, and finding that out beforehand is exactly what the MDE is for.

  • If you halve the MDE and change nothing else, what happens to the required sample size?
    It quadruples. Sample size scales with `1 / delta^2`, so cutting the target effect in half multiplies the traffic requirement by four. Read the other way, quadrupling traffic only halves the MDE, which is why sensitivity gets expensive fast and why an extra week rarely rescues a badly sized test.
  • Is the MDE the smallest difference that can ever come back significant?
    No. The MDE is defined on the true effect, not the measured one. A true effect below the MDE still has some chance of producing a significant result, and a true effect at the MDE fails to do so roughly one run in five at 80% power. It describes detection probability, not a cutoff.
  • Two teams quote a 5% MDE on different metrics. Are those comparable?
    Not without the baseline and variance. A 5% relative move on a 3% conversion rate is a tiny absolute shift with modest variance; a 5% move on average revenue per user sits on a heavy-tailed distribution and can need far more traffic. The MDE only means something alongside the metric it was computed for.

The MDE is the resolution of your measuring instrument. A bathroom scale that reads to the nearest kilogram cannot tell you that you lost 200 grams, however carefully you stand on it.

saying these in an interview costs you the question

  • Calls the MDE the lift the test actually measured
  • Treats MDE as a significance cutoff for the observed difference
  • Quotes an MDE without naming baseline, alpha or power
  • Thinks doubling traffic halves the MDE
  • Reverse-engineers the MDE from available traffic and calls it a target

context

open as a page

A test spec says 'detect a 5% lift' on a 3% conversion rate. Why is that underspecified?

level: middleimportance: must knowfreq 68%

basics

~10 s

It never says whether 5% means 5 percentage points, 3% to 8%, or a 5% relative lift, 3% to 3.15%. Those targets differ by hundreds of times in required traffic, so sizing cannot start.

open as a page

What MDE can a 50/50 test reach at a 5% baseline with 40,000 weekly visitors over two weeks?

level: seniorimportance: should knowfreq 54%

basics

~20 s

About 0.44 percentage points, roughly a 9% relative lift. Two weeks gives 80,000 visitors, 40,000 per arm; inverting the sizing rule at 5% significance and 80% power gives an absolute detectable difference near 0.0044 on a 5% baseline.

open as a page

How do you decide the MDE an experiment should be powered for when traffic is scarce?

level: principalimportance: should knowfreq 46%

basics

~20 s

Anchor the MDE on the smallest lift that would change the ship decision, then check whether the traffic supports it. If it does not, the honest options are a bigger bet, a more sensitive metric, or not testing.

open as a page

Where does the 16 come from in the per-arm sample rule n = 16 * sigma^2 / delta^2?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

The 16 is 2 x (1.96 + 0.84)^2 rounded up. 1.96 is the two-sided critical value at 5% significance, 0.84 corresponds to 80% power, and the factor 2 covers two equally sized arms each contributing sampling error.

open as a page