What MDE can a 50/50 test reach at a 5% baseline with 40,000 weekly visitors over two weeks?
answer
- same equation, other unknown
- per-arm sample, not total
- delta = 4 sigma over root n
- variance is p times one minus p
- answer near half a percentage point
basics
~20 sAbout 0.44 percentage points, roughly a 9% relative lift. Two weeks gives 80,000 visitors, 40,000 per arm; inverting the sizing rule at 5% significance and 80% power gives an absolute detectable difference near 0.0044 on a 5% baseline.
solid answer
~50 sRun the sizing rule backwards. Two weeks of 40,000 visitors is 80,000 total, 40,000 per arm on an even split. From `n = 16 * sigma^2 / delta^2`, the detectable effect is `delta = 4 * sigma / sqrt(n)`. For a 5% baseline, `sigma^2 = 0.05 x 0.95 = 0.0475`, so `delta = 4 * sqrt(0.0475 / 40000) = 0.0044` - about **0.44 percentage points, a 9% relative lift**. The judgment matters more than the arithmetic: report it back as "this window can only see lifts of roughly 9% or larger", and let the team decide whether a change plausibly that big is on the roadmap. Two caveats: only visitors who actually reach the tested surface count, and sensitivity improves only as the square root of traffic, so doubling the window buys about 29%, not half.
go deeper
Be ready to divide total traffic by the number of arms first, and to compute the variance of a rate as p times one minus p before touching the formula.
Explain the inversion cleanly: delta equals four sigma over the square root of the per-arm sample, and report the answer in both percentage points and relative terms.
Show the judgment around the number: check the exposed population, confirm the unit of analysis, and turn the result into a go or no-go recommendation rather than a bare figure.
Own the conversation where the achievable sensitivity does not match the ambition, and be ready to argue for concentrating traffic on fewer, larger bets instead of many underpowered ones.
## Inverting the sizing question Most sizing conversations start from an ambition and ask for traffic. The more useful direction, when traffic is what you actually have, is the reverse: **given the users I can get, what is the smallest lift I could detect?** It is the same equation solved for the other unknown. Starting from `n = 16 * sigma^2 / delta^2` (two arms, equal split, two-sided 5% significance, 80% power): `delta = sqrt(16 * sigma^2 / n) = 4 * sigma / sqrt(n)` where `n` is the **per-arm** sample. ## The worked case - Traffic: 40,000 visitors per week for two weeks = **80,000 total**. - Even split: **40,000 per arm**. - Baseline: a 5% signup rate, so `sigma^2 = p(1-p) = 0.05 x 0.95 = 0.0475`. `delta = 4 * sqrt(0.0475 / 40,000) = 4 * sqrt(0.0000011875) = 4 x 0.00109 = 0.00436` That is **0.44 percentage points** in absolute terms - a move from 5.00% to about 5.44%. As a relative lift: `0.00436 / 0.05 = 8.7%`, so call it **a 9% relative lift**. ## What to do with that number The arithmetic is the easy half. The senior behaviour is what happens next: 1. **State it as a constraint, not a result.** "With this traffic and window, we can reliably detect a lift of about 9% or bigger. Anything smaller will probably not show up." 2. **Ask whether 9% is plausible.** Most product changes on a mature funnel move things by low single-digit percentages. If nobody in the room believes the change could deliver 9%, the test as designed is close to a coin flip dressed as evidence, and that is worth knowing before the window is spent. 3. **Check the eligible population.** The 40,000 figure is site visitors. If only 60% of them reach the tested surface, the effective per-arm sample is 24,000 and the MDE rises to `4 * sqrt(0.0475/24,000) = 0.0056`, about 0.56pp or an 11% relative lift. Sizing on all traffic when only a subset is exposed is the most common way a plan turns out weaker than promised. 4. **Confirm the unit of analysis.** If randomisation is by visitor, the variance and the count must both be per visitor. Counting sessions inside a visitor-randomised test understates the true variance and overstates sensitivity. ## The square-root wall Because `delta` scales as `1 / sqrt(n)`: | Per-arm sample | MDE (absolute) | MDE (relative) | |---|---|---| | 20,000 | 0.62pp | 12.3% | | 40,000 | 0.44pp | 8.7% | | 80,000 | 0.31pp | 6.2% | | 160,000 | 0.22pp | 4.4% | Doubling the traffic improves sensitivity by only about 29%. Getting from a 9% MDE to a 2% MDE needs roughly **twenty times** the sample. This table is the single most useful thing to have internalised, because it kills the reflexive "let's just run it a bit longer" response - a bit longer buys almost nothing, and going from underpowered to adequately powered usually needs a different plan, not a longer one. ## Splitting the same traffic more ways If the same 80,000 visitors are spread over **four arms** instead of two, each arm holds 20,000 and the per-comparison MDE becomes `4 * sqrt(0.0475/20,000) = 0.0062`, about 0.62pp or a 12.3% relative lift. Halving the per-arm sample multiplies the MDE by `sqrt(2)`, roughly 40% worse - a real cost of testing four ideas at once on fixed traffic, before any adjustment for comparing several variants against the same control is considered. ## Sanity checks before quoting the number - **Is the baseline current?** A stale rate, or one measured on a different traffic mix, will misstate both the variance and any relative target derived from it. - **Is the metric actually binary?** Revenue per visitor is not a proportion; its variance must come from historical data and is often far larger relative to its mean, which pushes the MDE up sharply. - **Is the split really even?** A 90/10 holdout has a small arm that dominates the standard error, and the effective sensitivity is much closer to what the 10% arm alone can support. ## What a strong answer sounds like Compute it, state it in **both** absolute and relative terms (because stakeholders think in one and the mathematics in the other), name the assumptions you used - even split, two arms, 5% significance, 80% power, all visitors exposed - and finish with the decision the number forces: run it, redesign it, or spend the traffic on a bigger bet.
- The same 80,000 visitors are split across four arms instead of two. What happens to the MDE?Each arm drops from 40,000 to 20,000, and since the MDE scales as `1/sqrt(n)`, it grows by a factor of `sqrt(2)`, roughly 40%. On the 5% baseline it moves from about 0.44pp to about 0.62pp, a 12.3% relative lift. Testing four ideas at once costs real sensitivity on every one of them.
- Stakeholders ask to double the traffic to halve the MDE. What do you tell them?Doubling the sample improves the MDE by about 29%, not 50%, because sensitivity scales with the square root of n. Halving the MDE needs four times the traffic. That quadratic wall is usually the argument for changing the design - a bigger intended effect, a lower-variance metric, or a broader exposed population - rather than simply running longer.
- Only 60% of those visitors ever reach the tested surface. How does that change the answer?Only exposed users carry information, so the per-arm sample falls from 40,000 to 24,000 and the MDE rises to about 0.56pp, roughly an 11% relative lift. Sizing against total site traffic rather than the eligible, exposed population is the most common reason a test turns out less sensitive than the plan claimed.
saying these in an interview costs you the question
- Uses total traffic where the per-arm sample belongs
- Quotes only an absolute or only a relative number
- Assumes doubling traffic halves the MDE
- Sizes on all site visitors rather than exposed users
- Ignores that the split determines the smallest arm's power