skip to content

What MDE can a 50/50 test reach at a 5% baseline with 40,000 weekly visitors over two weeks?

level: seniorimportance: should knowfreq 54%

answer

  1. same equation, other unknown
  2. per-arm sample, not total
  3. delta = 4 sigma over root n
  4. variance is p times one minus p
  5. answer near half a percentage point

basics

~20 s

About 0.44 percentage points, roughly a 9% relative lift. Two weeks gives 80,000 visitors, 40,000 per arm; inverting the sizing rule at 5% significance and 80% power gives an absolute detectable difference near 0.0044 on a 5% baseline.

solid answer

~50 s

Run the sizing rule backwards. Two weeks of 40,000 visitors is 80,000 total, 40,000 per arm on an even split. From `n = 16 * sigma^2 / delta^2`, the detectable effect is `delta = 4 * sigma / sqrt(n)`. For a 5% baseline, `sigma^2 = 0.05 x 0.95 = 0.0475`, so `delta = 4 * sqrt(0.0475 / 40000) = 0.0044` - about **0.44 percentage points, a 9% relative lift**. The judgment matters more than the arithmetic: report it back as "this window can only see lifts of roughly 9% or larger", and let the team decide whether a change plausibly that big is on the roadmap. Two caveats: only visitors who actually reach the tested surface count, and sensitivity improves only as the square root of traffic, so doubling the window buys about 29%, not half.

go deeper

for a junior

Be ready to divide total traffic by the number of arms first, and to compute the variance of a rate as p times one minus p before touching the formula.

for a middle

Explain the inversion cleanly: delta equals four sigma over the square root of the per-arm sample, and report the answer in both percentage points and relative terms.

for a senior

Show the judgment around the number: check the exposed population, confirm the unit of analysis, and turn the result into a go or no-go recommendation rather than a bare figure.

for a principal

Own the conversation where the achievable sensitivity does not match the ambition, and be ready to argue for concentrating traffic on fewer, larger bets instead of many underpowered ones.

## Inverting the sizing question Most sizing conversations start from an ambition and ask for traffic. The more useful direction, when traffic is what you actually have, is the reverse: **given the users I can get, what is the smallest lift I could detect?** It is the same equation solved for the other unknown. Starting from `n = 16 * sigma^2 / delta^2` (two arms, equal split, two-sided 5% significance, 80% power): `delta = sqrt(16 * sigma^2 / n) = 4 * sigma / sqrt(n)` where `n` is the **per-arm** sample. ## The worked case - Traffic: 40,000 visitors per week for two weeks = **80,000 total**. - Even split: **40,000 per arm**. - Baseline: a 5% signup rate, so `sigma^2 = p(1-p) = 0.05 x 0.95 = 0.0475`. `delta = 4 * sqrt(0.0475 / 40,000) = 4 * sqrt(0.0000011875) = 4 x 0.00109 = 0.00436` That is **0.44 percentage points** in absolute terms - a move from 5.00% to about 5.44%. As a relative lift: `0.00436 / 0.05 = 8.7%`, so call it **a 9% relative lift**. ## What to do with that number The arithmetic is the easy half. The senior behaviour is what happens next: 1. **State it as a constraint, not a result.** "With this traffic and window, we can reliably detect a lift of about 9% or bigger. Anything smaller will probably not show up." 2. **Ask whether 9% is plausible.** Most product changes on a mature funnel move things by low single-digit percentages. If nobody in the room believes the change could deliver 9%, the test as designed is close to a coin flip dressed as evidence, and that is worth knowing before the window is spent. 3. **Check the eligible population.** The 40,000 figure is site visitors. If only 60% of them reach the tested surface, the effective per-arm sample is 24,000 and the MDE rises to `4 * sqrt(0.0475/24,000) = 0.0056`, about 0.56pp or an 11% relative lift. Sizing on all traffic when only a subset is exposed is the most common way a plan turns out weaker than promised. 4. **Confirm the unit of analysis.** If randomisation is by visitor, the variance and the count must both be per visitor. Counting sessions inside a visitor-randomised test understates the true variance and overstates sensitivity. ## The square-root wall Because `delta` scales as `1 / sqrt(n)`: | Per-arm sample | MDE (absolute) | MDE (relative) | |---|---|---| | 20,000 | 0.62pp | 12.3% | | 40,000 | 0.44pp | 8.7% | | 80,000 | 0.31pp | 6.2% | | 160,000 | 0.22pp | 4.4% | Doubling the traffic improves sensitivity by only about 29%. Getting from a 9% MDE to a 2% MDE needs roughly **twenty times** the sample. This table is the single most useful thing to have internalised, because it kills the reflexive "let's just run it a bit longer" response - a bit longer buys almost nothing, and going from underpowered to adequately powered usually needs a different plan, not a longer one. ## Splitting the same traffic more ways If the same 80,000 visitors are spread over **four arms** instead of two, each arm holds 20,000 and the per-comparison MDE becomes `4 * sqrt(0.0475/20,000) = 0.0062`, about 0.62pp or a 12.3% relative lift. Halving the per-arm sample multiplies the MDE by `sqrt(2)`, roughly 40% worse - a real cost of testing four ideas at once on fixed traffic, before any adjustment for comparing several variants against the same control is considered. ## Sanity checks before quoting the number - **Is the baseline current?** A stale rate, or one measured on a different traffic mix, will misstate both the variance and any relative target derived from it. - **Is the metric actually binary?** Revenue per visitor is not a proportion; its variance must come from historical data and is often far larger relative to its mean, which pushes the MDE up sharply. - **Is the split really even?** A 90/10 holdout has a small arm that dominates the standard error, and the effective sensitivity is much closer to what the 10% arm alone can support. ## What a strong answer sounds like Compute it, state it in **both** absolute and relative terms (because stakeholders think in one and the mathematics in the other), name the assumptions you used - even split, two arms, 5% significance, 80% power, all visitors exposed - and finish with the decision the number forces: run it, redesign it, or spend the traffic on a bigger bet.

  • The same 80,000 visitors are split across four arms instead of two. What happens to the MDE?
    Each arm drops from 40,000 to 20,000, and since the MDE scales as `1/sqrt(n)`, it grows by a factor of `sqrt(2)`, roughly 40%. On the 5% baseline it moves from about 0.44pp to about 0.62pp, a 12.3% relative lift. Testing four ideas at once costs real sensitivity on every one of them.
  • Stakeholders ask to double the traffic to halve the MDE. What do you tell them?
    Doubling the sample improves the MDE by about 29%, not 50%, because sensitivity scales with the square root of n. Halving the MDE needs four times the traffic. That quadratic wall is usually the argument for changing the design - a bigger intended effect, a lower-variance metric, or a broader exposed population - rather than simply running longer.
  • Only 60% of those visitors ever reach the tested surface. How does that change the answer?
    Only exposed users carry information, so the per-arm sample falls from 40,000 to 24,000 and the MDE rises to about 0.56pp, roughly an 11% relative lift. Sizing against total site traffic rather than the eligible, exposed population is the most common reason a test turns out less sensitive than the plan claimed.

saying these in an interview costs you the question

  • Uses total traffic where the per-arm sample belongs
  • Quotes only an absolute or only a relative number
  • Assumes doubling traffic halves the MDE
  • Sizes on all site visitors rather than exposed users
  • Ignores that the split determines the smallest arm's power

context