skip to content

How do you decide the MDE an experiment should be powered for when traffic is scarce?

level: principalimportance: should knowfreq 46%

answer

  1. value first, traffic second
  2. two numbers to compare
  3. smallest lift that changes the decision
  4. never reverse-engineer from traffic
  5. not testing is a valid outcome

basics

~20 s

Anchor the MDE on the smallest lift that would change the ship decision, then check whether the traffic supports it. If it does not, the honest options are a bigger bet, a more sensitive metric, or not testing.

solid answer

~50 s

Start from value, not from traffic. Ask what lift would make the change worth building, maintaining and carrying forward - that number is the MDE the decision actually needs. Then compute what the available traffic and window can detect and compare the two. If the achievable MDE is comfortably smaller, size and run. If it is much larger, say so plainly rather than quietly writing the achievable number into the plan as though it were a target: reverse-engineering the MDE from traffic manufactures a test that is powered only for effects nobody expects. The remaining levers are real but limited: concentrate traffic on fewer bets, widen the exposed population, pick a lower-variance or more proximate metric, pursue a larger design change, or accept the decision on judgment and evidence outside the test. Record the chosen MDE and its inputs in the plan so the design's sensitivity is on the record.

go deeper

for a junior

Be ready to say that the MDE should come from what the business would act on, and that the traffic calculation is a separate check performed afterwards.

for a middle

Explain how to turn a business threshold into an absolute effect using the baseline, and how to compare it against what the available per-arm sample can detect.

for a senior

Demonstrate the levers you reach for when the gap is large: concentrating traffic, widening exposure, choosing a lower-variance metric, or designing a bolder change.

for a principal

Own the allocation call across the programme: which surfaces are experimentable at all, how many concurrent tests a surface can support, and when a decision should be made without an experiment.

## The two MDEs, and why they must be compared Every experiment plan has two numbers that are easy to confuse: 1. **The MDE the decision needs** - the smallest true lift that would change what you do. Below it, you would not ship the change even if it were real, because it does not repay the build cost, the added complexity, the ongoing maintenance and the opportunity cost of the surface it occupies. 2. **The MDE the traffic supports** - what the available exposed users, split and window can actually detect at the chosen significance and power. Planning is the act of comparing them. When the achievable MDE is smaller than the needed one, the test is worth running. When it is larger, the test cannot answer the question, and the only defensible move is to say so. The common failure is silent substitution: someone computes what the traffic can see, writes that into the plan as "our MDE", and the document now reads as if a deliberate choice were made. Nobody lied, but the plan has lost the information that would let a reader judge whether the experiment was worth the window. ## How to derive the value-side number Make it concrete rather than philosophical: - **Payback framing.** Estimate the annualised value of a 1% relative lift on this metric. Divide the fully loaded build-and-maintain cost by that, and you have the lift at which the change breaks even. That is the floor. - **Decision framing.** Ask the decision-maker directly: "if the true lift were 1%, would you ship?" Then 2%, then 5%. The point where the answer flips is the MDE the decision needs, and it is often smaller than engineers assume and larger than product assumes. - **Portfolio framing.** If a surface receives dozens of changes a year, a 1% lift may genuinely matter in aggregate; if it receives one change every two years, only a large move justifies the effort at all. This conversation is worth having before any arithmetic, because it is the only input to sizing that statistics cannot supply. ## When the achievable MDE is too coarse The levers, roughly in order of how often they are the right answer: - **Concentrate traffic.** Running three experiments on one surface at once splits the population; running them sequentially, or accepting fewer bets per quarter, gives each one a real chance of a conclusion. A portfolio of underpowered tests produces mostly noise. - **Widen exposure.** Often only a fraction of traffic ever reaches the tested surface. Moving the change earlier in the funnel, or removing an unnecessary eligibility filter, can add more effective sample than any calendar change. - **Change the metric.** A metric closer to the intervention - a step completion rather than end-to-end revenue - usually has both a higher base rate and lower variance, so it is far cheaper to move detectably. The trade is that it is a weaker proxy for the outcome you actually care about, and that trade must be made explicitly. - **Aim bigger.** If only a large effect is detectable, build a change that could plausibly produce a large effect. A bold redesign is a better use of a scarce, low-traffic surface than a copy tweak. - **Relax the guarantees deliberately.** Accepting a higher significance level or lower power is a legitimate choice when the cost of a wrong decision is low and reversible, but it must be a stated decision with its consequences named, not a quiet edit to make the calculator return a friendly number. - **Do not test.** Some decisions are too small, too cheap to reverse, or too obviously correct to be worth a scarce experimental slot. Deciding on judgment and moving on is a legitimate outcome of a sizing exercise, and a leader should be comfortable saying it. ## Guardrails need their own answer A test is usually powered for a primary metric, while guardrails - latency, error rate, churn, refunds - ride along. Those are typically powered to detect only large regressions, which is a deliberate asymmetry: you are not trying to prove a guardrail unchanged, you are trying to catch a serious break. Stating that asymmetry in the plan prevents someone later reading unmoved guardrails as proof of safety at a sensitivity the test never had. ## Organisational hygiene - **Write the MDE and its inputs down** - baseline, variance, alpha, power, split, exposed population - in the plan before the test starts. - **Standardise the defaults** so that plans are comparable across teams, and make deviations explicit. - **Track the distribution of achievable MDEs** across your surfaces. If most sit far above what any realistic change would deliver, the constraint is structural, and the fix is a programme decision - fewer concurrent tests, better metrics, variance reduction, or accepting that some surfaces are not experimentable. ## What a strong answer sounds like A principal-level answer treats sizing as a resource-allocation decision, not a calculator exercise: it names the value-side number first, compares it to the achievable one, is explicit that the gap is a finding rather than an obstacle to negotiate away, and lands on a recommendation that includes "do not run this test" as a live option.

  • The traffic supports a 9% MDE but only a 2% lift is needed to justify shipping. What do you recommend?
    Say plainly that the experiment cannot answer the question as designed. Then work the levers: concentrate traffic on fewer concurrent bets, widen the exposed population, move to a lower-variance or more proximate metric, or aim at a larger design change. If none of those close the gap, recommend deciding on judgment rather than spending a window on a test that is mostly noise.
  • Someone proposes lowering power from 80% to 60% so the plan fits the traffic. How do you respond?
    It is a real lever, not a trick, but it must be a stated decision with its cost named: at 60% power, four in ten genuine effects at the target size go undetected. That can be acceptable for a cheap, reversible change, and it is indefensible for a decision that will stand for years. What is not acceptable is editing the input quietly so the calculator returns a comfortable number.
  • How should guardrail metrics be sized relative to the primary metric?
    Usually to catch only large regressions. You are not proving a guardrail unchanged, you are catching a serious break, and powering every guardrail to the same sensitivity as the primary metric is unaffordable. State that asymmetry in the plan so unmoved guardrails are never read as safety evidence at a sensitivity the design never had.

saying these in an interview costs you the question

  • Sets the MDE to whatever available traffic can detect
  • Treats sizing as a calculator step, not a decision
  • Never asks what lift would change the ship decision
  • Quietly lowers power to make the plan fit
  • Cannot accept 'do not run this test' as an outcome

context