skip to content

A test spec says 'detect a 5% lift' on a 3% conversion rate. Why is that underspecified?

level: middleimportance: must knowfreq 68%

answer

  1. two readings of one number
  2. percentage points versus percent of baseline
  3. 3% to 8%, or 3% to 3.15%
  4. sizing consumes the absolute difference
  5. traffic differs by orders of magnitude

basics

~10 s

It never says whether 5% means 5 percentage points, 3% to 8%, or a 5% relative lift, 3% to 3.15%. Those targets differ by hundreds of times in required traffic, so sizing cannot start.

solid answer

~40 s

The word "5% lift" has two readings on a rate metric. **Absolute**: add 5 percentage points, 3% to 8% - an enormous change, and sizing it needs only a few hundred users per arm. **Relative**: multiply by 1.05, 3% to 3.15%, an absolute shift of 0.15 percentage points - and that needs on the order of two hundred thousand users per arm at 5% significance and 80% power. Hundreds versus hundreds of thousands: nearly three orders of magnitude apart, from one ambiguous sentence. The fix is a house convention: state absolute effects in **percentage points** (pp), state relative effects as a percent **of a named baseline**, and always quote the baseline alongside. Sample-size arithmetic runs on the absolute difference, so a relative MDE is only meaningful once multiplied by the baseline rate.

go deeper

for a junior

Be ready to state both readings with actual numbers: 3% plus five points is 8%, 3% times 1.05 is 3.15%, and say which one the spec has to pick.

for a middle

Explain that sizing consumes the absolute difference and show the quadratic consequence: a 33-fold difference in delta means a roughly thousand-fold difference in traffic.

for a senior

Demonstrate that you catch this in review before traffic is committed, and that you re-derive the absolute target when the baseline drifts rather than trusting a stale plan.

for a principal

Own the reporting convention across teams, so relative and absolute language is unambiguous in specs, dashboards and post-hoc summaries alike.

## Two readings of the same sentence On a rate metric, "a 5% lift" is genuinely ambiguous, and the two readings are not close together. | Reading | Control | Treatment | Absolute difference | |---|---|---|---| | Absolute: +5 percentage points | 3.00% | 8.00% | 0.05 | | Relative: +5% of baseline | 3.00% | 3.15% | 0.0015 | The absolute reading is a change of more than 160% in the conversion rate - the kind of move a whole new checkout flow might produce once in a company's history. The relative reading is a routine, plausible product win. Nobody writing the spec thought they were saying something ambiguous, and that is exactly why it survives to the sizing meeting. ## Why the difference is so expensive Sample size depends on the **absolute** difference `delta`, squared, in the denominator. The ratio of the two absolute differences here is `0.05 / 0.0015 = 33.3`, and sizing scales with the square of that, so the traffic requirement differs by roughly a factor of a thousand before variance effects are considered. Using the standard per-arm rule for two proportions at 5% two-sided significance and 80% power, `n = 16 * p_bar * (1 - p_bar) / delta^2` where `p_bar` is the average of the two arm rates: - **Relative 5%** (3.00% vs 3.15%): `p_bar = 0.03075`, variance term `0.0298`, `delta = 0.0015` gives roughly **210,000 users per arm**. - **Absolute 5pp** (3.00% vs 8.00%): `p_bar = 0.055`, variance term `0.052`, `delta = 0.05` gives roughly **330 users per arm**. The gap narrows slightly from the raw thousand-fold because the variance term itself grows when the treatment rate jumps to 8%: a bigger rate carries more binomial variance. Even so, the two plans differ by hundreds of times. One is a test you can read out before lunch; the other may not be runnable at all on your traffic. ## The conversion rule `absolute delta = relative lift x baseline rate` This is why a relative MDE is meaningless without the baseline attached. "We can detect a 10% lift" is not a portable statement: on a 3% checkout rate it means 0.3 percentage points, on a 40% email open rate it means 4 percentage points, and the traffic bills for those two are wildly different. It also explains a common planning surprise. A team quotes a relative MDE, the baseline then drifts downward - a traffic-mix change, a stricter definition of the converting event - and the absolute effect they are powered for shrinks with it, while the variance per user barely moves. The plan silently becomes less achievable than it looked. ## A house convention that ends the argument 1. Write absolute effects as **percentage points**: "+0.3pp". Never "+0.3%" for an absolute move on a rate. 2. Write relative effects as a percent **with the baseline named**: "+10% relative to a 3.0% baseline". 3. Put the resulting absolute delta in the test plan too, because that is the number the arithmetic consumes. 4. Apply the same discipline to results reporting, so a shipped "+5%" is never read back as five points by someone reading the summary a quarter later. ## Which reading should a spec normally use? Product stakeholders usually think in **relative** terms - "a 10% better checkout" is how value is discussed and how it maps onto revenue forecasts. Statistics consumes **absolute** terms. Both belong in the plan: the relative number carries the business intent, the absolute number carries the arithmetic, and the baseline is what links them. The failure mode is a document that contains only one of the three. ## What a strong answer sounds like Name the ambiguity, give both endpoints (8% versus 3.15%), state that sizing consumes the absolute difference, quantify the gap as orders of magnitude rather than hand-waving "a lot more", and finish with the convention that prevents a repeat. Interviewers ask this because it is the single most common way a real experiment plan goes wrong in a way nobody notices until the readout.

  • How do you convert a relative MDE into the absolute one the arithmetic needs?
    Multiply by the baseline rate: absolute delta = relative lift x baseline. A 10% relative target on a 3% baseline is 0.3 percentage points. This is why a relative MDE quoted without its baseline is not a portable number - the same percentage means a very different absolute shift on a 3% rate and on a 40% rate.
  • The baseline rate drifts from 3.0% down to 2.5% before the test starts. What happens to a plan written as a 10% relative MDE?
    The absolute effect being targeted drops from 0.30pp to 0.25pp while the per-user variance barely moves, so the required sample rises by roughly 40%. A plan pinned to a relative number silently gets harder when the baseline falls; re-run the arithmetic against the current baseline rather than trusting the original traffic estimate.
  • Which convention should the written spec use so this cannot recur?
    State absolute effects in percentage points, state relative effects as a percent of an explicitly named baseline, and record both plus the baseline in the test plan. Reporting follows the same rule, so a shipped result is never re-read in the wrong units months later.

saying these in an interview costs you the question

  • Uses percent and percentage points interchangeably on rates
  • Assumes 'a 5% lift' obviously means relative without asking
  • Quotes a relative MDE with no baseline attached
  • Thinks the two readings need similar sample sizes
  • Sizes on the relative number without converting to absolute

context