skip to content

Where does the 16 come from in the per-arm sample rule n = 16 * sigma^2 / delta^2?

level: middleimportance: nice to knowfreq 33%

answer

  1. not a magic constant
  2. two arms, two z-values
  3. 1.96 and 0.84
  4. their sum squared, then doubled
  5. 15.68 rounded up

basics

~20 s

The 16 is 2 x (1.96 + 0.84)^2 rounded up. 1.96 is the two-sided critical value at 5% significance, 0.84 corresponds to 80% power, and the factor 2 covers two equally sized arms each contributing sampling error.

solid answer

~40 s

The general per-arm formula is `n = 2 * (z_alpha/2 + z_beta)^2 * sigma^2 / delta^2`. Plug in the two conventions almost everyone uses: `z = 1.96` for two-sided 5% significance and `z = 0.84` for 80% power. Their sum is 2.80, squared is 7.84, doubled for the two arms is 15.7, which rounds to **16**. So `n = 16 * sigma^2 / delta^2` is not a magic number, it is those two defaults baked in. Applied to a 3% checkout conversion and a 10% relative lift, `delta = 0.003` and `sigma^2 = p(1-p)` is about 0.0305 at the midpoint, giving roughly **54,000 visitors per arm**, about 110,000 in total. The constant only holds for those defaults: at 90% power the multiplier becomes about 21, and a one-sided test lowers it.

go deeper

for a junior

Be ready to use the rule correctly on a conversion metric: variance is p times one minus p, and the answer is per arm rather than the total.

for a middle

An interviewer expects the derivation on demand: 1.96 for two-sided 5%, 0.84 for 80% power, doubled for two arms, giving 15.68 rounded to 16.

for a senior

Show you use it as a sanity check on planning-tool output and that you notice when its assumptions break, such as an unbalanced split or tiny expected conversion counts.

for a principal

Be able to argue which defaults your organisation should standardise on, since moving from 80% to 90% power raises every experiment's traffic bill by about 30%.

## The rule and its derivation The back-of-envelope every experimenter should be able to write from memory is `n_per_arm = 16 * sigma^2 / delta^2` where `sigma^2` is the per-user variance of the metric and `delta` is the absolute effect you want to detect. The 16 is not arbitrary. The general form is `n_per_arm = 2 * (z_alpha/2 + z_beta)^2 * sigma^2 / delta^2` and it decomposes cleanly: - `z_alpha/2 = 1.96` - the standard normal critical value for a **two-sided test at 5% significance**. This is how far the observed difference must sit from zero, in standard errors, before it is called significant. - `z_beta = 0.84` - the standard normal value corresponding to **80% power**. This is the extra distance the true effect must sit beyond the critical value so that four runs in five land past it. - The **factor 2** - the comparison is between two independent arm means. The variance of a difference of two independent means is `sigma^2/n + sigma^2/n = 2 sigma^2/n`, and that 2 propagates straight into the numerator. Arithmetic: `1.96 + 0.84 = 2.80`; `2.80^2 = 7.84`; `2 x 7.84 = 15.68`, rounded up to **16**. The rounding is generous by about 2%, which is the right direction for a planning rule. ## What changes the constant The 16 is only correct for the specific defaults it encodes. Common variants: | Significance (two-sided) | Power | Constant | |---|---|---| | 5% | 80% | ~16 | | 5% | 90% | ~21 | | 1% | 80% | ~24 | | 1% | 90% | ~30 | For 90% power, `z_beta = 1.28`, so `2 x (1.96 + 1.28)^2 = 2 x 10.5 = 21`. Anyone quoting 16 for a 90%-power plan will under-size by about 30%. A one-sided test at 5% replaces 1.96 with 1.64 and lowers the constant to about 12, which is why the sidedness must be written down rather than assumed. ## Using it on a conversion rate For a proportion, the per-user variance is `sigma^2 = p * (1 - p)`. Take the leaf's worked case: a **3% baseline checkout conversion** and a target of a **10% relative lift**. 1. Convert the relative target to absolute: `delta = 0.10 * 0.03 = 0.003`. 2. Variance term at the midpoint between 3.0% and 3.3%: `p_bar = 0.0315`, so `p_bar (1 - p_bar) = 0.0315 x 0.9685 = 0.0305`. 3. Apply the rule: `n = 16 x 0.0305 / 0.003^2 = 0.488 / 0.000009 ~= 54,000` per arm. 4. Two arms: roughly **108,000 visitors** must reach the checkout surface. Notice how brutal the `delta^2` denominator is. Tighten the ambition to a 5% relative lift and `delta` halves to 0.0015, so the requirement quadruples to about 216,000 per arm. Loosen it to a 20% relative lift and it falls to about 13,500 per arm. ## Why memorise it Three reasons interviewers like this question: - **It makes you calculator-free.** You can sanity-check any planning tool's output in your head to within a few percent, which catches the classic input mistakes: relative entered as absolute, total traffic entered where per-arm was meant, variance taken from a different metric. - **It exposes the assumptions.** Someone who knows the 16 is `2 x (1.96 + 0.84)^2` automatically knows the rule is tied to two arms, an equal split, a two-sided test, 5% significance and 80% power. Someone who memorised only the digit "16" will apply it to a four-arm test or a 90%-power plan without blinking. - **It shows the normal approximation.** The formula is a large-sample result. With very small expected counts - a handful of conversions per arm - the normal approximation degrades, and the planning number should be treated as indicative rather than exact. ## Caveats worth stating out loud - **Per arm, not total.** The rule gives the sample for *each* arm; double it for a two-arm total. - **Equal split assumed.** An unbalanced split needs more total traffic for the same sensitivity, because the smaller arm dominates the standard error. - **Variance must match the metric.** For proportions, `p(1-p)` is trivially available; for revenue-like metrics, the variance has to come from historical data, and outliers make it much larger than intuition suggests. - **The unit of analysis must match randomisation.** If you randomise visitors, `sigma^2` must be the variance of the per-visitor value, not the per-session one.

  • What does the constant become at 90% power, and why?
    About 21. The value for 90% power is 1.28 rather than 0.84, so `2 x (1.96 + 1.28)^2 = 2 x 10.5 ~= 21`. Using 16 on a plan that promises 90% power under-sizes the test by roughly 30%, which is a silent and common planning error.
  • Apply the rule to a 3% baseline and a 10% relative lift target.
    The absolute effect is `0.10 x 0.03 = 0.003`. With a variance term of about `0.0315 x 0.9685 = 0.0305`, the rule gives `16 x 0.0305 / 0.000009 ~= 54,000` per arm, so about 108,000 visitors must reach the surface in total.
  • Where does the factor of 2 in the formula come from?
    From comparing two independent arm means. The variance of their difference is `sigma^2/n + sigma^2/n = 2 sigma^2/n`, so the 2 rides straight into the numerator. It is also the reason the rule gives a per-arm number rather than a total.

saying these in an interview costs you the question

  • Treats 16 as a universal constant for any alpha or power
  • Applies the per-arm number as the total sample
  • Cannot say which z-values the 16 encodes
  • Uses it with a one-sided test unchanged
  • Plugs in a relative lift where the absolute delta belongs

context