Where does the 16 come from in the per-arm sample rule n = 16 * sigma^2 / delta^2?
answer
- not a magic constant
- two arms, two z-values
- 1.96 and 0.84
- their sum squared, then doubled
- 15.68 rounded up
basics
~20 sThe 16 is 2 x (1.96 + 0.84)^2 rounded up. 1.96 is the two-sided critical value at 5% significance, 0.84 corresponds to 80% power, and the factor 2 covers two equally sized arms each contributing sampling error.
solid answer
~40 sThe general per-arm formula is `n = 2 * (z_alpha/2 + z_beta)^2 * sigma^2 / delta^2`. Plug in the two conventions almost everyone uses: `z = 1.96` for two-sided 5% significance and `z = 0.84` for 80% power. Their sum is 2.80, squared is 7.84, doubled for the two arms is 15.7, which rounds to **16**. So `n = 16 * sigma^2 / delta^2` is not a magic number, it is those two defaults baked in. Applied to a 3% checkout conversion and a 10% relative lift, `delta = 0.003` and `sigma^2 = p(1-p)` is about 0.0305 at the midpoint, giving roughly **54,000 visitors per arm**, about 110,000 in total. The constant only holds for those defaults: at 90% power the multiplier becomes about 21, and a one-sided test lowers it.
go deeper
Be ready to use the rule correctly on a conversion metric: variance is p times one minus p, and the answer is per arm rather than the total.
An interviewer expects the derivation on demand: 1.96 for two-sided 5%, 0.84 for 80% power, doubled for two arms, giving 15.68 rounded to 16.
Show you use it as a sanity check on planning-tool output and that you notice when its assumptions break, such as an unbalanced split or tiny expected conversion counts.
Be able to argue which defaults your organisation should standardise on, since moving from 80% to 90% power raises every experiment's traffic bill by about 30%.
## The rule and its derivation The back-of-envelope every experimenter should be able to write from memory is `n_per_arm = 16 * sigma^2 / delta^2` where `sigma^2` is the per-user variance of the metric and `delta` is the absolute effect you want to detect. The 16 is not arbitrary. The general form is `n_per_arm = 2 * (z_alpha/2 + z_beta)^2 * sigma^2 / delta^2` and it decomposes cleanly: - `z_alpha/2 = 1.96` - the standard normal critical value for a **two-sided test at 5% significance**. This is how far the observed difference must sit from zero, in standard errors, before it is called significant. - `z_beta = 0.84` - the standard normal value corresponding to **80% power**. This is the extra distance the true effect must sit beyond the critical value so that four runs in five land past it. - The **factor 2** - the comparison is between two independent arm means. The variance of a difference of two independent means is `sigma^2/n + sigma^2/n = 2 sigma^2/n`, and that 2 propagates straight into the numerator. Arithmetic: `1.96 + 0.84 = 2.80`; `2.80^2 = 7.84`; `2 x 7.84 = 15.68`, rounded up to **16**. The rounding is generous by about 2%, which is the right direction for a planning rule. ## What changes the constant The 16 is only correct for the specific defaults it encodes. Common variants: | Significance (two-sided) | Power | Constant | |---|---|---| | 5% | 80% | ~16 | | 5% | 90% | ~21 | | 1% | 80% | ~24 | | 1% | 90% | ~30 | For 90% power, `z_beta = 1.28`, so `2 x (1.96 + 1.28)^2 = 2 x 10.5 = 21`. Anyone quoting 16 for a 90%-power plan will under-size by about 30%. A one-sided test at 5% replaces 1.96 with 1.64 and lowers the constant to about 12, which is why the sidedness must be written down rather than assumed. ## Using it on a conversion rate For a proportion, the per-user variance is `sigma^2 = p * (1 - p)`. Take the leaf's worked case: a **3% baseline checkout conversion** and a target of a **10% relative lift**. 1. Convert the relative target to absolute: `delta = 0.10 * 0.03 = 0.003`. 2. Variance term at the midpoint between 3.0% and 3.3%: `p_bar = 0.0315`, so `p_bar (1 - p_bar) = 0.0315 x 0.9685 = 0.0305`. 3. Apply the rule: `n = 16 x 0.0305 / 0.003^2 = 0.488 / 0.000009 ~= 54,000` per arm. 4. Two arms: roughly **108,000 visitors** must reach the checkout surface. Notice how brutal the `delta^2` denominator is. Tighten the ambition to a 5% relative lift and `delta` halves to 0.0015, so the requirement quadruples to about 216,000 per arm. Loosen it to a 20% relative lift and it falls to about 13,500 per arm. ## Why memorise it Three reasons interviewers like this question: - **It makes you calculator-free.** You can sanity-check any planning tool's output in your head to within a few percent, which catches the classic input mistakes: relative entered as absolute, total traffic entered where per-arm was meant, variance taken from a different metric. - **It exposes the assumptions.** Someone who knows the 16 is `2 x (1.96 + 0.84)^2` automatically knows the rule is tied to two arms, an equal split, a two-sided test, 5% significance and 80% power. Someone who memorised only the digit "16" will apply it to a four-arm test or a 90%-power plan without blinking. - **It shows the normal approximation.** The formula is a large-sample result. With very small expected counts - a handful of conversions per arm - the normal approximation degrades, and the planning number should be treated as indicative rather than exact. ## Caveats worth stating out loud - **Per arm, not total.** The rule gives the sample for *each* arm; double it for a two-arm total. - **Equal split assumed.** An unbalanced split needs more total traffic for the same sensitivity, because the smaller arm dominates the standard error. - **Variance must match the metric.** For proportions, `p(1-p)` is trivially available; for revenue-like metrics, the variance has to come from historical data, and outliers make it much larger than intuition suggests. - **The unit of analysis must match randomisation.** If you randomise visitors, `sigma^2` must be the variance of the per-visitor value, not the per-session one.
- What does the constant become at 90% power, and why?About 21. The value for 90% power is 1.28 rather than 0.84, so `2 x (1.96 + 1.28)^2 = 2 x 10.5 ~= 21`. Using 16 on a plan that promises 90% power under-sizes the test by roughly 30%, which is a silent and common planning error.
- Apply the rule to a 3% baseline and a 10% relative lift target.The absolute effect is `0.10 x 0.03 = 0.003`. With a variance term of about `0.0315 x 0.9685 = 0.0305`, the rule gives `16 x 0.0305 / 0.000009 ~= 54,000` per arm, so about 108,000 visitors must reach the surface in total.
- Where does the factor of 2 in the formula come from?From comparing two independent arm means. The variance of their difference is `sigma^2/n + sigma^2/n = 2 sigma^2/n`, so the 2 rides straight into the numerator. It is also the reason the rule gives a per-arm number rather than a total.
saying these in an interview costs you the question
- Treats 16 as a universal constant for any alpha or power
- Applies the per-arm number as the total sample
- Cannot say which z-values the 16 encodes
- Uses it with a one-sided test unchanged
- Plugs in a relative lift where the absolute delta belongs