How do you compute the standard error of the difference between two independent sample means?
answer
- uncertainty accumulates when you subtract
- variances add, standard errors do not
- square, add, then square-root
- each group contributes s squared over n
- pairing breaks the independence assumption
basics
~10 sVariances add, standard errors do not. The standard error of the difference between two independent sample means is sqrt(SE1^2 + SE2^2), which expands to sqrt(s1^2/n1 + s2^2/n2).
solid answer
~40 sSquare, add, square-root. For two **independent** samples, `Var(mean1 - mean2) = Var(mean1) + Var(mean2)`, because independent variances add even when the statistics are subtracted. Taking the square root gives `SE_diff = sqrt(SE1^2 + SE2^2) = sqrt(s1^2/n1 + s2^2/n2)`. Say two schools report mean test scores of 512 and 498 with standard errors 4 and 3; the difference of 14 points carries a standard error of `sqrt(16 + 9) = 5`, not `4 + 3 = 7`. Adding the standard errors directly is the classic error and always overstates the uncertainty. Two consequences follow. First, the noisier or smaller group dominates: the group with the larger `s^2/n` sets the floor, so topping up the already-large sample barely helps. Second, the formula needs independence - for paired measurements on the same subjects it does not apply.
go deeper
Know that comparing two group averages has its own uncertainty, and that you combine the two standard errors by squaring, adding and taking the square root rather than by adding them directly.
Explain why: independent variances add even under subtraction, giving sqrt(s1 squared over n1 plus s2 squared over n2). Work a numeric example and show that combining 4 and 3 yields 5, not 7.
Demonstrate practical judgment - spot paired data being analysed as independent, recognise that the smaller arm dominates the combined standard error, and steer extra sampling effort to where it actually buys precision.
Own the experiment design implications: allocation between arms, when pairing or blocking is worth its operational cost, and how to explain to stakeholders that a big control group cannot rescue an underpowered treatment arm.
## What is being estimated Comparisons are almost never about one group; they are about a **difference**. The estimate is `mean1 - mean2` and, like any statistic, it has its own sampling distribution and therefore its own standard error, describing how much that difference would move if both samples were redrawn. ## The rule: variances add For independent random variables `A` and `B`: ``` Var(A - B) = Var(A) + Var(B) ``` The minus sign does **not** become a minus on the right. Subtracting a noisy quantity injects noise just as adding one does; uncertainty accumulates in either direction. Applying this to two sample means: ``` Var(mean1 - mean2) = s1^2/n1 + s2^2/n2 SE_diff = sqrt(s1^2/n1 + s2^2/n2) = sqrt(SE1^2 + SE2^2) ``` The recipe is: **square each standard error, add, take the square root.** Standard errors combine in quadrature, like the sides of a right triangle, never by simple addition. ## A worked comparison Two independent schools sit the same test. | school | mean | s | n | SE = s / sqrt(n) | |---|---|---|---|---| | A | 512 | 40 | 100 | 4 | | B | 498 | 30 | 100 | 3 | - Difference in means: `512 - 498 = 14` points. - Standard error of that difference: `sqrt(4^2 + 3^2) = sqrt(25) = 5` points. So the observed 14-point gap is estimated with a wobble of about 5 points per redraw. Note the arithmetic: 5 is smaller than the 7 you would get by adding the standard errors, and larger than either one alone. Both bounds are worth remembering as a sanity check - the combined standard error always sits between `max(SE1, SE2)` and `SE1 + SE2`. ## Why adding standard errors is wrong Adding assumes the two errors always push in the same direction, which is exactly what independence rules out. Sometimes school A's sample runs high and B's runs high too, and the errors partly cancel in the difference. Because cancellation is possible, the honest combined uncertainty is less than the sum. Adding standard errors inflates uncertainty, which sounds conservative but leads to declaring real differences unmeasurable. ## The smaller group dominates Unequal sample sizes have a consequence people repeatedly miss. Suppose the control group has 1,000,000 observations with `s = 40` (`SE = 0.04`) and the treatment group has 400 with `s = 40` (`SE = 2`): ``` SE_diff = sqrt(0.04^2 + 2^2) = sqrt(0.0016 + 4) = 2.0004 ``` The giant group contributes essentially nothing. Precision of the comparison is set by the smaller arm, so pouring more data into the already-large group is wasted effort; the lever is the small arm. This is why balanced allocation is roughly optimal when the two groups have similar variance. ## Unequal variances Nothing above required `s1 = s2`. The formula `sqrt(s1^2/n1 + s2^2/n2)` handles unequal spreads directly and is the safe default. An alternative, the **pooled** standard error, first combines the two samples into a single variance estimate and is only valid when the population variances really are equal; when they are not and the sample sizes differ, pooling can badly misstate the standard error. Using the separate-variance version costs almost nothing when variances happen to be equal, so it is the sensible habit. ## When independence fails: paired data If the same students are measured before and after a programme, or the same users see both designs, the two means are **not** independent. The general identity is ``` Var(A - B) = Var(A) + Var(B) - 2 * Cov(A, B) ``` With positive correlation - a strong student is strong on both occasions - the covariance term is positive and subtracts, so the true standard error of the difference is **smaller** than the independent formula suggests. The clean way to handle it is to stop treating them as two samples: compute each subject's own difference, then take the standard error of those `n` differences as an ordinary mean, `s_d / sqrt(n)`. That automatically captures the covariance. Applying the independent formula to paired data throws away the pairing and leaves real effects looking noisier than they are - a genuine loss of statistical efficiency. ## Sanity checks to carry into an interview - The combined standard error is never smaller than the larger of the two inputs and never larger than their sum. - Halving the standard error of a difference still needs roughly four times the data, in **both** arms. - If one arm is far larger, quote its contribution and show that it is negligible; interviewers like seeing the term dropped for the right reason rather than by feel.
- Why can't you simply add the two standard errors together?Because that assumes the two sampling errors always move in the same direction. Under independence they sometimes offset, so part of the noise cancels in the difference and the honest combined standard error is smaller than the sum. Adding gives 7 where the correct combination of 4 and 3 gives 5, systematically overstating uncertainty and hiding real differences.
- The two measurements are taken on the same subjects before and after - what changes?Independence fails, so `Var(A - B) = Var(A) + Var(B) - 2 * Cov(A, B)`. With positively correlated measurements the covariance term shrinks the true standard error below what the independent formula reports. The clean fix is to compute each subject's own difference and take the standard error of those differences, `s_d / sqrt(n)`, which captures the pairing automatically.
- One group has 400 observations and the other has a million - where should extra data go?Into the small group. The combined standard error is `sqrt(s1^2/n1 + s2^2/n2)`, and the term from the million-row group is negligible, so it already contributes nothing measurable. Precision of the comparison is pinned by the smaller arm, and adding to the large arm changes essentially nothing at any cost.
saying these in an interview costs you the question
- Adds the two standard errors instead of combining in quadrature
- Subtracts variances because the means are subtracted
- Applies the independent formula to before-and-after measurements on the same subjects
- Assumes equal variances and pools without checking
- Thinks the larger sample determines the precision of the comparison