skip to content

In a two-proportion z-test comparing 42/300 with 61/300, when is the standard error pooled?

level: seniorimportance: should knowfreq 42%

answer

  1. the test and the interval answer different questions
  2. the null says the two rates are equal
  3. use that assumption while computing the statistic
  4. an interval must allow non-zero differences
  5. combined rate is 103 out of 600

basics

~20 s

Pool for the hypothesis test, not for the confidence interval. The test assumes the null that both rates are equal, so one combined proportion estimates their shared variance; an interval must allow them to differ.

solid answer

~50 s

The test statistic is built under the null hypothesis, and that null says the two response rates are the same. If they are the same, the best estimate of that common rate uses every observation: `p_pool = (42 + 61) / 600 = 0.172`, giving `SE_0 = sqrt(p_pool * (1 - p_pool) * (1/300 + 1/300)) = 0.0308`. With a difference of `0.203 - 0.140 = 0.063`, that is `z = 2.06` and a two-sided p-value just under 0.05. A confidence interval for the difference is a different question: it must cover values other than zero, so assuming equality would be circular, and you use the unpooled `sqrt(p1*(1-p1)/n1 + p2*(1-p2)/n2)`, here about 0.0307. In this balanced example the two standard errors barely differ, but with unequal group sizes and rates far apart they can diverge enough that the interval and the test disagree at the boundary.

go deeper

for a junior

Be ready to compute the two sample proportions and state that the test compares their difference against its standard error using a normal reference distribution.

for a middle

Explain why the null hypothesis of equal rates licenses a single combined estimate, and be able to write both the pooled and unpooled standard error formulas.

for a senior

Show judgment when the two disagree: recognise a borderline result, report the difference with its interval rather than a bare p-value, and know which design features drive the two standard errors apart.

for a principal

Own the reporting standard. Decide what your organisation publishes for a comparison of rates, how borderline results are described, and how to stop teams from selecting whichever formula crosses a threshold.

## The setup Two survey arms each contacted 300 people. Arm A returned 42 responses, arm B returned 61. ``` p1 = 42/300 = 0.140 p2 = 61/300 = 0.2033 difference = 0.0633 ``` The question is whether that gap is bigger than sampling noise, and the two-proportion z-test answers it with the usual skeleton: difference divided by its standard error, compared against the standard normal. ## Why the test pools A hypothesis test computes "how surprising is this data **if the null were true**". The null here is `p1 = p2`: one shared response rate produced both arms. Under that assumption, splitting the data into two separate variance estimates throws information away — every one of the 600 contacts is a draw from the *same* Bernoulli process. So you estimate the shared rate from everything: ``` p_pool = (42 + 61) / (300 + 300) = 103/600 = 0.17167 SE_0 = sqrt( p_pool * (1 - p_pool) * (1/n1 + 1/n2) ) = sqrt( 0.17167 * 0.82833 * (1/300 + 1/300) ) = sqrt( 0.142193 * 0.0066667 ) = 0.03079 z = 0.06333 / 0.03079 = 2.06 ``` The two-sided p-value is about 0.04, so at a 5% level this is a significant difference — though only just, which is worth saying out loud rather than reporting "significant" flatly. Pooling matters because it makes the test's null distribution exactly the one the p-value claims. If you compute the statistic with an unpooled standard error and then compare it to the standard normal, the reference distribution no longer matches how the statistic behaves under the null, and the error rate drifts from its nominal level — usually slightly, but the pooled version is the one with the clean justification. ## Why the interval does not pool A confidence interval is not computed under the null. Its whole purpose is to display the range of differences consistent with the data, including differences far from zero. Estimating the standard error under an assumption that the difference is zero would be circular: you would be using "they are equal" to describe how uncertain you are about how unequal they are. So each arm contributes its own variance: ``` SE_unpooled = sqrt( p1*(1-p1)/n1 + p2*(1-p2)/n2 ) = sqrt( 0.140*0.860/300 + 0.2033*0.7967/300 ) = sqrt( 0.00040133 + 0.00053996 ) = 0.03068 CI(95%) = 0.0633 +/- 1.96 * 0.03068 = (0.0032, 0.1235) ``` The interval excludes zero, consistent with the test rejecting — as you would expect when both are computed correctly on the same data. ## When do the two standard errors actually differ? In this example, hardly at all: 0.0308 versus 0.0307. That is typical of a **balanced** design with proportions that are not far apart. The pooled and unpooled standard errors separate when: - **Group sizes are unequal.** Pooling weights the combined rate by sample size, so a large arm drags the assumed common rate toward its own value. - **The proportions are far apart.** Pooling assumes one rate; the further the truth is from that, the worse the assumption describes either arm. - **Rates are near 0 or 1**, where `p * (1 - p)` changes fastest with `p`. When they do differ, you can get the awkward boundary case where a 95% interval excludes zero while the pooled test returns `p > 0.05`, or the reverse. That is not a contradiction to be argued away — it is the honest consequence of the two procedures answering different questions. Report both, say which one your decision rests on, and do not go shopping for whichever crosses the threshold. ## Conditions for the normal approximation The z-test approximates a discrete binomial by a continuous normal, which needs enough events in each cell. The usual rough guideline is at least about 10 successes and 10 failures per group. Here 42 and 258, and 61 and 239, are comfortable. With single-digit counts in any cell the approximation degrades and an exact method is more appropriate. ## Practical checklist 1. Confirm each arm is a count of independent successes out of a fixed number of trials. 2. Check the counts are large enough for a normal approximation in every cell. 3. For the **test**: pool, `z = (p1 - p2) / sqrt(p_pool*(1-p_pool)*(1/n1 + 1/n2))`. 4. For the **interval**: do not pool, `(p1 - p2) +/- 1.96 * sqrt(p1*(1-p1)/n1 + p2*(1-p2)/n2)`. 5. Report the difference and its interval, not only the p-value — a 6.3 percentage-point gap with an interval running from 0.3 to 12.4 points is a much more useful sentence than "p = 0.04". ## What interviewers are checking That you understand a null hypothesis is not just a threshold to compare against but an assumption you are allowed to *use* while computing the test statistic — and that the same assumption becomes illegitimate the moment you switch from testing to estimating. Candidates who have only ever called a routine rarely notice that these two standard errors are different quantities at all.

  • What is the pooled proportion for 42/300 versus 61/300, and what z does it give?
    The pooled rate is 103/600 = 0.1717. The null standard error is sqrt(0.1717 * 0.8283 * (1/300 + 1/300)) = 0.0308, and the observed difference is 0.2033 - 0.140 = 0.0633, so z = 2.06 with a two-sided p-value of about 0.04. It clears a 5% threshold, but only just, which is worth stating alongside the verdict.
  • Can the pooled test and the unpooled confidence interval disagree, and what do you do then?
    Yes, near the boundary — an interval can exclude zero while the pooled test returns p just above 0.05, or the reverse. They are computed under different assumptions, so this is expected rather than an error. Report both, state which one the decision rests on before you looked, and treat the result as borderline evidence rather than picking whichever crosses the line.
  • When do the pooled and unpooled standard errors differ enough to matter?
    When the group sizes are unequal, when the two proportions are far apart, or when either rate sits near 0 or 1 where p*(1-p) changes fastest. With balanced arms and similar rates, as in 42/300 versus 61/300, the two standard errors agree to three decimals and the choice changes nothing numerically — only the reasoning.

saying these in an interview costs you the question

  • Uses the pooled standard error to build the confidence interval
  • Says pooling versus not pooling is an arbitrary style choice
  • Cannot state what the null hypothesis assumes about the two rates
  • Reports significance without the size of the difference
  • Ignores whether each cell has enough events for a normal approximation

context