skip to content

For the difference between two proportions, should the confidence interval use a pooled standard error?

level: seniorimportance: nice to knowfreq 32%

answer

  1. pooling encodes an assumption
  2. which procedure assumes the null is true
  3. the test and the interval need not match
  4. each group keeps its own rate in the interval

basics

~10 s

No. Pooling assumes the two proportions are equal, which is the null hypothesis a test assumes but an interval must not. Build the interval with each group's own observed proportion in the standard error.

solid answer

~50 s

The interval is `(p1_hat - p2_hat) ± z × sqrt( p1_hat(1 - p1_hat)/n1 + p2_hat(1 - p2_hat)/n2 )`, using each group's own proportion — the unpooled form. The pooled proportion `(x1 + x2) / (n1 + n2)` belongs to the two-proportion z test statistic, where it is legitimate precisely because the test is computed *assuming the null is true* and under that assumption both groups share one common rate. An interval makes no such assumption; it describes the difference the data actually show. Building the null into it is a category error, and it can make the test and the interval disagree near the decision boundary. The same small-count weakness as the single-proportion Wald interval applies here — with few events in either arm, prefer a score-based interval for the difference such as Newcombe's method.

go deeper

for a junior

Know that a difference between two proportions gets its own interval, and that it is built from each group's observed rate rather than from the two groups merged.

for a middle

Explain the two standard-error forms and state clearly which belongs to the test and which to the interval, with the reason in one sentence.

for a senior

Show you have handled the awkward case — a test and an interval disagreeing at the margin, or an arm with almost no events — and know which method you would switch to.

for a principal

Own the reporting convention: which summary the organisation leads with for rate comparisons, and how differences are communicated so nobody reads percentage points as a relative change.

## Two formulas that look almost the same For two groups with successes `x1` of `n1` and `x2` of `n2`, write `p1_hat = x1 / n1` and `p2_hat = x2 / n2`. There are two standard-error expressions in circulation. **Unpooled**, using each group's own estimate: ``` SE_unpooled = sqrt( p1_hat(1 - p1_hat)/n1 + p2_hat(1 - p2_hat)/n2 ) ``` **Pooled**, using the combined rate `p_hat = (x1 + x2) / (n1 + n2)`: ``` SE_pooled = sqrt( p_hat(1 - p_hat) × (1/n1 + 1/n2) ) ``` They are numerically close when the two proportions are close and the groups are similar in size, which is why the choice is easy to get wrong without consequence for a while — and then to get wrong with consequence. ## The rule - The **test statistic** for the two-proportion z test uses the **pooled** standard error. - The **confidence interval** for the difference uses the **unpooled** standard error. ## Why the test may pool A hypothesis test is computed under the assumption that the null is true. For a two-proportion test the null is `p1 = p2`: one common rate produced both groups' data. Under that assumption, the best estimate of that single shared rate uses all the data, so combining the successes and combining the trials gives a more efficient estimate than either group alone. Pooling is not a convenience here; it is the correct thing to do given what the test assumes while computing its reference distribution. ## Why the interval must not An interval is not computed under the null. Its job is to describe the range of differences compatible with the data actually observed — including differences far from zero. Pooling would build the assumption *there is no difference* into a quantity whose entire purpose is to characterise the difference. If the two rates genuinely differ, the pooled expression is estimating the variance of something that does not exist: a common rate shared by two groups that do not share one. There is a practical consequence beyond the philosophical one. Because the two standard errors are not equal, a pooled test and an unpooled interval can disagree near the boundary: the test can reject at 5% while the interval, built the standard way, still contains zero — or the reverse. This is a known and accepted quirk of the classical two-proportion procedures, not a bug in your arithmetic. What you must not do is report both and let the reader assume they agree; pick the summary you will lead with and be consistent. ## Interpreting the result The interval is on the **difference** scale — a difference in percentage points, not a relative change. An interval of `[0.01, 0.06]` on a baseline of 0.04 is a very different story from the same interval on a baseline of 0.40, and stakeholders routinely read one as the other. If the relative scale is what matters, build the interval for the ratio rather than converting the endpoints of a difference interval by dividing, which does not give a correct interval for the ratio. ## Where the recipe breaks down The unpooled formula inherits the weaknesses of the single-proportion Wald interval. With few events in either arm, or a rate close to 0 or 1, the normal approximation behind it is poor, coverage falls below nominal, and the endpoints can fall outside the logically possible range of `[-1, 1]`. Two practical responses: - Use a **score-based interval for the difference** — Newcombe's hybrid method builds the difference interval from each group's Wilson limits and behaves far better with small counts. - If one arm has zero events, do not report a difference interval as if it were reliable; note the zero-event arm explicitly and bound its rate separately. ## What a good answer sounds like State the rule in one sentence, justify it with *the test assumes the null, the interval does not*, and then add the operational note that the two can disagree at the margin. That combination — the rule, the reason, and the consequence — is what distinguishes a candidate who memorised a formula from one who understands what the formula is for.

  • Why is pooling the right choice inside the two-proportion z test statistic?
    Because the test statistic is computed assuming the null hypothesis of equal proportions is true. Under that assumption a single common rate generated both groups, so combining all successes over all trials estimates it more efficiently than either group alone. The interval makes no such assumption and therefore cannot borrow it.
  • Can a pooled test and an unpooled interval reach different conclusions on the same data?
    Yes, near the decision boundary the test can reject while the interval still contains zero, or the reverse, because the two standard errors differ. It is a known quirk of the classical procedures rather than an arithmetic mistake. Decide in advance which summary you lead with and report it consistently.
  • What would you do if one of the two arms has almost no events?
    Stop trusting the normal-approximation formula. With very few events the coverage degrades badly and endpoints can fall outside the possible range. Use a score-based interval for the difference, such as Newcombe's hybrid built from each group's Wilson limits, and state the sparse arm's own bound explicitly rather than hiding it inside a difference.

saying these in an interview costs you the question

  • Pools for the interval because it uses more data
  • Assumes the test and the interval must share one standard error
  • Trusts a Wald difference interval with only a handful of events
  • Reads a difference in percentage points as a relative change
  • Divides difference-interval endpoints to get a ratio interval

context