skip to content

Why does a 2x2 chi-square test give the same p-value as a two-proportion z-test?

level: seniorimportance: should knowfreq 36%

answer

  1. one degree of freedom is a squared normal
  2. all four cell deviations share a magnitude
  3. squaring discards the sign
  4. pooled standard error, no correction
  5. z keeps the interval, chi-square scales up

basics

~20 s

On a 2x2 table the Pearson chi-square statistic is algebraically identical to the square of the pooled-variance two-proportion z statistic. A chi-square variable on one degree of freedom is a squared standard normal, so the two-sided tail areas coincide exactly.

solid answer

~50 s

They are the same test written two ways. For a 2x2 table the Pearson statistic simplifies to exactly `z^2`, where z is the two-proportion statistic computed with the pooled proportion in the standard error. And a chi-square distribution on 1 df is by definition the distribution of the square of a standard normal, so P(chi-square > z^2) equals P(|Z| > |z|) — the same two-sided p-value, to the last digit. The practical consequences follow from squaring: the chi-square form throws away the sign, so it cannot express a one-sided alternative and gives you no direction or confidence interval, while the z form hands you a signed difference in proportions and a CI for it. Two caveats: the equality assumes the uncorrected Pearson statistic and the pooled standard error, so a continuity correction or an unpooled standard error breaks it.

go deeper

for a junior

Know the headline fact: on a 2x2 table the chi-square statistic equals the z statistic squared, and both report the same two-sided p-value. You are not expected to derive it.

for a middle

Explain the one-degree-of-freedom fact that a chi-square variable on 1 df is a squared standard normal, and note that all four cell deviations in a 2x2 table share one magnitude.

for a senior

Turn the identity into a choice: use the proportion form when you need direction, a one-sided alternative or an interval, and the chi-square when the table grows beyond two by two. Name the conditions that break the equality.

for a principal

Own the reporting standard. Insist that count-table results reach stakeholders as a difference with an interval rather than a bare significance verdict, and set the team convention on continuity corrections and pooling so results stay comparable.

## The claim Take a 2x2 table: two groups, each observation a success or a failure. You can analyse it two ways. 1. **Two-proportion test:** compare the success proportions p1 and p2 with a statistic `z = (p1 - p2) / SE`, where the standard error is built from the **pooled** proportion p = (total successes)/(total observations): `SE = sqrt(p(1-p)(1/n1 + 1/n2))`. 2. **Chi-square test of independence:** compute expected counts from the margins and sum `(O - E)^2 / E` over the four cells, on `(2-1)(2-1) = 1` degree of freedom. These give the same p-value, because **chi-square = z^2** identically, cell by cell, for every 2x2 table. ## Why the distributions line up The distributional half is a definition. A chi-square random variable on k degrees of freedom is the sum of k squared independent standard normals; with k = 1 it is simply `Z^2` for `Z ~ N(0,1)`. So the event `Z^2 > c` is the same event as `|Z| > sqrt(c)`, and the upper-tail area of the chi-square at `z^2` equals the two-sided normal tail area at `z`. Nothing approximate happens at this step — given that the statistics are equal, the p-values must be too. ## Why the statistics are equal The algebraic half is a small piece of bookkeeping worth being able to sketch. Label the table cells a, b (group 1 successes and failures) and c, d (group 2). Because the expected counts are built from the margins, all four deviations `O - E` have the **same absolute value**; they alternate in sign so that every row and column deviation cancels. Call that common magnitude D. Working it through, `D = (ad - bc)/N`, and summing `D^2/E` over the four cells collapses to chi-square = N (ad - bc)^2 / [(a+b)(c+d)(a+c)(b+d)] Expanding the pooled-variance z statistic for the same table produces exactly the same expression once squared. The reason the two derivations meet is that both are driven by a single number — the one free deviation in the table — which is also why the test has exactly one degree of freedom. ## What this tells you to do in practice Because the relationship is squaring, everything that survives squaring is shared and everything that does not is lost: - **Direction disappears.** The chi-square statistic is the same whether group 1 or group 2 is ahead. If you need a one-sided alternative — "is the treatment arm's conversion *higher*?" — you must use the z form, or halve the chi-square p-value while separately confirming the direction from the table. - **Effect size disappears.** The z form starts from `p1 - p2`, a quantity in interpretable units, and extends naturally to a confidence interval for that difference. The chi-square gives a magnitude of surprise with no units. When a stakeholder asks "how much better?", the chi-square has no answer. - **The chi-square generalises; the z does not.** Beyond 2x2 — three groups, or an outcome with four levels — the chi-square keeps working with `(r-1)(c-1)` degrees of freedom, while the two-proportion statistic has no direct extension. ## The conditions for exact equality The identity holds for the **uncorrected** Pearson statistic and the **pooled** standard error. Two common variations break it: - **Yates' continuity correction** subtracts 0.5 from each absolute deviation before squaring. The corrected statistic no longer equals the uncorrected z squared; it is smaller, so the p-value is larger. If you compare a corrected chi-square against an uncorrected z you will see a discrepancy and it is not a bug. - **An unpooled standard error**, using p1 and p2 separately rather than the pooled p, gives a slightly different z. Pooling is the natural choice for a null of equal proportions, since under that null there is one common proportion to estimate; unpooled standard errors belong with confidence intervals for the difference, where no null is assumed. A likelihood-ratio statistic (sometimes called G-squared) is a third route to the same table. It is asymptotically equivalent but not algebraically identical, so it generally returns a slightly different number. ## How this shows up in interviews The question is usually a probe for whether you understand that named tests are not separate machinery. A strong answer states the identity, explains the 1-df squared-normal fact in one sentence, and then makes the practical point: prefer the z form on a 2x2 because it preserves sign and yields an interval, and reach for the chi-square when the table is bigger than 2x2. A weak answer treats the agreement as a coincidence, or claims the tests differ in power.

  • If they agree exactly, when would you still prefer the two-proportion form on a 2x2 table?
    Whenever direction or magnitude matters, which is nearly always. The z form gives a signed difference in proportions, supports a one-sided alternative, and extends to a confidence interval for the difference. The chi-square gives only a two-sided verdict with no units, so it answers whether but never how much.
  • Why does applying Yates' continuity correction break the exact equality?
    The correction shrinks each absolute deviation by 0.5 before squaring, so the corrected statistic is strictly smaller than the uncorrected one and no longer equals z squared. The result is a larger, more conservative p-value. The identity is stated for the plain Pearson statistic only.
  • Does the identity extend to a 2x3 table compared against three separate pairwise z-tests?
    No. Beyond 2x2 the table has more than one degree of freedom, so no single normal statistic captures it, and the chi-square becomes an omnibus test of any departure from independence. Running pairwise comparisons instead raises a multiplicity problem that the single omnibus test does not have.

It is the same distance reported two ways: the z statistic says how far and in which direction, the chi-square reports only the square of the distance. Squaring keeps the magnitude and forgets which way you walked.

saying these in an interview costs you the question

  • Calls the agreement a numerical coincidence
  • Claims one of the two tests has more power
  • Uses a chi-square to justify a one-sided directional claim
  • Thinks the identity holds beyond a 2x2 table
  • Compares a Yates-corrected statistic with an uncorrected z

context