In a chi-square test of independence on a 3x4 contingency table, how do you get expected counts and degrees of freedom?
answer
- independence means multiply the marginals
- row total times column total
- divide by the grand total
- margins are fixed, not free
- two times three for a 3x4 table
basics
~20 sEach expected count is that cell's row total times its column total, divided by the grand total. Degrees of freedom are (rows minus one) times (columns minus one), so a 3x4 table has (3-1)(4-1) = 6.
solid answer
~50 sThe null hypothesis is that the two categorical variables are independent — say job role (3 levels) and preferred work mode (4 levels) for N surveyed people. Under independence, the probability of landing in a cell is the product of its marginal probabilities, so the expected count is `E_ij = (row_i total * col_j total) / N`. That formula falls straight out of `N * (row_i/N) * (col_j/N)`. You then form the usual `sum((O - E)^2 / E)` over all 12 cells. Degrees of freedom are `(r-1)(c-1) = 2 * 3 = 6`: the row and column totals are fixed by the data, so once you know six cells the remaining six are determined. A large statistic says role and work-mode preference are associated in this sample — it does not say which way the association runs or that one causes the other.
go deeper
Memorise the two formulas and be able to apply them to a small table: expected equals row total times column total over grand total, and df equals (rows minus one) times (columns minus one).
Derive both formulas on demand. Expect to explain that independence means multiplying marginal probabilities, and to justify (r-1)(c-1) by counting how many cells are still free once the margins are fixed.
Demonstrate that you check the design before the arithmetic: one tally per subject, independent observations, expected counts large enough. Then translate a rejection into a description of the table rather than a bare p-value.
Own the decision of whether a two-way significance test answers the actual question. With very large samples almost every table rejects, so the useful output is the size and shape of the association and what action it would justify.
## The question a test of independence asks Given a cross-tabulation of two categorical variables, is the distribution of one variable the same regardless of the level of the other? Take a survey of N respondents cross-tabbed as **job role** (engineer / analyst / manager — 3 rows) by **preferred work mode** (fully remote / hybrid-two-days / hybrid-four-days / fully on-site — 4 columns). That is a 3x4 contingency table with 12 cells. The null hypothesis is independence: knowing someone's role tells you nothing about their work-mode preference. ## Expected counts under independence Probability theory does the work. If two events are independent, P(A and B) = P(A) * P(B). Estimating the marginals from the data, the probability of row i is row_i_total / N and the probability of column j is col_j_total / N. So under the null the probability of cell (i, j) is their product, and the expected number of people there is E_ij = N * (row_i_total / N) * (col_j_total / N) = (row_i_total * col_j_total) / N Hence the familiar rule: **row total times column total over grand total**. Worked concretely, if 240 of 600 respondents are engineers and 150 of 600 prefer fully remote, the engineer/fully-remote cell expects 240 * 150 / 600 = 60 people. Expected counts are usually fractional and should be left that way — never round them before computing the statistic. By construction the expected counts reproduce both sets of margins exactly. That is the whole point: the test holds the marginal distributions fixed at their observed values and asks only whether the *interior* of the table is arranged as independence would predict. ## The statistic chi-square = sum over all cells of (O_ij - E_ij)^2 / E_ij summed over every one of the 12 cells. Each term is a squared deviation scaled by the noise you expect in a cell of that size. Larger totals mean larger absolute deviations are normal, and the division by E accounts for it. ## Where (r-1)(c-1) comes from Degrees of freedom count how many cell deviations are genuinely free once the constraints are imposed. There are r * c cells. The margins are not free — they were used to build the expected counts — so: - fixing the row totals costs r constraints, - fixing the column totals costs c constraints, - but one of those is redundant, since both sets must sum to the same N. That leaves r*c - (r + c - 1) = (r-1)(c-1). For the 3x4 table: 12 - (3 + 4 - 1) = 12 - 6 = **6**, matching (3-1)(4-1) = 6. A hands-on way to see it: in a 3x4 table with all margins known, you can fill any 2x3 block of cells freely, and every remaining cell is then forced by subtraction. Six free cells, six degrees of freedom. ## Assumptions and health checks - **Counts, one cell per subject.** Each respondent contributes exactly one tally. If people can pick two work modes, the totals exceed N and the test is invalid. - **Independent observations.** Surveying five people per team, or the same person twice, breaks the sampling assumption; the test has no idea and will happily return a small p-value. - **Adequate expected counts.** The chi-square reference distribution is asymptotic. Sparse expectations degrade it, and the usual remedy is to merge substantively similar categories — for example collapsing the two hybrid columns into one — or to use an exact method. ## Reading the outcome A significant result says only "the variables are associated in this sample". It does not name a direction, a magnitude, or a cause, and it is symmetric — the test cannot tell you whether role shapes preference or the reverse. To describe the association, go back to the table: compare row percentages across columns, and inspect standardised residuals `(O - E)/sqrt(E)` to see which cells over- and under-shoot. Interpret those descriptively, because you are now scanning twelve cells. Also note what independence testing and goodness of fit share and where they differ. Both use the same statistic and the same reference family; goodness of fit compares one variable's counts against *externally specified* probabilities with df = k - 1, while independence compares a two-way table against margin-derived expectations with df = (r-1)(c-1). Candidates who memorise a single df formula for "the chi-square test" get caught here. Finally, sample size cuts both ways. With N in the millions, an association far too weak to act on will still reject; with a few dozen respondents, a strong association can go undetected. Always report the table alongside the p-value.
- Why is one of the r + c margin constraints redundant when counting degrees of freedom?Because the row totals and the column totals must both add up to the same grand total N. Once you know all r row totals and c-1 column totals, the last column total is forced. So the r + c constraints impose only r + c - 1 independent restrictions, leaving r*c - (r + c - 1) = (r-1)(c-1) free deviations.
- The test rejects independence for job role against work-mode preference. What do you report next?The table itself. Show row percentages so the reader sees how preference shifts across roles, and flag the cells with the largest standardised residuals as the drivers. State plainly that the test is symmetric and observational: it establishes association within this sample, not direction and not causation.
- How does a chi-square test of homogeneity differ from a test of independence?Arithmetically they are identical — same expected counts, same statistic, same (r-1)(c-1) degrees of freedom. The difference is the sampling design and therefore the claim: independence draws one sample and cross-tabs two variables, while homogeneity draws fixed-size samples from several populations and asks whether one categorical distribution matches across them.
saying these in an interview costs you the question
- Uses categories minus one for a two-way table
- Rounds expected counts to whole numbers before summing
- Reads a significant result as evidence of causation
- Computes expected counts from the grand mean of the cells
- Applies the test when respondents can appear in several cells