How is Cramer's V computed from a contingency table's chi-square statistic, and what does its scale mean?
answer
- normalise the chi-square statistic
- divide by n first
- then by the smaller dimension minus one
- square root at the end
- strength only, never a direction
basics
~10 sCramer's V equals sqrt(chi2 / (n x (min(r, c) - 1))), with n the total count and r, c the table dimensions. It rescales the chi-square statistic onto 0 to 1, with no direction.
solid answer
~40 sCramer's V turns a chi-square statistic into a bounded effect size for two categorical variables: `V = sqrt(chi2 / (n * (min(r, c) - 1)))`. For a 2x3 device-by-plan table with 500 respondents and a chi-square statistic of 20, min(2, 3) - 1 = 1, so V = sqrt(20/500) = 0.20. The divisor is the maximum chi-square a table of that shape can reach, n x (min(r, c) - 1), which is what pins V at 1 when one variable perfectly determines the other and 0 when the rows are proportionally identical. On a 2x2 table the divisor is just n, and V equals the magnitude of the phi coefficient. V is unsigned: nominal categories have no ordering, so it reports strength only, never direction, and it never says which cells drive the association.
go deeper
Recall the formula and that V runs from 0 to 1 with no sign. Be able to plug in a chi-square statistic, a sample size and a table shape and get a number.
Explain why each divisor is there: n removes sample-size scaling, min(r, c) - 1 pins the ceiling at 1, and the square root undoes the squaring. Show that 2x2 reduces to phi.
Demonstrate judgment on interpretation: check n and table shape before believing a V, look at conditional row proportions to find where the association sits, and flag near-1 values as likely duplicate encodings.
Own how association is reported across the org: which coefficient is standard for categorical pairs, whether bias correction is applied by default, and how category-collapsing rules are fixed so numbers stay comparable between teams.
## The problem V solves The chi-square statistic on a contingency table measures how far the observed cell counts sit from what proportionally identical rows would give. It is useful but unbounded and it scales with the sample: double every count in a table and the chi-square statistic doubles, even though the pattern of association is unchanged. That makes it useless for answering how strong, and useless for comparing two tables. Cramer's V is the standardisation that fixes both problems. ## The formula For a table with r rows, c columns and n total observations: `V = sqrt( chi2 / (n * (min(r, c) - 1)) )` Three pieces matter. Dividing by n removes the sample-size scaling, so V stays put when you duplicate the dataset. Dividing by min(r, c) - 1 removes the table-shape scaling: the largest chi-square value a table of that shape can attain is n x (min(r, c) - 1), reached when each level of the smaller dimension pins down the other variable exactly. Taking the square root brings the quantity back from a squared scale to something comparable to a correlation magnitude. Worked example: a 2x3 table of device type against subscription plan on 500 respondents with a chi-square statistic of 20. Here min(2, 3) - 1 = 1, so V = sqrt(20 / (500 x 1)) = sqrt(0.04) = 0.20. If the same pattern held on 5,000 respondents the chi-square statistic would be 200 and V would still be 0.20, which is the whole point. ## Phi as the 2x2 special case On a 2x2 table min(r, c) - 1 = 1, so V = sqrt(chi2 / n), which is the phi coefficient. Phi has a second, equivalent definition: code each binary variable 0/1 and compute Pearson's correlation between the two columns. That version carries a sign, which is meaningful for binary flags where the coding has a natural direction, for example email-opened against purchased: a positive phi says opening and purchasing co-occur more than proportionality would give. Cramer's V is the absolute value, generalised to tables bigger than 2x2 where a sign would be meaningless because there is no ordering of nominal levels to be positive or negative about. ## Reading the scale V = 0 means every row has the same conditional distribution as every other, that is, no association in this sample. V = 1 means the smaller dimension is perfectly determined: knowing the row tells you the column with certainty (or vice versa, whichever dimension is smaller). In between, V behaves like a correlation magnitude but with no direction, and its interpretation is not shape-free: conventional small/medium/large cutoffs are usually quoted separately for each table shape, so 0.2 on a 2x2 is not automatically the same evidence as 0.2 on a 6x6. ## What V does not tell you It does not say where the association lives. A 4x5 table with V = 0.3 might be driven by one unusual row, with the other three behaving like the marginal; only inspecting the conditional row proportions reveals that. It does not tell you direction, because nominal levels have no order. It says nothing about causation: a strong association between country and currency is definitional, not causal. And it is not invariant to how you build the table: merging two rare categories into an other bucket changes both the chi-square statistic and min(r, c), so it changes V, which means the category scheme is an analysis choice you must disclose. ## Small-sample inflation V is biased upward. Even with two completely unrelated columns the observed counts wobble, producing a positive chi-square statistic and therefore a positive V. The inflation grows with table size and shrinks with n, so a large sparse table on a small sample can report a respectable-looking V from noise alone. Sanity-check any V against the table's shape and sample size before treating it as a finding, and prefer a bias-corrected variant when comparing tables of different shapes.
- Why is the divisor n x (min(r, c) - 1) rather than simply n?Because sqrt(chi2/n) only tops out at 1 on a 2x2 table. On bigger tables the attainable maximum chi-square is n x (min(r, c) - 1), so dividing by n alone would let the statistic exceed 1 and would make tables of different shapes incomparable. The extra factor is exactly what normalises the ceiling to 1 for any shape.
- How does Cramer's V relate to the phi coefficient?On a 2x2 table they agree in magnitude: min(r, c) - 1 = 1 makes V = sqrt(chi2/n), which is phi. Phi additionally has a signed form, being Pearson's correlation between the two variables coded 0/1, so it can say whether the two flags co-occur or exclude each other. V drops the sign and extends the idea to any table shape.
- Can Cramer's V be compared directly across tables of different shapes?Only with care. Both the upward small-sample bias and the conventional interpretation cutoffs depend on the table's dimensions, so a 0.25 on a 2x2 with 10,000 rows is far stronger evidence than a 0.25 on a 6x6 with 200 rows. Compare bias-corrected values, and always report n and the table shape alongside V.
Chi-square is distance in raw miles; V converts it to a fraction of the longest trip that table could possibly make.
saying these in an interview costs you the question
- Reports Cramer's V as if it had a direction or sign
- Forgets the square root and quotes chi2 over n
- Divides by min(r, c) instead of min(r, c) minus one
- Treats a high V as evidence of causation
- Assumes V is unaffected by sample size in small samples