skip to content

Association & Correlation

How two columns move together, measured by covariance, Pearson's r and rank coefficients, and the traps that let one number claim far too much. Interviewers push hard on correlation versus causation.

on this pageshow

explore

questions

25

In a cross-tab of device type by plan, what is the difference between joint, marginal and conditional proportions?

level: juniorimportance: must knowfreq 62%

answer

  1. same table, different denominators
  2. cell over grand total
  3. row or column total over grand total
  4. conditioning changes the denominator
  5. row percent is not column percent

basics

~20 s

A joint proportion divides a cell count by the grand total. A marginal proportion divides a row or column total by the grand total. A conditional proportion divides a cell by its own row or column total.

solid answer

~40 s

All three come off the same table; only the denominator changes. Take 500 respondents cross-tabbed by device (mobile, desktop) against plan (free, basic, pro), with 30 mobile users on pro, 300 mobile users overall and 80 pro users overall. The joint proportion of mobile-and-pro is 30/500 = 0.06, the share of the whole sample. The marginal proportion on pro is 80/500 = 0.16, the plan mix ignoring device. The conditional proportion `P(pro | mobile)` is 30/300 = 0.10, the plan mix inside the mobile group. Comparing groups is a conditional question, so I read row percentages within each device. And conditioning is not symmetric: `P(mobile | pro)` is 30/80 = 0.375, a different number answering a different question.

go deeper

for a junior

Be ready to compute all three off a small table live and say in words what each answers. Practise naming the denominator every time you speak a percentage.

for a middle

Explain the identity joint = marginal times conditional, and show how comparing each row's conditional distribution with the column marginal is the definition of independence in a sample.

for a senior

Demonstrate that you choose the denominator from the business question, and that you catch a colleague reporting the reversed conditional in a deck before it drives a decision.

for a principal

Own the reporting convention: whether dashboards default to row percentages, column percentages or both, and how base rates are surfaced so teams do not read a rare-group rate as a population fact.

## One table, three denominators A cross-tab (contingency table) counts how many observations fall into each combination of two categorical variables. Suppose 500 survey respondents are cross-tabulated by device type against subscription plan: | | Free | Basic | Pro | Row total | | --- | --- | --- | --- | --- | | Mobile | 180 | 90 | 30 | 300 | | Desktop | 80 | 70 | 50 | 200 | | Column total | 260 | 160 | 80 | 500 | Every proportion you can quote from this table is a count divided by something. What distinguishes the three families is which something. **Joint proportion** = cell count / grand total. Mobile-and-pro is 30/500 = 0.06. It answers: what share of everybody is in this exact combination? The nine joint proportions of the six cells sum to 1 across the whole table. **Marginal proportion** = row or column total / grand total. Mobile is 300/500 = 0.60; pro is 80/500 = 0.16. Marginals describe one variable on its own, collapsing the other away. The row marginals sum to 1, and separately the column marginals sum to 1. **Conditional proportion** = cell count / its own row or column total. Inside mobile, the plan split is 180/300 = 0.60 free, 90/300 = 0.30 basic, 30/300 = 0.10 pro; those sum to 1 because they are a distribution within a single row. Inside desktop it is 0.40, 0.35, 0.25. Conditional proportions answer group-comparison questions: do mobile and desktop users pick different plans? ## How they hang together Joint = marginal x conditional. `P(mobile and pro) = P(mobile) x P(pro | mobile) = 0.60 x 0.10 = 0.06`, matching the direct calculation. Rearranged, a conditional is a joint divided by a marginal, which is the whole content of conditional probability applied to counts. ## Conditioning direction matters `P(pro | mobile) = 30/300 = 0.10` and `P(mobile | pro) = 30/80 = 0.375` share a numerator and nothing else. The first says pro is a rare choice among mobile users; the second says most pro subscribers happen to be on mobile. Both are true, and swapping them is the single most common cross-tab error, because the marginal group sizes (300 mobile vs 80 pro) are very different. ## Why this is the foundation for association measures Independence between the two variables means every row's conditional distribution equals the column marginal distribution. Here the pro marginal is 0.16, while pro is 0.10 among mobile users and 0.25 among desktop users. The rows differ from the marginal and from each other, so the variables are associated in this sample. The size of that departure, aggregated across all cells, is exactly what a contingency-table coefficient such as phi or Cramer's V compresses into a single number between 0 and 1. ## Practical habits Always state the denominator out loud when reporting a percentage: 10 percent of mobile users chose pro is unambiguous, while 10 percent of mobile pro users is meaningless. Decide first whether the question is about the whole population (joint), one variable alone (marginal) or a comparison between groups (conditional), then pick the denominator that answers it. When group sizes differ sharply, raw cell counts are actively misleading: 180 free mobile users versus 80 free desktop users looks like a device gap until you notice mobile is 60 percent of the sample and the free rates are 0.60 versus 0.40 in the other direction from what the raw counts suggest at a glance.

  • Why does flipping the conditioning direction change the number so much?
    Because the denominators are different marginal groups. With 30 mobile pro users, 300 mobile users and 80 pro users, `P(pro | mobile)` is 30/300 = 0.10 but `P(mobile | pro)` is 30/80 = 0.375. Whenever one marginal is far larger than the other, the two conditionals diverge, which is why base rates have to be quoted alongside any conditional percentage.
  • How can you tell from the cross-tab alone whether the two variables look associated?
    Convert each row to conditional proportions and compare them with the column marginal. If every row's plan mix matched the overall plan mix, the variables would be independent in the sample. Here pro is 16 percent overall but 10 percent of mobile users and 25 percent of desktop users, so the rows depart from the marginal. Summarising the size of that departure is what phi and Cramer's V do.
  • When is a joint proportion the right thing to report rather than a conditional one?
    When the question is about volume across the whole population rather than rates within a group. Sizing a support queue or a revenue segment is joint: mobile pro users are 6 percent of all respondents. Asking whether device type influences plan choice is conditional, because it needs each group's rate on its own denominator.

Same photograph, three crops: the whole frame, one edge strip, or a zoom inside one row.

saying these in an interview costs you the question

  • Reads row percentages as if they were column percentages
  • Confuses P(A given B) with P(B given A)
  • Quotes a percentage without naming its denominator
  • Compares raw cell counts while group sizes differ wildly
  • Thinks joint proportions within a row sum to 1

context

open as a page

Ice-cream sales correlate with drowning deaths — what can and cannot be concluded from that?

level: juniorimportance: must knowfreq 85%

basics

~20 s

A correlation only says the two series move up and down together in the observed data. It cannot say ice cream causes drownings: something else, such as hot weather, can drive both, and correlation carries no direction while causation does.

open as a page

What is the difference between sample covariance and Pearson's correlation coefficient r?

level: juniorimportance: must knowfreq 86%

basics

~20 s

Sample covariance measures whether two variables move together, but it carries the product of their units, so its size means little alone. Pearson's r divides covariance by both standard deviations, giving a unitless number between -1 and 1.

open as a page

What is regression to the mean, and when should you expect to see it?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Regression to the mean is the tendency for an extreme measurement to be followed by a less extreme one on remeasurement. Expect it whenever two measurements are correlated but not perfectly, because chance helped produce the extreme.

open as a page

What does Spearman's rank correlation coefficient measure between two variables?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Spearman's rho is a correlation computed on the ranks of the data instead of the raw values. It measures how well two variables move together in a consistently increasing or decreasing pattern, not just along a straight line.

open as a page

How is Cramer's V computed from a contingency table's chi-square statistic, and what does its scale mean?

level: middleimportance: must knowfreq 66%

basics

~10 s

Cramer's V equals sqrt(chi2 / (n x (min(r, c) - 1))), with n the total count and r, c the table dimensions. It rescales the chi-square statistic onto 0 to 1, with no direction.

open as a page

What does Anscombe's quartet show about trusting Pearson's r and other summary statistics?

level: middleimportance: must knowfreq 62%

basics

~20 s

Anscombe's quartet is four eleven-point datasets sharing the same means, variances, correlation near 0.82 and fitted line to two decimals, yet their scatterplots look completely different. Summary statistics do not identify a dataset - plot before concluding.

open as a page

Study hours and exam score correlate at r = 0.8, so what does r-squared = 0.64 mean?

level: middleimportance: must knowfreq 68%

basics

~20 s

Squaring Pearson's r gives shared variance: 0.64 means 64 percent of the variation in exam scores moves with variation in study hours, leaving 36 percent unaccounted for. It is a proportion of variance, not of scores.

open as a page

With r = 0.6 between two exams, what score do you predict for a student who scored 2 SD above the mean?

level: middleimportance: must knowfreq 46%

basics

~20 s

About 1.2 standard deviations above the mean. In standard-deviation units the expected second score is the correlation times the first score, so 0.6 times 2 gives 1.2. The prediction is pulled toward the mean because the correlation is below 1.

open as a page

How is Kendall's tau computed from concordant and discordant pairs?

level: middleimportance: must knowfreq 58%

basics

~20 s

Kendall's tau compares every pair of observations. A pair is concordant when both variables order it the same way, discordant when they disagree. Tau is concordant minus discordant pairs, divided by the total number of pairs.

open as a page

What does the point-biserial correlation between a binary group flag and a continuous test score measure?

level: middleimportance: should knowfreq 48%

basics

~10 s

Point-biserial r is just Pearson's correlation with the binary variable coded 0 and 1. It rescales the gap between the two group means, and shrinks as the two groups become more unequal in size.

open as a page

Why is Pearson's r zero when y equals x squared over a range symmetric about zero?

level: middleimportance: should knowfreq 52%

basics

~20 s

Pearson's r measures straight-line association only. On a parabola centred at zero, the rising and falling halves contribute deviation products of opposite sign that cancel exactly, so r is zero even though y is perfectly determined by x.

open as a page

What do the diagonal and off-diagonal entries of a sample covariance matrix hold?

level: middleimportance: should knowfreq 44%

basics

~20 s

The diagonal holds each variable's sample variance, because a variable's covariance with itself is its variance. Off-diagonal entries hold the pairwise sample covariances, and the matrix is symmetric since covariance does not depend on pair order.

open as a page

For x = 1..10 and y = e^x, why is Spearman's rho exactly 1 while Pearson's r is only about 0.72?

level: middleimportance: should knowfreq 52%

basics

~20 s

The relationship is perfectly monotone but strongly curved. Spearman's rho sees only the ordering, which matches exactly, so it hits 1. Pearson's r measures closeness to a straight line, and an exponential curve is far from straight.

open as a page

Salary varies across 12 departments — how do you quantify how much of salary variation department accounts for?

level: seniorimportance: should knowfreq 38%

basics

~10 s

Use the correlation ratio eta. Eta-squared is the between-group sum of squares over the total sum of squares: the share of salary variance accounted for by department membership. It runs 0 to 1.

open as a page

Why does SAT score correlate weakly with college GPA when measured only among admitted students?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Admission selects on the score, so admitted students span a narrow score range. Correlation scales with how much the predictor varies relative to the leftover scatter, so truncating that spread shrinks r even though the underlying relationship is unchanged.

open as a page

How can a single extreme point push Pearson's r from near zero to 0.8?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Pearson's r sums cross-products of deviations, so one point far from both means can contribute a term that outweighs all the others and dominates the statistic. Plot the scatter and recompute r with each row removed to catch it.

open as a page

Crashes fell 35% at junctions where speed cameras were installed after a record-bad year — how much credit do the cameras deserve?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Some of the drop is regression to the mean: the junctions were chosen for having an extreme year. Estimate the camera effect against untreated junctions selected by the same rule, not against the treated sites' own worst year.

open as a page

How do heavy ties on a 1-5 satisfaction scale affect Spearman's rho and Kendall's tau?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Ties break the simple formulas. Spearman's rho needs averaged midranks for each tied block and a full correlation on them. Kendall's tau needs the tie-corrected tau-b, because tau-a cannot reach 1 once many pairs are tied.

open as a page

Your org pilots every programme on its worst-performing segment and reports the rebound as impact — how do you fix that?

level: principalimportance: should knowfreq 34%

basics

~20 s

Change the evaluation, not the targeting. Selecting the worst segment guarantees a rebound whatever the programme does, so require an untreated comparison chosen by the same extreme rule and a published expected-rebound figure any claimed effect must beat.

open as a page

Flight instructors report that praise after a great landing precedes a worse one — what explains this?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Regression to the mean, not the praise. Exceptional landings are partly luck, so the next attempt is usually closer to the trainee's normal standard whatever the instructor says. The same logic makes criticism after a terrible landing look effective.

open as a page

A teammate reports Cramer's V of 0.35 between two categorical columns — what do you check before acting on it?

level: seniorimportance: nice to knowfreq 26%

basics

~10 s

Check sample size and table shape first: Cramer's V is biased upward, so a large sparse table on few rows gives sizeable values from noise. Then read the cell counts and conditional row proportions.

open as a page

Why can a state-level correlation between income and vote share mislead about individual voters?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

A state-level correlation describes states, not people. Averaging discards within-state variation, so the state-level coefficient can be stronger than, or opposite in sign to, the individual-level one. Reading it as a fact about voters is the ecological fallacy.

open as a page

Pearson's r = 0.9: why does that number not tell you how big the effect is in real units?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

Pearson's r is standardised, so it reports how tightly points hug a straight line, not how steep that line is. Dividing by both spreads removes the units, leaving r free to accompany any slope at all.

open as a page

Why does Kendall's tau usually come out smaller in magnitude than Spearman's rho on the same data?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

They are different scales, not competing estimates of one quantity. Tau is a difference of pair-agreement probabilities; Spearman's rho is a correlation of rank positions that weights large rank displacements heavily. Tau typically lands lower.

open as a page