skip to content

In a cross-tab of device type by plan, what is the difference between joint, marginal and conditional proportions?

level: juniorimportance: must knowfreq 62%

answer

  1. same table, different denominators
  2. cell over grand total
  3. row or column total over grand total
  4. conditioning changes the denominator
  5. row percent is not column percent

basics

~20 s

A joint proportion divides a cell count by the grand total. A marginal proportion divides a row or column total by the grand total. A conditional proportion divides a cell by its own row or column total.

solid answer

~40 s

All three come off the same table; only the denominator changes. Take 500 respondents cross-tabbed by device (mobile, desktop) against plan (free, basic, pro), with 30 mobile users on pro, 300 mobile users overall and 80 pro users overall. The joint proportion of mobile-and-pro is 30/500 = 0.06, the share of the whole sample. The marginal proportion on pro is 80/500 = 0.16, the plan mix ignoring device. The conditional proportion `P(pro | mobile)` is 30/300 = 0.10, the plan mix inside the mobile group. Comparing groups is a conditional question, so I read row percentages within each device. And conditioning is not symmetric: `P(mobile | pro)` is 30/80 = 0.375, a different number answering a different question.

go deeper

for a junior

Be ready to compute all three off a small table live and say in words what each answers. Practise naming the denominator every time you speak a percentage.

for a middle

Explain the identity joint = marginal times conditional, and show how comparing each row's conditional distribution with the column marginal is the definition of independence in a sample.

for a senior

Demonstrate that you choose the denominator from the business question, and that you catch a colleague reporting the reversed conditional in a deck before it drives a decision.

for a principal

Own the reporting convention: whether dashboards default to row percentages, column percentages or both, and how base rates are surfaced so teams do not read a rare-group rate as a population fact.

## One table, three denominators A cross-tab (contingency table) counts how many observations fall into each combination of two categorical variables. Suppose 500 survey respondents are cross-tabulated by device type against subscription plan: | | Free | Basic | Pro | Row total | | --- | --- | --- | --- | --- | | Mobile | 180 | 90 | 30 | 300 | | Desktop | 80 | 70 | 50 | 200 | | Column total | 260 | 160 | 80 | 500 | Every proportion you can quote from this table is a count divided by something. What distinguishes the three families is which something. **Joint proportion** = cell count / grand total. Mobile-and-pro is 30/500 = 0.06. It answers: what share of everybody is in this exact combination? The nine joint proportions of the six cells sum to 1 across the whole table. **Marginal proportion** = row or column total / grand total. Mobile is 300/500 = 0.60; pro is 80/500 = 0.16. Marginals describe one variable on its own, collapsing the other away. The row marginals sum to 1, and separately the column marginals sum to 1. **Conditional proportion** = cell count / its own row or column total. Inside mobile, the plan split is 180/300 = 0.60 free, 90/300 = 0.30 basic, 30/300 = 0.10 pro; those sum to 1 because they are a distribution within a single row. Inside desktop it is 0.40, 0.35, 0.25. Conditional proportions answer group-comparison questions: do mobile and desktop users pick different plans? ## How they hang together Joint = marginal x conditional. `P(mobile and pro) = P(mobile) x P(pro | mobile) = 0.60 x 0.10 = 0.06`, matching the direct calculation. Rearranged, a conditional is a joint divided by a marginal, which is the whole content of conditional probability applied to counts. ## Conditioning direction matters `P(pro | mobile) = 30/300 = 0.10` and `P(mobile | pro) = 30/80 = 0.375` share a numerator and nothing else. The first says pro is a rare choice among mobile users; the second says most pro subscribers happen to be on mobile. Both are true, and swapping them is the single most common cross-tab error, because the marginal group sizes (300 mobile vs 80 pro) are very different. ## Why this is the foundation for association measures Independence between the two variables means every row's conditional distribution equals the column marginal distribution. Here the pro marginal is 0.16, while pro is 0.10 among mobile users and 0.25 among desktop users. The rows differ from the marginal and from each other, so the variables are associated in this sample. The size of that departure, aggregated across all cells, is exactly what a contingency-table coefficient such as phi or Cramer's V compresses into a single number between 0 and 1. ## Practical habits Always state the denominator out loud when reporting a percentage: 10 percent of mobile users chose pro is unambiguous, while 10 percent of mobile pro users is meaningless. Decide first whether the question is about the whole population (joint), one variable alone (marginal) or a comparison between groups (conditional), then pick the denominator that answers it. When group sizes differ sharply, raw cell counts are actively misleading: 180 free mobile users versus 80 free desktop users looks like a device gap until you notice mobile is 60 percent of the sample and the free rates are 0.60 versus 0.40 in the other direction from what the raw counts suggest at a glance.

  • Why does flipping the conditioning direction change the number so much?
    Because the denominators are different marginal groups. With 30 mobile pro users, 300 mobile users and 80 pro users, `P(pro | mobile)` is 30/300 = 0.10 but `P(mobile | pro)` is 30/80 = 0.375. Whenever one marginal is far larger than the other, the two conditionals diverge, which is why base rates have to be quoted alongside any conditional percentage.
  • How can you tell from the cross-tab alone whether the two variables look associated?
    Convert each row to conditional proportions and compare them with the column marginal. If every row's plan mix matched the overall plan mix, the variables would be independent in the sample. Here pro is 16 percent overall but 10 percent of mobile users and 25 percent of desktop users, so the rows depart from the marginal. Summarising the size of that departure is what phi and Cramer's V do.
  • When is a joint proportion the right thing to report rather than a conditional one?
    When the question is about volume across the whole population rather than rates within a group. Sizing a support queue or a revenue segment is joint: mobile pro users are 6 percent of all respondents. Asking whether device type influences plan choice is conditional, because it needs each group's rate on its own denominator.

Same photograph, three crops: the whole frame, one edge strip, or a zoom inside one row.

saying these in an interview costs you the question

  • Reads row percentages as if they were column percentages
  • Confuses P(A given B) with P(B given A)
  • Quotes a percentage without naming its denominator
  • Compares raw cell counts while group sizes differ wildly
  • Thinks joint proportions within a row sum to 1

context