skip to content

From a joint table of device and plan tier, how do you compute a marginal and a conditional distribution?

level: juniorimportance: should knowfreq 58%

answer

  1. sum along one direction
  2. row and column totals
  3. slice, then renormalise to 1
  4. divide the cell by the row total
  5. compare a cell with the product of margins

basics

~20 s

Sum a row or column of joint probabilities to get a marginal: P(mobile) = P(mobile, free) + P(mobile, pro). Divide a single joint cell by that marginal to get a conditional: P(pro | mobile) = P(mobile, pro) / P(mobile).

solid answer

~40 s

A joint table gives `P(X = x, Y = y)` for every combination, and the whole table sums to 1. **Marginalising** means summing out the variable you no longer care about: the row totals give the distribution of device, the column totals give the distribution of plan tier. **Conditioning** means restricting to one row (or column) and renormalising so it sums to 1: `P(pro | mobile) = P(mobile, pro) / P(mobile)`. With cells `P(mobile, free) = 0.40`, `P(mobile, pro) = 0.20`, `P(desktop, free) = 0.15`, `P(desktop, pro) = 0.25`, the marginals are `P(mobile) = 0.60` and `P(pro) = 0.45`, and `P(pro | mobile) = 0.20 / 0.60 = 1/3` while `P(pro | desktop) = 0.25 / 0.40 = 0.625`. Because those two conditionals differ, the variables are not independent.

go deeper

for a junior

Practise until row totals, column totals and a divide-by-the-row-total conditional are automatic on a 2x2 table. Say the denominator out loud so you never invert the conditioning.

for a middle

Be ready to state the continuous analogue: marginals come from integrating out the other variable, and the conditional density is the joint divided by the marginal wherever that marginal is positive.

for a senior

Show that you check independence cell by cell rather than eyeballing totals, and explain why aggregating a joint table over a third variable can flip the apparent direction of a conditional.

for a principal

Own what the organisation stores and reports: per-variable summaries permanently discard the dependence structure, so decide up front which cross-tabulations must be retained for the questions the business will ask later.

## The joint table A joint distribution over two categorical variables is a table of probabilities, one per cell, that sums to 1. Take device (mobile, desktop) against plan tier (free, pro): | | free | pro | row total | |---------|------|------|-----------| | mobile | 0.40 | 0.20 | 0.60 | | desktop | 0.15 | 0.25 | 0.40 | | col total | 0.55 | 0.45 | 1.00 | Every inner cell is a joint probability, e.g. `P(device = mobile, tier = pro) = 0.20`. ## Marginalising: summing out The **marginal** distribution of one variable is obtained by summing the joint probabilities over all values of the other: `P(mobile) = P(mobile, free) + P(mobile, pro) = 0.40 + 0.20 = 0.60` `P(pro) = P(mobile, pro) + P(desktop, pro) = 0.20 + 0.25 = 0.45` The name comes from writing these totals in the margins of the table. Each marginal is itself a valid distribution: `P(mobile) + P(desktop) = 1` and `P(free) + P(pro) = 1`. The continuous analogue replaces the sum with an integral. If (X, Y) has joint density `f(x,y)`, then `f_X(x) = integral over y of f(x,y) dy` and symmetrically for `f_Y(y)`. ## Conditioning: slice and renormalise The **conditional** distribution of tier given device is one row of the table, divided by that row's total so it sums to 1: `P(pro | mobile) = P(mobile, pro) / P(mobile) = 0.20 / 0.60 = 1/3 ~= 0.333` `P(free | mobile) = 0.40 / 0.60 = 2/3` Those two add to 1, as a distribution must. Doing the same on the second row: `P(pro | desktop) = 0.25 / 0.40 = 0.625` Conditioning in the other direction slices a column instead: `P(mobile | pro) = P(mobile, pro) / P(pro) = 0.20 / 0.45 = 4/9 ~= 0.444` Note that `P(pro | mobile) = 1/3` and `P(mobile | pro) = 4/9` are different quantities. Swapping which variable is on the right of the bar changes the question being asked, and it changes the denominator from a row total to a column total. The continuous version is `f_{Y|X}(y | x) = f(x,y) / f_X(x)`, defined wherever `f_X(x) > 0`. ## Reading independence off the table X and Y are independent exactly when every cell equals the product of its two margins: `P(x, y) = P(x) * P(y)` for all pairs. Here `P(mobile) * P(pro) = 0.60 * 0.45 = 0.27`, while the actual cell is 0.20. One mismatch is enough — these variables are dependent. Equivalently, independence means the conditional distribution of tier is the same in every row; we found 1/3 versus 0.625, so it is not. A practical shortcut: for a 2x2 table, checking a single cell against the product of its margins settles it, because the other three cells are then forced by the row and column totals. ## The direction of information loss Marginals are computable from the joint, but the joint is **not** recoverable from the marginals. Many different joint tables share the same row and column totals — the totals 0.60/0.40 and 0.55/0.45 above are also consistent with the independent table (0.33, 0.27 / 0.22, 0.18) and with a much more skewed one. Only under an explicit independence assumption can you multiply the marginals back into a joint. This is why summary statistics per variable can never answer questions about how the variables move together. ## Sanity checks to run out loud - Do all cells sum to 1? - Does each conditional distribution sum to 1 across the values of the conditioned variable? - Is every conditional probability at least as large as the corresponding joint cell? It must be, because dividing by a number below 1 can only increase it. - Are you dividing by the right margin? `P(A | B)` divides by `P(B)`, the thing you are told, not by `P(A)`.

  • How would you check independence directly from that joint table?
    Compare each cell with the product of its two marginal totals. With `P(mobile) = 0.60` and `P(pro) = 0.45`, independence would require the cell to be `0.27`; it is `0.20`, so they are dependent. Equivalently, compare the conditional distributions across rows: `P(pro | mobile) = 1/3` versus `P(pro | desktop) = 0.625`. Independence means every row has the identical conditional profile.
  • Given only the two marginal distributions, can you reconstruct the joint table?
    No. Infinitely many joint tables share the same row and column totals, differing in how much the variables move together. Only if you assume independence can you fill it in, by multiplying the marginals cell by cell. This is why per-variable summaries can never answer a question about association — the dependence information is destroyed by marginalising.
  • Why is P(pro | mobile) different from P(mobile | pro)?
    They share the same numerator, the joint cell 0.20, but divide by different totals: the first by `P(mobile) = 0.60` giving 1/3, the second by `P(pro) = 0.45` giving 4/9. Conditioning restricts attention to a different subpopulation in each case. Treating the two as interchangeable is the standard inversion error.

A marginal is the shadow the table casts on one wall; a conditional is one slice of the table held up on its own and rescaled to full height.

saying these in an interview costs you the question

  • Divides the joint cell by the wrong marginal total
  • Treats P(A|B) and P(B|A) as the same number
  • Thinks the marginals determine the joint table
  • Forgets that a conditional distribution must sum to 1
  • Reads a joint cell as if it were a conditional probability

context