skip to content

Law of Total Probability

Splitting a probability across a partition of cases, drawing the tree diagram, and the matching law of total expectation. It is the step most candidates skip on the way to Bayes.

on this pageshow

questions

5

Using the law of total probability, what is P(defective) if line A makes 60% of units at 2% defective and line B 40% at 5%?

level: juniorimportance: must knowfreq 78%

answer

  1. split the sample space into cases first
  2. each case carries a weight
  3. weights are the production shares
  4. multiply along a branch, add across branches
  5. not the plain average of 2% and 5%

basics

~10 s

The overall defect rate is 3.2%. The law of total probability weights each line's defect rate by that line's share of production: 0.60 * 0.02 + 0.40 * 0.05 = 0.032.

solid answer

~40 s

Split the units by which line made them — that is a partition, because every unit comes from exactly one line. The law of total probability then says `P(defective) = P(defective | A) P(A) + P(defective | B) P(B)`, so `0.02 * 0.60 + 0.05 * 0.40 = 0.012 + 0.020 = 0.032`, i.e. 3.2%. The sanity check is a count: out of 1000 units, 600 come from A and yield 12 defects, 400 come from B and yield 20, giving 32 defects per 1000. The answer is not the plain average of 2% and 5% (3.5%) — that would only be right if the two lines produced equal volumes. Note also that the weights are production shares, not defect shares, and they must sum to 1.

go deeper

for a junior

Be ready to write the formula and do the arithmetic out loud without hesitating, and to state the two conditions on the cases: mutually exclusive and exhaustive. Saying 3.5% is a common and costly slip.

for a middle

Explain why the branch products may be added — they are disjoint joint probabilities — and show the frequency check with 1000 units. Expect to be asked what happens if a case is missing or if two cases overlap.

for a senior

Show judgment about where the weights come from in real data: actual output shares over the period in question, not nameplate capacity or a stale plan. Point out that a shift in the mix moves the overall rate even when no line changed.

for a principal

Frame this as the tool for reasoning about aggregates you cannot measure directly, and be clear about its limits: the decomposition is only as good as the partition, and reporting a single blended rate hides mix effects that drive very different decisions.

## What a partition is A **partition** of the sample space is a set of events `B1, B2, ..., Bk` that are **mutually exclusive** (no two can happen together) and **exhaustive** (at least one must happen). Every possible outcome falls into exactly one `Bi`. "Which production line made this unit" is a partition as long as every unit comes from exactly one line and the lines listed cover all output. ## The law itself Take any event `A` — here, "this unit is defective". Because the `Bi` are disjoint and cover everything, they cut `A` into disjoint slices: ``` P(A) = P(A and B1) + P(A and B2) + ... + P(A and Bk) ``` Each slice can be rewritten with the multiplication rule, `P(A and Bi) = P(A | Bi) * P(Bi)`, which gives the **law of total probability**: ``` P(A) = sum over i of P(A | Bi) * P(Bi) ``` In words: take the probability of `A` inside each case, and average those numbers **weighted by how often each case occurs**. ## The worked numbers With `P(A) = 0.60`, `P(B) = 0.40`, `P(defective | A) = 0.02`, `P(defective | B) = 0.05`: ``` P(defective) = 0.02 * 0.60 + 0.05 * 0.40 = 0.012 + 0.020 = 0.032 -> 3.2% ``` The two products, 0.012 and 0.020, are **joint** probabilities: 1.2% of all units are "from line A and defective", 2.0% are "from line B and defective". They are not conditional probabilities, and that is exactly why they can be added. A frequency check makes it concrete. Imagine 1000 units: 600 from A with 2% defective gives 12 bad units; 400 from B with 5% gives 20 bad units; 32 bad units out of 1000 is 3.2%. ## Why 3.5% is wrong Averaging 2% and 5% to get 3.5% silently assumes each line supplies half the output. The unweighted average is the special case where every weight is `1/k`. Because line A — the cleaner line — produces more, the true rate is pulled below the midpoint. Whenever the group with the extreme conditional rate is also the smaller group, the naive average is biased toward that extreme. ## The tree diagram view Draw one branch per element of the partition, labelled with `P(Bi)`. From each, draw sub-branches labelled with the conditional probabilities `P(A | Bi)` and `P(not A | Bi)`. Then: - **multiply along a path** to get that path's joint probability, - **add across the paths** that end in the outcome you care about. The four path probabilities here are 0.588, 0.012, 0.380 and 0.020, and they sum to 1 — a useful arithmetic check. ## Where people go wrong 1. **Using the wrong weights.** The weights must be the probabilities of the conditioning events — actual output shares, not machine capacity, not each line's share of the *defects*. 2. **Overlapping cases.** If a unit could be counted under two categories (say "night shift" and "line B" used together as if they were a partition), pieces are double-counted and the total exceeds the truth. 3. **Non-exhaustive cases.** If a third line quietly supplies 5% of output and is left out, the weights sum to 0.95 and the answer is too small. 4. **Reversing the conditioning.** `P(defective | A)` and `P(A | defective)` are different numbers; only the first belongs in this formula. Here `P(defective | A) = 0.02` while `P(A | defective) = 0.012 / 0.032 = 0.375`. ## Why the law matters beyond the arithmetic It is the standard move for turning an unconditional question that is hard into conditional questions that are easy: you rarely know the overall defect rate directly, but you usually know each line's rate and its volume. It is also the quantity that normalises any inversion of the conditioning — the total probability of the observed event is the denominator you divide by. ## Two extensions worth knowing **Continuous conditioning.** When the conditioning variable `X` is continuous with density `f(x)`, the sum becomes an integral: `P(A) = integral of P(A | X = x) f(x) dx`. The structure is identical — a weighted average of conditional probabilities. **Expectations.** The same decomposition holds for means. If `Y` is the repair cost of a unit, then `E[Y] = E[Y | A] P(A) + E[Y | B] P(B)`. This is the law of total expectation, and it is the workhorse for computing averages case by case. ## Fast sanity checks - The result must lie between the smallest and largest conditional probability (2% and 5% here). - The weights must sum to 1. - If all conditional probabilities are equal, the answer must equal that shared value.

  • Why can't you simply average 2% and 5% to get 3.5%?
    The plain average is the special case where both lines produce equal volumes. Here line A supplies 60% of units, so its 2% rate gets more weight and the true rate falls below the midpoint. Whenever the group with the extreme conditional rate is the smaller group, the unweighted average is pulled toward that extreme and overstates or understates the total.
  • How does the calculation change if a third line supplies part of the output?
    You add one more term: `P(defective) = sum over lines of P(defective | line) P(line)`. The requirement is that the partition stays mutually exclusive and exhaustive, so every unit belongs to exactly one line and the shares sum to 1. If you leave a line out, the weights sum to less than 1 and the total is understated.
  • What does the law of total probability look like when the conditioning variable is continuous?
    The sum becomes an integral. If `X` is continuous with density `f(x)`, then `P(A) = integral of P(A | X = x) f(x) dx`. It is still a weighted average of conditional probabilities, with the density supplying the weights instead of a finite list of case probabilities.
  • Which numbers in this problem are joint probabilities rather than conditional ones?
    The branch products are joint: 0.012 is `P(line A and defective)` and 0.020 is `P(line B and defective)`. The inputs 0.02 and 0.05 are conditional — defect rates *given* the line. Only joint probabilities may be added together, which is why the multiplication step has to come first.

It is a weighted grade: each production line is a course with its own score, and its share of output is the credit weight. A high score in a one-credit course barely moves the GPA.

saying these in an interview costs you the question

  • Averages the conditional rates without weighting by output share
  • Uses overlapping categories that double-count some units
  • Adds conditional probabilities directly instead of joint ones
  • Leaves out a case, so the weights do not sum to one
  • Confuses P(defective given line A) with P(line A given defective)
  • Weights by each line's share of defects instead of its share of production

context

open as a page

What is the expected number of fair coin flips until the first HH, by first-step conditioning?

level: middleimportance: must knowfreq 50%

basics

~20 s

Six flips on average. Condition on the next flip using two states — no progress, or a trailing head. Solving E0 = 1 + 0.5E1 + 0.5E0 and E1 = 1 + 0.5*E0 gives E1 = 4 and E0 = 6.

open as a page

How does the law of total variance split customer spend variance into within- and between-segment parts?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It writes Var(Y) = E[Var(Y | S)] + Var(E[Y | S]): the average spread inside segments plus the spread of the segment means. The first term is within-segment noise, the second is how far apart the segments sit.

open as a page

How do you choose the segment partition when forecasting an aggregate rate as a weighted sum?

level: principalimportance: should knowfreq 30%

basics

~20 s

Pick segments that are mutually exclusive, exhaustive, assignable before the outcome is known, and whose conditional rates are more stable than the aggregate. Then forecast rates and mix separately, and stop splitting once cells get too thin.

open as a page

In gambler's ruin on a fair game, why is the chance of reaching N before 0 from k equal to k/N?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Conditioning on the next bet gives P(k) = 0.5P(k-1) + 0.5P(k+1), so each value is the average of its neighbours — a straight line. The boundaries P(0) = 0 and P(N) = 1 force P(k) = k/N.

open as a page