skip to content

How does the law of total variance split customer spend variance into within- and between-segment parts?

level: seniorimportance: should knowfreq 38%

answer

  1. two terms, and both are non-negative
  2. average of variances plus variance of averages
  3. a conditional mean is itself random
  4. equal segment means kills one term
  5. p(1-p) times the squared mean gap

basics

~20 s

It writes Var(Y) = E[Var(Y | S)] + Var(E[Y | S]): the average spread inside segments plus the spread of the segment means. The first term is within-segment noise, the second is how far apart the segments sit.

solid answer

~50 s

The law of total variance — often called Eve's law — states `Var(Y) = E[Var(Y | S)] + Var(E[Y | S])`, where `S` is the segment label. The first term averages each segment's own variance weighted by segment size: that is the variation you cannot explain by knowing which segment a customer belongs to. The second term is the variance of the conditional means: how far the segment averages sit from the overall average. With two segments at shares `p` and `1-p`, the between term simplifies to `p(1-p)(m1 - m2)^2`. Concretely, if 70% of customers average 50 with variance 100 and 30% average 90 with variance 400, the within term is `0.7*100 + 0.3*400 = 190`, the between term is `0.7*0.3*40^2 = 336`, and total variance is 526 — far more than the weighted average of segment variances alone.

go deeper

for a junior

Recall that total variance splits into two non-negative pieces — spread inside groups and spread between group averages — and that the second piece is what a plain average of group variances leaves out.

for a middle

Write the identity and compute both terms on a two-segment example, including the p(1-p)(m1 - m2)^2 shortcut, and verify against a direct calculation of E[Y^2] - E[Y]^2.

for a senior

Use the split diagnostically: say what a dominant between term implies for reporting and targeting, flag the bimodality risk, and explain why a pooled standard deviation built from segment variances alone mis-sizes intervals.

for a principal

Own the decision the decomposition informs — whether a segmentation earns its place in the metric stack, how much explanatory power justifies the extra reporting complexity, and when to hunt for a better partition instead.

## Start with the mean Before variance, get the expectation version straight. The **law of total expectation** says that conditioning on a partition and averaging back recovers the unconditional mean: ``` E[Y] = sum over segments s of E[Y | S = s] * P(S = s) ``` written compactly as `E[Y] = E[E[Y | S]]`. The inner `E[Y | S]` is itself a random quantity: it takes the value `m1` when the customer is in segment 1 and `m2` in segment 2. That is the key mental shift — a conditional mean is a random variable, because which condition you land in is random. ## The law of total variance Once you accept that `E[Y | S]` is a random variable, it has a variance of its own, and ``` Var(Y) = E[Var(Y | S)] + Var(E[Y | S]) ``` The two pieces have plain-English readings: - **`E[Var(Y | S)]` — within-segment (unexplained) variance.** Take each segment's own variance, average them weighted by segment size. This is the spread that remains after you are told which segment the customer is in. - **`Var(E[Y | S])` — between-segment (explained) variance.** Treat each customer as if they spent exactly their segment's average, and measure the variance of *that*. It is large when the segments have very different means. Every unit of variance in `Y` lands in exactly one bucket, and the buckets add. ## A worked mixture Two segments: - Segment A: 70% of customers, mean spend 50, variance 100 - Segment B: 30% of customers, mean spend 90, variance 400 **Overall mean:** `0.7*50 + 0.3*90 = 35 + 27 = 62`. **Within term:** `E[Var(Y | S)] = 0.7*100 + 0.3*400 = 70 + 120 = 190`. **Between term:** `Var(E[Y | S]) = 0.7*(50 - 62)^2 + 0.3*(90 - 62)^2 = 0.7*144 + 0.3*784 = 100.8 + 235.2 = 336`. The two-group shortcut confirms it: `p(1-p)(m1 - m2)^2 = 0.7*0.3*1600 = 336`. **Total:** `190 + 336 = 526`, so the standard deviation is about 22.9. Check it directly: `E[Y^2] = 0.7*(100 + 2500) + 0.3*(400 + 8100) = 1820 + 2550 = 4370`, and `4370 - 62^2 = 4370 - 3844 = 526`. The decomposition and the brute-force computation agree, as they must. ## The mistake this prevents The common error is to report the weighted average of segment variances, 190, as the variance of overall spend. That is only correct when every segment has the same mean. Here it understates the true variance by nearly two thirds, and the corresponding standard deviation would be 13.8 instead of 22.9 — a difference that would badly mis-size any interval or capacity plan. The opposite error is adding the raw variances, `100 + 400 = 500`, which happens to land close by coincidence and is structurally wrong: it uses no weights and no means. ## Reading the split The ratio `Var(E[Y | S]) / Var(Y)` is the share of variation explained by the segmentation — here `336 / 526 = 0.64`. Interpretation: - **Between term dominates** (as here): the segment label is a strong predictor of spend. Reporting a single blended average is misleading, and segment-level targets make sense. - **Within term dominates:** customers inside a segment differ far more than the segments differ from one another. The segmentation is cosmetic for this metric, and you should look for a different partition. A large between term also warns that the overall distribution may be bimodal — one hump per segment — so the overall mean can describe almost nobody. ## Refining the partition Splitting segments into finer ones can never increase the within term at the population level. Formally, if `S'` refines `S`, then `E[Var(Y | S')] <= E[Var(Y | S)]`, with the between term rising by the same amount because the total is fixed. Conditioning on more information moves variance from the unexplained bucket to the explained one. In finite samples this is not guaranteed — estimated conditional means and variances get noisier as groups shrink, so a measured within term can wobble upward. ## Where it is used - Deciding whether a segmentation is worth carrying in reporting at all. - Explaining why an overall spend distribution has fatter spread than any single segment. - Any hierarchical or grouped setting where variation naturally decomposes into a group-level part and a within-group part. ## Sanity checks - Both terms are non-negative, so `Var(Y)` is always at least the weighted average of the segment variances. - If all segment means are equal, the between term is exactly zero and `Var(Y) = E[Var(Y | S)]`. - If every segment has zero internal variance, all variance is between-segment.

  • What does it tell you when the between-segment term dominates?
    That knowing the segment tells you most of what you can know about spend. The segment label is a strong predictor, the pooled average may describe nobody, and the overall distribution is likely bimodal. Practically it argues for segment-level targets and reporting rather than a single blended number.
  • How is this connected to the law of total expectation?
    They are the same conditioning move applied to different moments. The law of total expectation gives `E[Y] = E[E[Y | S]]`, and you need those conditional means before you can compute either variance term. The between term is literally the variance of the random variable that the law of total expectation averages out.
  • Can splitting into finer segments ever raise the within-segment term?
    Not at the population level. If the finer partition refines the coarser one, `E[Var(Y | S')] <= E[Var(Y | S)]` and the between term rises by the same amount, since total variance is fixed. In a finite sample it can appear to rise, because conditional means and variances estimated from small groups are noisy.
  • Why is the weighted average of the segment variances not the overall variance?
    Because it discards the spread between the segment means. It equals the total only in the special case where all segments share the same mean. In the 70/30 example it gives 190 against a true 526, understating the standard deviation by roughly 40%.

Think of two towns' income data. Within-variance is how much neighbours differ inside each town; between-variance is how far the two town averages sit apart. National spread needs both.

saying these in an interview costs you the question

  • Reports the weighted average of segment variances as the total
  • Adds the two segment variances together with no weights
  • Ignores the between-group term when the segment means differ
  • Treats a conditional mean as a constant rather than a random variable
  • Confuses the conditional variance with the variance of the conditional mean
  • Claims one of the two terms can be negative

context