skip to content

A dashboard reports overall conversion as the unweighted mean of three teams' rates — why is that wrong?

level: seniorimportance: must knowfreq 60%

answer

  1. a rate has a denominator you discarded
  2. every team counted equally, sizes were not
  3. reconstruct the totals instead
  4. weights are the group sizes
  5. safe only with equal denominators

basics

~10 s

Averaging rates weights every team equally regardless of size, so a tiny team counts as much as a huge one. The correct overall rate is total conversions over total users — a size-weighted mean.

solid answer

~50 s

An unweighted mean of rates answers the question what is the average team, not what is the overall rate. Take three teams converting 5 of 10, 30 of 100 and 8 of 40 users: the rates are 50%, 30% and 20%, whose plain mean is 33.3%. The real overall rate is `(5 + 30 + 8) / (10 + 100 + 40) = 43 / 150 = 28.7%`. The gap exists because the 10-user team carries a third of the reported average while contributing under 7% of the users. The fix is to recompute from the raw numerator and denominator, or equivalently to take a weighted mean with each team's user count as the weight — the two are algebraically the same number. The unweighted mean is correct only when the denominators are equal, so a dashboard should carry counts next to every rate to make that check possible.

go deeper

for a junior

Remember that percentages cannot simply be averaged. Recompute a rate from total numerator over total denominator, and always show the counts a percentage was built from.

for a middle

Explain the weighted-mean identity: weighting each rate by its denominator reconstructs the numerators, so it equals the pooled rate exactly, and the unweighted version coincides only when the denominators match.

for a senior

Show how you would find this in a live pipeline — reconcile the metric against raw totals, hunt for two-stage aggregation, and judge whether the reporting unit is the user or the group before choosing the fix.

for a principal

Own the metric definition. Decide and document the unit of analysis for each headline number, and push storage toward additive numerator and denominator columns so no downstream consumer can average rates by accident.

## The defect A rate is a ratio: `rate = numerator / denominator`. Rates do not add or average the way plain numbers do, because averaging them throws away the denominators. The **average-of-averages error** is what happens when a pipeline computes a rate per group, then averages those rates as if they were ordinary measurements. A worked example makes the size of the error visible. | Team | Conversions | Users | Rate | |------|-------------|-------|------| | A | 5 | 10 | 50.0% | | B | 30 | 100 | 30.0% | | C | 8 | 40 | 20.0% | | **All** | **43** | **150** | **28.7%** | The unweighted mean of the three rates is `(50 + 30 + 20) / 3 = 33.3%`. The true overall rate is `43 / 150 = 28.7%`. The dashboard is overstating conversion by more than four and a half percentage points, and the direction of the error is not random — it is whichever way the small groups happen to lean. ## Why the weighted mean is the same computation The correct overall rate can be written as a weighted mean of the group rates, with weights equal to the denominators: `overall = sum(w_i * r_i) / sum(w_i)`, where `w_i` is group i's user count and `r_i` its rate. Check it: `(10*0.50 + 100*0.30 + 40*0.20) / 150 = (5 + 30 + 8) / 150 = 43 / 150 = 28.7%`. This is not a second method that happens to agree — it is the same arithmetic rearranged. `w_i * r_i` reconstructs group i's numerator, so the weighted mean is literally total numerator over total denominator. That identity gives the exact condition under which the naive version is safe. If every `w_i` is the same value w, then `sum(w * r_i) / (n * w) = sum(r_i) / n`, the unweighted mean. **Equal denominators is the only structural guarantee**; equal rates makes them agree too, but that is a coincidence of the data rather than a property you can rely on. ## Where it shows up in real work - **Rolling up team, region or segment metrics** into a company number. - **Averaging daily rates over a month.** Weekends and holidays carry far less traffic but the same one-thirtieth weight in a plain mean, so a quiet day with a freak rate can move the monthly headline. Weight by daily traffic, or simply divide the month's total conversions by the month's total sessions. - **Averaging per-user rates.** A user with three sessions and a user with three hundred contribute equally to a mean of per-user rates. Whether that is wrong depends on the question: if you genuinely want the typical *user*, per-user weighting is the intent; if you want the overall *session* conversion rate, it is a defect. - **Two-stage aggregation in a query or job.** Any `GROUP BY` that produces a rate followed by an outer aggregate over those rates is the pattern to look for in a review. ## The point that separates a good answer The fix is not always to weight. It is to state the population you are averaging over and then compute consistently with it. - *What is our conversion rate?* — the unit is the user or session. Pool the raw counts. - *How does a typical team perform?* — the unit is the team. The unweighted mean is now the right statistic, and its weakness is different: a ten-user team's rate is a noisy estimate, so the average team rate is dominated by sampling noise from the small teams. Report it with the spread, or with counts alongside. Stating that distinction out loud is what an interviewer is listening for. A candidate who says only that weighting is required has learned a rule; a candidate who asks which unit the metric is defined over has understood the problem. ## Diagnosing it in the wild Three cheap checks: 1. **Reconcile against the total.** Compute the metric from raw numerator and denominator and compare. Any disagreement means an intermediate average exists somewhere. 2. **Look for wildly unequal group sizes.** Equal denominators make the bug invisible; skewed group sizes make it large. Both facts follow from the weighted-mean identity. 3. **Put counts next to rates everywhere.** A rate with no denominator on screen cannot be sanity-checked by the reader, and a 50% built on ten users deserves different trust from a 50% built on ten thousand. ## Guardrails worth owning Store and ship numerator and denominator as separate additive columns and derive the rate at read time. Additive quantities roll up correctly under any grouping; rates do not. That one modelling decision removes the whole class of bug from the reporting layer rather than relying on every analyst to remember the rule.

  • When is the unweighted mean of group rates actually the right number to report?
    When the unit of analysis is the group itself — for example, how the typical team or store performs, where each one should count once regardless of size. Its weakness then shifts to precision rather than bias: rates from tiny groups are noisy estimates, so report counts and spread alongside, and be explicit that the figure describes teams, not users.
  • How would you aggregate a month of daily conversion rates into one monthly number?
    Divide the month's total conversions by the month's total sessions, which is a traffic-weighted mean of the daily rates. A plain mean of thirty daily rates gives a quiet holiday the same weight as the heaviest trading day, letting a low-volume outlier day move the headline several points.
  • What data model change prevents this bug from recurring?
    Persist the numerator and denominator as separate additive columns and compute the rate only at presentation time. Sums roll up correctly under any grouping, so any slice recomputes its own rate from raw counts. Storing precomputed rates invites a downstream average of them, which is exactly the failure mode.

Two shops sell at 100% and 0% conversion. If one served a single customer and the other served a thousand, calling the chain 50% converts a rounding-scale shop into half the business.

saying these in an interview costs you the question

  • Says averaging the group rates is fine because they are all percentages
  • Believes equal group counts are unnecessary for the naive mean to work
  • Reports rates on dashboards without the denominators
  • Thinks the error is small whenever the rates look similar
  • Assumes weighting is always right without asking what the unit of analysis is

context