A ratio divides one column's mean by another's, and each column has holes in different rows — why is that ratio not over one population, and what fixes it?
answer
- one column, the two routes agree
- skipping deletes per column
- removal deletes per row
- numerator and denominator, different populations
- a shared population costs sample size
basics
~20 sBecause skipping happens per column: each mean drops its own column's holes, so the numerator and denominator are averages over different, overlapping sets of rows. Restricting to the rows present in both columns first gives the two numbers one shared population.
solid answer
~50 sFor a single column, stepping over the holes and removing the incomplete rows beforehand are the same operation: same set of values, same answer. The moment a second column enters they part, because skipping deletes per column while removal deletes per row. One mean is then over the 9,100 rows that recorded spend and the other over the 8,700 that recorded visits, while only 8,200 recorded both. The ratio is a number from one population divided by a number from another, and if the holes have a cause related to the values — the largest accounts failing to report spend, say — it is biased rather than merely noisier. The fix is to choose the population explicitly: keep the rows present in both columns and compute both means there. It costs sample size, 8,200 instead of 9,100, and buys two numbers that are actually comparable.
code
pseudocode · 13 linesrows = 10000
present(spend) = 9100 # 900 rows never recorded spend
present(visits) = 8700 # 1300 rows never recorded visits
present(spend and visits) = 8200
# per-column skipping: each mean picks its own population
a = mean_of(spend, over = present(spend)) # divided by 9100
b = mean_of(visits, over = present(visits)) # divided by 8700
ratio_per_column = a / b # 900 rows are in a's set and not in b's
# one shared population, chosen before either mean is taken
kept = rows where spend is present and visits is present # 8200
ratio_shared = mean_of(spend, over = kept) / mean_of(visits, over = kept)go deeper
Recall that each column has its own holes, so two averages from one table need not describe the same records. Knowing that the question exists is enough at this level.
State the arity rule and demonstrate it with counts: identical for one column, two populations for two, and the restriction to rows present in both as the way to force one population.
Diagnose it in work that already shipped — a rate whose two halves come from different row sets — and weigh the fix honestly, naming the sample size given up and whether the absence has a cause related to the values.
Decide what the organisation guarantees about published two-column metrics: a declared population per metric, counts carried in the output, and a rule for what happens when the overlap between two fields falls below a level the metric can support.
## One column: the two routes are the same operation Take a mean over a single column. Route A steps over the holes and folds what is left. Route B removes the incomplete rows from the table first and then folds. The values skipped by A and the rows removed by B are the same set, so the two divide by the same count and return the same number. The only difference is **blast radius**: A leaves those rows in the table for every later step, B removes them permanently. That distinction matters for the rest of the pipeline, but it does not move this number. ## Two columns: two populations Add a second column and the equivalence breaks, for one reason: **skipping deletes per column, removal deletes per row.** Worked, on 10,000 rows: | Set | Rows | |---|---| | `spend` present | 9,100 | | `visits` present | 8,700 | | both present | 8,200 | | `spend` present, `visits` absent | 900 | | `visits` present, `spend` absent | 500 | | neither present | 400 | Compute each mean with per-column skipping and the numerator is an average over 9,100 rows while the denominator is an average over 8,700. Those two sets overlap in 8,200 rows and disagree about 1,400 of them. The ratio is not "the ratio for this dataset" — it is a number from one population divided by a number from another, and there is no population it describes. The same is true of any pairing of two columns: a ratio of totals, a per-record rate assembled from two fields, any statistic that reads two columns at once. The arity is the whole test. ## Noise, or bias If the holes fell at random, the two populations would be random samples of the same thing and the ratio would simply be noisier than it looks. That assumption is usually wrong, because absence has a cause and the cause is often related to the value that is missing: - Accounts too large or too complex to report on time are missing from the numerator's population and present in the denominator's. - A sensor that fails under load records nothing exactly when the reading would have been extreme. - A field that became mandatory last quarter is present for new records and absent for old ones, so the two populations differ by cohort as well as by size. In each case the number is wrong **in a direction**, and reporting the two counts does not repair it — it only tells the reader that the two numbers are not commensurable. ## Choosing a population, and paying for it 1. **One shared population.** Keep the rows present in both columns, then compute both means over those rows. The two numbers are now comparable, and the ratio describes an identifiable set of records. 2. **No ratio.** Report each mean with its own count and decline to divide them. Sometimes this is the honest deliverable, particularly when the overlap is small. 3. **State the population in the output.** Whichever of the two you pick, the count travels with the number, because a reader cannot reconstruct it. Option 1 costs sample size: 8,200 rather than 9,100 and 8,700. It also throws away rows that were informative about one of the two columns. That is a real cost and you should say so out loud rather than pretend the restriction is free — the point is that it buys the one property the ratio needs, which is a denominator and a numerator describing the same records. ## Where this hides in practice - **A matrix of pairwise statistics over many columns.** Where a two-column surface restricts to the rows present in both, every cell of the matrix has its own population, so the cells are not mutually comparable and the matrix need not even be internally consistent. Designs differ in the rule they apply, so check rather than assume. - **A ratio of two totals**, which looks safer than a ratio of two means and is not: each total is summed over its own set of rows. - **A rate defined per record**, where the numerator was counted over the rows that reported one field and the denominator over the rows that reported another. - **A metric that was correct when written**, because at the time both columns were complete, and became a population mismatch the first month one of them degraded upstream. ## What to say in the interview Lead with the arity rule — one column, identical; two columns, two populations — then put numbers on it, then name the fix and its cost in the same breath. If you want one more beat, note that the skip-absent flag does not help here: it governs one call over one column, and the problem is a relationship between two calls.
- When are the two routes genuinely interchangeable?When exactly one column is read. The values an aggregate steps over and the rows you would have removed are then the same set, so both divide by the same count and return the same number. Arity is the entire test, and it is worth stating as a rule rather than checking case by case.
- What should make you suspicious of a matrix of pairwise statistics across many columns?That each cell may carry its own population. Where the surface restricts each pair to the rows present in both, every cell keeps a different set of rows, so the cells are not comparable with each other and the matrix need not be internally consistent. Designs differ, so establish which rule is in force.
- Why does the direction of the error matter more than its size here?Size you can bound by publishing the counts. Direction depends on why the values are absent, and if the records failing to report one field are systematically the largest, that mean is low and the ratio is wrong predictably. No denominator footnote repairs a systematic gap.
saying these in an interview costs you the question
- Says skipping and removing incomplete rows always give the same answer
- Treats the population gap as rounding rather than as possible bias
- Assumes writing the skip flag on each call repairs the mismatch
- Publishes a ratio with neither mean's denominator attached
- Believes holes in two columns fall on the same rows
- Thinks a ratio of two totals is safer than a ratio of two means