skip to content

A mean over a 10,000-row column with 1,600 cells recording nothing returns a number — which denominator did it use?

level: middleimportance: must knowfreq 78%

answer

  1. a number came back is already evidence
  2. it divides by what survived the skip
  3. present values, not rows
  4. 8,400, not 10,000
  5. report the count beside the number

basics

~20 s

The count of present values, 8,400 — not the 10,000 rows. An aggregate that steps over the holes removes them first, so the reported number is an average of what was recorded, over a population smaller than the table.

solid answer

~50 s

Getting a number back at all tells you the surface stepped over the holes rather than handing absence back. Having stepped over them, it divides by what is left: the count of present values. With 1,600 holes in 10,000 rows, the mean is the sum of the 8,400 recorded values divided by 8,400 — sixteen percent of the table contributed nothing to the numerator and is absent from the denominator too. That is a sound estimator of the recorded population and a wrong answer to "the average per record", which would need every record to carry a value. The two differ by however much the holes would have pulled the average, and nothing in the output says which one you are holding — which is why the honest report is the number plus the count it was computed over.

go deeper

for a junior

Recall that holes are stepped over rather than counted, so the average describes only the records that reported something. Saying "it divided by the values it had, not by the rows" is the whole junior answer.

for a middle

Do the arithmetic out loud: sum of present values over count of present values, 8,400 rather than 10,000, and name the second average — holes treated as zero — that the tool did not give you.

for a senior

Show how you make this impossible to miss in production: the count travels with the number, completeness is asserted, and a threshold alert is never fed a bare mean from a column you have not audited.

for a principal

Weigh what the organisation reports as standard. Carrying a denominator beside every published metric costs schema and screen space; not carrying it costs an unbounded number of quiet population swaps nobody can audit afterwards.

## What "a number came back" already tells you Before the denominator, notice the fact of the result. An aggregate over a column containing holes has two defensible behaviours: step over them and compute on what remains, or refuse and hand absence back on the grounds that the answer genuinely is unknown. Which one you got is decided by the **skip-absent flag** — the per-call setting deciding whether an aggregate steps over the holes or hands back absence — and by the default its surface's author chose. A number in your hand means that, on this surface, the default was to skip. ## The denominator, in arithmetic Having skipped, the mean divides by what survived: - rows in the table: **10,000** - cells recording nothing: **1,600** - present values: **8,400** - reported mean = (sum of the 8,400 present values) ÷ **8,400** If those present values sum to 840,000, the reported mean is **100**. Divide the same total by the row count and you get **84**. Both are arithmetically fine, and they answer different questions; the output is a bare number that does not say which question it answered. Wherever holes are skipped, the denominator is the count of present values, never the row count. This is not a variation between tools — it is what skipping *means*. A tool that skipped the holes and then divided by the rows would be reporting a total-per-record, not a mean. ## Three averages, and you were handed one of them 1. **The mean of what was recorded.** Denominator: the present values. This is what a skipping mean gives you. It is the right number when a hole means "this record is out of scope", or when you are deliberately describing the population that reported. 2. **The mean per record, treating a hole as a measured zero.** Denominator: the row count, numerator unchanged. This is a different estimator and a skipping mean will never produce it; you get it only by substituting a value first. It is smaller by exactly the share of holes — with sixteen percent absent it is eighty-four percent of the skipping mean. 3. **The mean per record where a hole means "unknown".** Nobody can hand you this from the column alone. The honest outputs are the first number with its count attached, or an estimate that states what it assumed. The distance between (1) and (2) is bounded by the hole fraction, which is exactly why the per-column hole census is worth having before you quote any average. ## Why nothing warns you - Skipping is **silent by design**: no exception, no warning, nothing in the result that marks it. - The result has the same type and the same shape as a mean over a complete column. - Test fixtures usually contain no holes, so the behaviour never shows up in tests. - The consumer downstream — a chart, a tile, a spreadsheet cell — has no channel to carry a denominator even if it knew one. - The number is **plausible**. A mean over eighty-four percent of a population is close enough to the full one to pass a sniff test and far enough away to cross a threshold. ## Making the denominator visible | Practice | What it catches | |---|---| | Report the count of present values beside every mean | the denominator swap, directly and permanently | | Compute both denominators and compare the two answers | how much the holes are actually worth | | Assert a minimum completeness per column before reporting | a column that quietly degraded upstream | | Write the skip-absent flag out where results are reported | a surface whose default runs the other way | None of these is expensive, and the last costs one keystroke while removing the entire class of surprise on that call. ## One caution about generalising The **direction** of the default is a design decision, not a law. Where a surface defaults to skipping you get a number and pass a flag to make absence propagate. Where it defaults to propagating, a single hole makes the whole answer absent until you ask for the holes to be removed — and that is an entire ecosystem's default rather than an obscure corner of one. On a plain rectangle of numbers with no row labels, the ordinary reduction commonly propagates and a separately named absence-aware reduction is the one that skips. What does not vary is the arithmetic. Wherever the holes were skipped, the denominator is the count of the values that were present, and the reader who assumes otherwise is reading a number about a population that was never described to them.

  • How large is the gap between a skipping mean and an average per record that treats holes as zero?
    Exactly the share of present values. Substituting zero keeps the numerator and grows the denominator to the row count, so the second number is the first multiplied by the present fraction — with 1,600 holes in 10,000 rows, eighty-four percent of it. They differ by construction, not by rounding.
  • What would you need to know before deciding which of those averages is the right one?
    What a hole means in this column: not applicable, measured as nothing, or not recorded. The first two license restricting or substituting. For the third the honest output is the recorded-population mean with its count attached, so a reader can see what share of the table it speaks for.

A survey goes to 10,000 people and 1,600 skip the income question. The reported average is the average over the 8,400 who answered. Presenting it as the population's average quietly swaps one population for another, and nothing on the page shows the swap.

saying these in an interview costs you the question

  • Says a skipping mean divides by the number of rows
  • Assumes the result is an average per record
  • Believes the tool would raise if too many values were absent
  • Reports a mean without the count it was computed over
  • Thinks skipping and substituting zero give the same mean