A grouped result emits a category that has no rows in it — what do a count, a sum and a mean each return?
answer
- counting has a true answer for nothing
- a fold may or may not return its identity
- no denominator, so no honest mean
- carry the group's size beside the totals
basics
~20 sThe size of the group is zero, and that is the only unambiguous one. A sum over nothing folds to zero in some designs and comes back absent in others, and a mean has no denominator, so it comes back absent rather than zero.
solid answer
~50 sOnce a group with no rows is emitted at all — which needs the key column to declare its allowed values and the tool to emit the unobserved ones — each reduction has to answer for an empty set. **The size of the group is zero**: there are no rows, and that answer is the same everywhere. **A sum is a fork**: some designs fold to the additive identity and return zero, others return an absent marker on the grounds that nothing was summed. **A mean has no denominator at all**, so returning zero would assert a measurement nobody made; designs return an absent marker or signal an error. Extremes behave like the mean — there is no candidate value, so you get an absent marker, an error, or in a few designs the neutral end of the type's range, which is a number nobody observed. The practical upshot: only the group's size distinguishes "no rows here" from "rows here that happened to total zero".
go deeper
Know that a group with no rows has a size of zero, and that a mean cannot be computed over it because there is nothing to divide by.
Sort the reductions into three classes: counting has a true answer, a fold may or may not return its identity depending on the design, and anything that divides or selects has no honest answer.
Demonstrate the consumer-side move: carry the size of the group so a zero in a report can be read, and refuse to fill an absent aggregate with zero unless the fill is a stated decision.
Set the convention once for the team: whether shipped results emit empty groups at all, and what an empty cell is allowed to contain, so two reports built by different people cannot disagree about what zero means.
## Getting to an empty group in the first place Most grouped results never contain an empty group, because a group is formed from rows and no rows means no group. An empty group exists only when something outside the data supplies the key value: the key column carries a **declared set of allowed values** and the tool emits the members of that set that no row used. The question therefore only arises in that setting — but once you have asked for a complete shape, every reduction in the result has to produce something for a set with nothing in it, and that is where the surprises are. ## What each reduction has to do with nothing | Reduction | Result for a group with no rows | Why | |---|---|---| | The size of the group | zero, everywhere | there are no rows, and counting nothing is unambiguous | | The count of present values in a column | zero, everywhere | no values exist to be present | | A sum | zero in some designs, absent in others | a fold over an empty set can return its identity, or decline to answer | | A mean | absent, or an error | the denominator is zero and no honest number exists | | The largest or smallest value | absent, an error, or the neutral end of the range | there is no candidate value to return | | A distinct-value count | zero, everywhere | the set of distinct values is itself empty | The pattern underneath the table is worth naming, because it makes the whole row set memorable: - **Reductions that count** have a true answer for nothing. Zero rows is a fact, not a guess. - **Reductions that fold with an identity** (a sum being the obvious one) *could* return that identity, and designers disagree about whether they should. Returning zero says "the total is nothing"; returning an absent marker says "there was nothing to total". Both are coherent positions, so treat the behaviour as something to check rather than something to know. - **Reductions that must select or divide** (a mean, a largest value, a median) have no honest answer. There is no denominator and no candidate value, so the only truthful outputs are an absent marker or a refusal. A design that returns the neutral end of a numeric range for the largest value of nothing is returning a number that was never in the data, and that number will propagate into anything downstream that reads it. ## Why the zero in a report is the dangerous one A reader looking at a row that says `0` cannot tell, from that cell alone, which of two things happened: 1. Rows existed for that category and their values genuinely added to zero. 2. No rows existed for that category and the reduction folded to its identity. Those two mean completely different things to a business — one is "we did this and it netted out", the other is "we did not do this" — and no amount of staring at the total distinguishes them. The distinguishing column is **the size of the group**, which is zero in the second case and non-zero in the first. This is the concrete reason to carry the group's size beside the aggregates in any result whose empty groups are emitted: it is the only field that separates the two readings, and it costs almost nothing to produce alongside the reduction that was already running. ## A mean over an empty group is not zero This is the single most common wrong answer, and it comes from treating a mean as "a sum with a divide bolted on". Over an empty group there is a sum (possibly), but the divisor is zero, and no convention makes zero over zero into a number. Designs therefore return an absent marker, or they signal an error. If you see zero in that cell, either the tool returned the sum's identity and something downstream rounded the story, or a later step filled the absent marker in with zero — and that later fill is a decision somebody made, not a property of the mean. ## What a good answer sounds like Separate the three classes rather than reciting outcomes: counting reductions have a true answer, folding reductions have a defensible answer that designs disagree about, and selecting or dividing reductions have no answer at all. Then close on the consumer: state that you would carry the size of the group next to the aggregates, so that a zero in the report can be read correctly, and say that you would not fill an absent mean with zero without a stated reason, because that turns "we do not know" into a measurement.
- Why is the size of the group worth carrying beside the aggregates?It is the only field that separates a category whose rows summed to zero from a category that had no rows at all. Both show a total of zero, and they mean opposite things to a reader. The size costs nothing extra to produce while the split is already formed, and it makes every zero in the report self-explaining.
- Is filling an absent mean with zero a reasonable default?Rarely. An absent mean says no measurement exists; zero says a measurement exists and it was zero. Filling converts one into the other silently, and everything downstream — an average of averages, a ranking, a threshold alert — then treats a gap as a real low value. If a fill is needed, it should be an explicit, documented decision made where the report is assembled.
saying these in an interview costs you the question
- Saying a mean over an empty group is zero
- Assuming every design folds an empty sum to zero
- Treating a zero total as proof that rows existed
- Expecting the largest value to fall back to a sensible number
- Filling an absent aggregate with zero without saying so