A 10,000-row table has 50 distinct grouping-key values; how many rows come back from a collapse against a summary aligned onto every row?
answer
- three shapes, one split
- the row count is the tell
- fifty groups against ten thousand rows
- what the per-group step hands back
- whole groups, never part of one
basics
~20 sCollapsing to one row per group returns 50 rows, one per distinct key value. Aligning that same summary back onto its rows returns all 10,000, with each row carrying its own group's value beside it.
solid answer
~40 sA grouped operation splits rows by key, computes once per key and reassembles the answers — and the reassembly can take more than one shape. **Collapsing to one row per group** folds every input row away: 50 groups, 50 rows. **Aligning a group summary back onto its rows** keeps the input intact and puts each row's own group value beside it: 10,000 rows, with one value repeated across all the rows of a key. A third shape, a condition evaluated once per group, keeps or drops whole groups, so it returns somewhere between 0 and 10,000 rows — always the sum of the surviving groups' sizes, and never part of a group. Saying the expected row count out loud before running catches most mistakes in this area.
go deeper
Be able to say the row counts out loud: one row per distinct key for a collapse, one row per input row for a summary aligned back. That is the answer a first screen is listening for.
Explain what decides the shape — the value handed back per group — and where the key values land afterwards, which depends on whether the tool carries a row-label slot at all.
Predict the result before running it and check the count you got, especially where a grouped step feeds a report: a per-row result summed as though it were per-group inflates every total by the group sizes.
Fix the convention across the team: which shape a shared transform returns, and whether the grouping key comes back addressable by name. Two teams disagreeing on that produce two subtly different numbers from one source.
## One split, more than one shape of answer A **grouped operation** — split, apply, combine — is a single pass that gives every row a key, runs a computation once per key, and reassembles the answers into a result. The **grouping key** is the column, columns or derived expression whose value decides which rows belong together. Splitting is the easy half and nearly every candidate can describe it. What separates answers is predicting the *result*: how many rows come back, and what one of those rows stands for. Three of the things you can ask for after a split are these, and they differ first of all in row count. ## Three result shapes | What you ask for | Rows in the result (10,000 rows, 50 keys) | What one result row stands for | |---|---|---| | Collapse to one row per group | 50 | a whole group, folded to a value | | Align a group summary back onto its rows | 10,000 | one original row, plus its group's value | | Keep or drop whole groups | 0 to 10,000, in whole groups | an original row whose group passed | - **Collapsing to one row per group** means the result has exactly as many rows as there are groups; every input row has been folded away. Fifty distinct key values give fifty rows, whatever the group sizes were, and whatever the largest group held. - **Aligning a group summary back onto its rows** means the result has the same number of rows as the input, and each row now carries its own group's summary beside it. The value repeats within a group — all 300 rows of one key see that key's single number. That repetition is the point: it puts a group-level quantity next to row-level ones so the two can be compared row by row. - **Keeping or dropping whole groups** applies a condition evaluated once per group: all of that group's rows survive together or leave together. The row count is the sum of the surviving groups' sizes, so it is not a number you can pick freely — a key is present entirely or not at all. These three are not an exhaustive inventory of everything a grouped operation can produce, but they are the three that get confused with one another, and the confusion shows up every time as a wrong row count. ## What decides which one you get The shape follows from what the per-group step hands back, not from which words were typed: 1. Hand back **one value per group** and there is nothing left to put 10,000 rows on, so the result collapses. 2. Hand back **a value for each row of the group** and the result can only be as long as the input, so it aligns back. 3. Hand back **a verdict for the group** — pass or fail — and nothing is computed into the result at all; the original rows are kept or dropped in blocks. ## Where the key values end up After a collapse the key values have to live somewhere, and this is a point where designs genuinely differ. In a design that carries a row-label slot, the key lands there unless you ask otherwise, and a later step that selects by **column name** will not find it. In a design with no row-label concept at all, the key is simply another output column and there is nothing to move. Both are normal. What is not normal is writing the next step as though only one of them exists — if that step addresses the key by name, check that the key is a column in the design actually in front of you. ## A condition once per group is not a condition once per row The third shape is the one that gets mis-named. A condition evaluated **once per group** tests something about the group as a whole — its size, its total, its span — and decides the fate of all its rows at once. A condition evaluated **once per row** tests each record on its own values and can leave a group partly present: some rows of that key gone, others kept. The two read almost identically in English, and in several ecosystems the same call name does both, with the difference being only what it was applied to. The result tells you which one ran: after a group-level condition, every surviving key still has every row it started with. Note also what the group-level form returns. It returns the original rows of the passing groups, unfolded and at full width — not one row per surviving group. Collapsing is a separate operation, and if you want one row per survivor you do both. ## Getting it wrong - Treating a per-row result as if it were one row per group, and inflating a total by the size of each group. - Expecting a collapse to preserve a column the reduction was never asked for — those rows are gone, and nothing can carry an unreduced column through. - Assuming the aligned-back shape must be produced by computing a small summary and matching it back; what it produces is a value per input row, computed from that row's group, and how a given tool gets there is its own business. - Reading 'the group is smaller than I expected' as a bug in the split, when a condition applied to rows upstream removed records before the key was ever assigned.
- How many rows can a condition evaluated once per group return?Anything from zero to the full input, but only in whole groups: the count is the sum of the surviving groups' sizes. If a key is present at all, every one of its rows is present. A result holding only some rows of a key came from a condition evaluated once per row, not a group one.
- In the aligned-back shape, why does the same number appear on many rows?Because the value is computed once for the group and then carried onto each of that group's rows. A key with 300 rows shows its one summary 300 times. That is what makes the shape useful: a group-level quantity now sits beside row-level ones and the two can be compared record by record.
- Does a collapse guarantee the 50 result rows come back in key order?No. Any ordering in the result is a by-product of how the groups were formed: a split that orders the key column leaves an ordering behind, a hash-bucket split leaves none and returns keys in the order they were first seen. If the order matters to what comes next, ask for it rather than inheriting it.
A school register. Collapsing is one line per class showing that class's average mark; aligning back is the full register, every pupil's own mark with their class average beside it; a condition evaluated once per group strikes whole classes off the register and leaves every remaining pupil's line exactly as it was.
saying these in an interview costs you the question
- Says every grouped result has one row per group
- Thinks aligning a summary back changes the row count
- Calls a condition over groups a condition over rows
- Believes a group-level condition can keep part of a group
- Cannot state the result's row count before running the step