skip to content

Split, Apply, Combine

Splitting rows by key, doing work per group and putting the answers back. Naming the phases is easy; the output's row count, its labels and its missing groups are where people slip.

on this pageshow

questions

25

A 10,000-row table has 50 distinct grouping-key values; how many rows come back from a collapse against a summary aligned onto every row?

level: juniorimportance: must knowfreq 72%

answer

  1. three shapes, one split
  2. the row count is the tell
  3. fifty groups against ten thousand rows
  4. what the per-group step hands back
  5. whole groups, never part of one

basics

~20 s

Collapsing to one row per group returns 50 rows, one per distinct key value. Aligning that same summary back onto its rows returns all 10,000, with each row carrying its own group's value beside it.

solid answer

~40 s

A grouped operation splits rows by key, computes once per key and reassembles the answers — and the reassembly can take more than one shape. **Collapsing to one row per group** folds every input row away: 50 groups, 50 rows. **Aligning a group summary back onto its rows** keeps the input intact and puts each row's own group value beside it: 10,000 rows, with one value repeated across all the rows of a key. A third shape, a condition evaluated once per group, keeps or drops whole groups, so it returns somewhere between 0 and 10,000 rows — always the sum of the surviving groups' sizes, and never part of a group. Saying the expected row count out loud before running catches most mistakes in this area.

go deeper

for a junior

Be able to say the row counts out loud: one row per distinct key for a collapse, one row per input row for a summary aligned back. That is the answer a first screen is listening for.

for a middle

Explain what decides the shape — the value handed back per group — and where the key values land afterwards, which depends on whether the tool carries a row-label slot at all.

for a senior

Predict the result before running it and check the count you got, especially where a grouped step feeds a report: a per-row result summed as though it were per-group inflates every total by the group sizes.

for a principal

Fix the convention across the team: which shape a shared transform returns, and whether the grouping key comes back addressable by name. Two teams disagreeing on that produce two subtly different numbers from one source.

## One split, more than one shape of answer A **grouped operation** — split, apply, combine — is a single pass that gives every row a key, runs a computation once per key, and reassembles the answers into a result. The **grouping key** is the column, columns or derived expression whose value decides which rows belong together. Splitting is the easy half and nearly every candidate can describe it. What separates answers is predicting the *result*: how many rows come back, and what one of those rows stands for. Three of the things you can ask for after a split are these, and they differ first of all in row count. ## Three result shapes | What you ask for | Rows in the result (10,000 rows, 50 keys) | What one result row stands for | |---|---|---| | Collapse to one row per group | 50 | a whole group, folded to a value | | Align a group summary back onto its rows | 10,000 | one original row, plus its group's value | | Keep or drop whole groups | 0 to 10,000, in whole groups | an original row whose group passed | - **Collapsing to one row per group** means the result has exactly as many rows as there are groups; every input row has been folded away. Fifty distinct key values give fifty rows, whatever the group sizes were, and whatever the largest group held. - **Aligning a group summary back onto its rows** means the result has the same number of rows as the input, and each row now carries its own group's summary beside it. The value repeats within a group — all 300 rows of one key see that key's single number. That repetition is the point: it puts a group-level quantity next to row-level ones so the two can be compared row by row. - **Keeping or dropping whole groups** applies a condition evaluated once per group: all of that group's rows survive together or leave together. The row count is the sum of the surviving groups' sizes, so it is not a number you can pick freely — a key is present entirely or not at all. These three are not an exhaustive inventory of everything a grouped operation can produce, but they are the three that get confused with one another, and the confusion shows up every time as a wrong row count. ## What decides which one you get The shape follows from what the per-group step hands back, not from which words were typed: 1. Hand back **one value per group** and there is nothing left to put 10,000 rows on, so the result collapses. 2. Hand back **a value for each row of the group** and the result can only be as long as the input, so it aligns back. 3. Hand back **a verdict for the group** — pass or fail — and nothing is computed into the result at all; the original rows are kept or dropped in blocks. ## Where the key values end up After a collapse the key values have to live somewhere, and this is a point where designs genuinely differ. In a design that carries a row-label slot, the key lands there unless you ask otherwise, and a later step that selects by **column name** will not find it. In a design with no row-label concept at all, the key is simply another output column and there is nothing to move. Both are normal. What is not normal is writing the next step as though only one of them exists — if that step addresses the key by name, check that the key is a column in the design actually in front of you. ## A condition once per group is not a condition once per row The third shape is the one that gets mis-named. A condition evaluated **once per group** tests something about the group as a whole — its size, its total, its span — and decides the fate of all its rows at once. A condition evaluated **once per row** tests each record on its own values and can leave a group partly present: some rows of that key gone, others kept. The two read almost identically in English, and in several ecosystems the same call name does both, with the difference being only what it was applied to. The result tells you which one ran: after a group-level condition, every surviving key still has every row it started with. Note also what the group-level form returns. It returns the original rows of the passing groups, unfolded and at full width — not one row per surviving group. Collapsing is a separate operation, and if you want one row per survivor you do both. ## Getting it wrong - Treating a per-row result as if it were one row per group, and inflating a total by the size of each group. - Expecting a collapse to preserve a column the reduction was never asked for — those rows are gone, and nothing can carry an unreduced column through. - Assuming the aligned-back shape must be produced by computing a small summary and matching it back; what it produces is a value per input row, computed from that row's group, and how a given tool gets there is its own business. - Reading 'the group is smaller than I expected' as a bug in the split, when a condition applied to rows upstream removed records before the key was ever assigned.

  • How many rows can a condition evaluated once per group return?
    Anything from zero to the full input, but only in whole groups: the count is the sum of the surviving groups' sizes. If a key is present at all, every one of its rows is present. A result holding only some rows of a key came from a condition evaluated once per row, not a group one.
  • In the aligned-back shape, why does the same number appear on many rows?
    Because the value is computed once for the group and then carried onto each of that group's rows. A key with 300 rows shows its one summary 300 times. That is what makes the shape useful: a group-level quantity now sits beside row-level ones and the two can be compared record by record.
  • Does a collapse guarantee the 50 result rows come back in key order?
    No. Any ordering in the result is a by-product of how the groups were formed: a split that orders the key column leaves an ordering behind, a hash-bucket split leaves none and returns keys in the order they were first seen. If the order matters to what comes next, ask for it rather than inheriting it.

A school register. Collapsing is one line per class showing that class's average mark; aligning back is the full register, every pupil's own mark with their class average beside it; a condition evaluated once per group strikes whole classes off the register and leaves every remaining pupil's line exactly as it was.

saying these in an interview costs you the question

  • Says every grouped result has one row per group
  • Thinks aligning a summary back changes the row count
  • Calls a condition over groups a condition over rows
  • Believes a group-level condition can keep part of a group
  • Cannot state the result's row count before running the step
open as a page

Computing the largest order value per customer over a long table, what does the run keep for each customer as it reads?

level: juniorimportance: must knowfreq 58%

basics

~20 s

One running value per customer — the largest amount seen so far, replaced when a bigger one arrives. Each record is folded in and then dropped, so the room used tracks the number of customers, not the number of orders.

open as a page

After grouping orders by both region and payment method, what identifies each row of the collapsed result, and how many rows are there?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Each result row is identified by a pair, one value from each grouping key, and the result holds exactly one row per pair that actually occurs in the data, not one row per pair the two columns could form.

open as a page

A grouped count over a table of orders returns 11 rows though the business defines 14 categories — why?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A grouped result carries one row per key value that actually occurred, so the three unlisted categories had no rows in this input. They are absent from the result rather than present with a zero.

open as a page

A grouped aggregation over a table is described as three phases — what happens in each?

level: juniorimportance: must knowfreq 70%

basics

~10 s

Split, apply, combine. The split gives every row a grouping key and records which rows share each key; the apply runs one computation per key; the combine assembles those per-key answers into a result.

open as a page

A sales table is narrowed to customers whose total exceeds 1,000; how does that differ from keeping only the individual orders above 1,000?

level: middleimportance: must knowfreq 68%

basics

~20 s

The customer total is a condition evaluated once per group: every order of a qualifying customer survives and none of a failing one. The order amount is a condition evaluated once per row, which can leave a customer partly present.

open as a page

Which grouped computations finish with one fixed-size running value per group, and which need the group's values available at once?

level: middleimportance: must knowfreq 55%

basics

~20 s

A computation folds when a fixed-size state plus one more record yields the new state — a group's total, its row count, an extreme. It cannot fold when the answer turns on the group's whole distribution, like an exact middle value.

open as a page

Rows in each group are numbered by descending amount and the top three kept — how does the tie disposition change which rows survive?

level: middleimportance: must knowfreq 72%

basics

~20 s

It decides both how many rows survive and which ones. Distinct consecutive numbers keep exactly three but pick arbitrarily among tied rows; shared numbers keep every row tied at the cut, so a group can return four or more.

open as a page

A table of 50,000 rows is grouped by one key column and the group sizes total 48,600 — where did the rest go?

level: middleimportance: must knowfreq 64%

basics

~20 s

Those 1,400 rows carry no value in the grouping key. A tool either gathers such rows into one group of their own or leaves them out of the computation entirely; here it left them out, so the group sizes no longer reconcile with the table.

open as a page

Grouping 50 million rows by a key returns immediately — what has actually been built at that point?

level: middleimportance: must knowfreq 55%

basics

~20 s

Normally just bookkeeping: a record of which row positions carry which key value. No per-key table has been built and no computation has run. Rows are copied only when something later demands a materialised group.

open as a page

Orders are grouped by customer and each row numbered inside its group by date; what does a five-order customer receive?

level: juniorimportance: should knowfreq 60%

basics

~20 s

Five numbers, one per row, running 1 to 5 and restarting at every customer. A position inside a group is a per-row value, not a single value for the group, and the chosen ordering decides which row is 1.

open as a page

When a grouped step must have its group's values at once, is every column of those rows held or only the ones read?

level: middleimportance: should knowfreq 38%

basics

~20 s

It depends on what the surface hands the step. Some pass only the values of the column being reduced; some materialise a table of the group's rows across every column, including ones the step never reads — and then the group's width is part of the bill.

open as a page

Swapping the order of two grouping keys changes what about the collapsed result, and what does it leave untouched?

level: middleimportance: should knowfreq 46%

basics

~20 s

Key order changes the layout - which key is the outer part of the identity, and so how the result is arranged and reads - but not a single aggregated value, because the set of pairs is identical either way.

open as a page

Where do the two grouping-key values end up after a collapse, and can a later step select them by column name?

level: middleimportance: should knowfreq 54%

basics

~20 s

Either into the result's row-label slot as one label made of two parts, or into two ordinary columns; which happens is a property of the tool. Only the column form is visible to a step that selects by name.

open as a page

A running total is carried down each group's rows under a chosen ordering — how does it differ from the group's total?

level: middleimportance: should knowfreq 57%

basics

~20 s

A running total is one value per row — that row plus every earlier row of the same group under the chosen ordering. The group's total is one value for the whole group, and it is what the group's last row already holds.

open as a page

A grouped result emits a category that has no rows in it — what do a count, a sum and a mean each return?

level: middleimportance: should knowfreq 44%

basics

~20 s

The size of the group is zero, and that is the only unambiguous one. A sum over nothing folds to zero in some designs and comes back absent in others, and a mean has no denominator, so it comes back absent rather than zero.

open as a page

A grouped total returns its keys in no obvious order — what does that say about the split?

level: middleimportance: should knowfreq 50%

basics

~20 s

That the split formed groups by hashing key values into buckets rather than by ordering them. Any ordering in a grouped result is a by-product of the split mechanism, not a guarantee — order the result explicitly if you need it.

open as a page

A per-region average changed after a row condition was added upstream; how does that differ from dropping whole regions with a group-level condition?

level: seniorimportance: should knowfreq 51%

basics

~20 s

A row condition upstream changes which rows enter each group, so every surviving region's average is recomputed over fewer rows. A group-level condition removes regions outright and leaves every remaining average exactly as it was.

open as a page

In a grouped operation, how does what a hand-written per-group body returns decide whether the result has one row per group or one row per input row?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The shape of the return decides it. One value per group collapses to one row per group; a value for each row of the group comes back aligned onto those rows; a pass-or-fail verdict keeps or drops the group whole.

open as a page

A grouped run over a 4 GB table exhausts a 32 GB machine's memory. What do you check about the groups first?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The distribution of group sizes, and specifically its maximum. When the per-key step needs its group present, the peak follows the largest group rather than the table, so one key holding most of the rows can exhaust a machine many times the input's size.

open as a page

A per-group numbering returned different rows this month with no code change — what should you check about its ordering?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Check whether the numbering states its own ordering or inherits the row order it happens to find. An inherited ordering is owned by whatever step ran before it, so an unrelated upstream change silently renumbers every group.

open as a page

A monthly report grouped on two key columns lost one of its rows this month though nobody changed the code — what happened?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Group identity is the pair of key values, and only pairs that actually occurred get a row. A pair that had rows last month had none this month, so it produced nothing at all — not a zero, nothing — and the result came back one row narrower.

open as a page

One customer appears as two rows in a grouped report — what does the split compare to decide group identity?

level: seniorimportance: should knowfreq 58%

basics

~20 s

The value the key expression produced for each row, compared for equality — not how that value prints. A trailing space, a difference in case, or a timestamp carrying a time of day all produce distinct key values from rows you consider identical.

open as a page

A grouped total over 80 million rows with 50 million distinct keys is slow — which phase is costing you?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The split. With a near-unique key, organising tens of millions of key values dominates: each apply touches only one or two rows, and the combine emits almost as many result rows as there were inputs.

open as a page

A monthly report groups spend by region and by a size band derived from the amount column; what must be true of that derivation for two months' reports to be comparable?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

The derivation must be deterministic per record, against band edges fixed outside the run. Edges computed from each month's own data make the bands mean different things each time, so one band label no longer denotes the same records.

open as a page