In a grouped operation, how does what a hand-written per-group body returns decide whether the result has one row per group or one row per input row?
answer
- the return shape is the contract
- one value, many values, or a verdict
- built-in reductions declare their shape
- a shape that moves with the data
basics
~20 sThe shape of the return decides it. One value per group collapses to one row per group; a value for each row of the group comes back aligned onto those rows; a pass-or-fail verdict keeps or drops the group whole.
solid answer
~50 sA reduction the tool implements itself declares its own result shape, so the row count is known before anything runs. A **hand-written per-group body** — a function you supply that the library cannot look inside — declares nothing, and the result's shape is inferred from what the body actually handed back for each group. Return a single value and the result collapses to one row per group; return as many values as the group had rows and the result comes back aligned onto those rows; return a verdict and the group is kept or dropped whole. The hazard is a body whose return shape depends on the data — one value for groups of one, several for larger groups. Nothing errors. The row count simply becomes a function of the input, and it moves the first month the input changes character.
code
pseudocode · 16 lines# a grouped operation hands each group's rows to a body you wrote,
# and the SHAPE of what the body returns decides the result's shape
body_one(group_rows) -> one value # result: one row per group
body_many(group_rows) -> value per group row # result: one row per input row
body_test(group_rows) -> true or false # result: group kept, or dropped whole
# the unstable case: the shape depends on the data
function body_unstable(group_rows):
if count(group_rows) == 1:
return single_value_of(group_rows) # one value
else:
return two_largest(group_rows) # two values
# singleton groups contribute one row, every other group contributes two,
# so the result's row count now moves with the input and nothing reports itgo deeper
Know that a function you supply per group returns something, and that what it returns is what appears in the result: a single value per group is not the same as a value for every row.
Explain the three return shapes and the result each one produces, and name the expected row count for each before the step runs.
Spot the body whose return shape depends on the data, say why it passes every test and then moves in production, and keep every branch of a body returning one shape.
Decide where hand-written bodies are allowed at all. Generality bought at the price of a data-dependent result shape is a standing cost, and it is paid by whoever debugs the row count a quarter later.
## The return value is the contract A **grouped operation** gives every row a key, runs a computation once per key, and reassembles the answers. When the computation is one the tool implements itself, its result shape is fixed by definition: a reduction over a group produces one value, so the result has one row per group, and you can state the row count before anything runs. When the computation is **a hand-written per-group body** — a function you supply that the library cannot look inside — that guarantee is gone. The tool sees only what came back, and assembles a result to fit it. ## Three returns, three results | What the body returns for one group | What the result looks like | Row count | |---|---|---| | A single value | one row per group | the number of distinct keys | | One value for each row of the group | a value on each input row | unchanged from the input | | A verdict, pass or fail | the group's original rows, or none of them | sum of the surviving groups' sizes | Read the table the other way round and it becomes a design tool: decide the result you want, then write a body that returns that shape. A per-row share of a group total wants the middle line; a per-group headline number wants the first; a cohort cut wants the third. ## What varies between tools, and what does not Two things here are genuinely different between designs, and a candidate who states either as universal is wrong somewhere: - **What the body is handed.** Some surfaces call the body once per record; others hand it a whole group's column in one call, sometimes under a very similar name. That distinction — not the fact that the body was hand-written — is what decides whether a crossing out of the library's own execution is paid for every record or once per group. Reason about what the callback *receives*, never about who wrote it. - **Whether grouping persists.** In some designs grouping is a state the table carries until it is explicitly dropped, and a later step still operates per group; in others the split produces a transient handle that the very next call consumes, after which the result is an ordinary table with no memory of the split. The first model carries a hazard the second does not: a computation written as though the table were ungrouped, and quietly evaluated once per group instead. What does not vary is the rule at the top. Whatever the body was handed and however long the grouping lives, the shape it hands back determines the shape of the assembled result. ## When the shape is not constant The failure worth interviewing on is a body whose return shape depends on the group's contents. 1. A body returns the group's mean when the group has one row, and its two largest values otherwise. Singleton groups contribute one row; every other group contributes two. 2. A body returns a value per row in the common case, but a single value on an early-exit path taken for degenerate input. 3. A body returns a sequence per group, but a *shorter* one than it was given — trimmed or de-duplicated — so the result is neither the collapsed shape nor the aligned one, and matches nothing downstream expects. All three run without complaint. The result's row count becomes a function of the data, which means it is right today, right in the test fixture, and different next month when the first singleton group appears. No code changed, and nothing in the change log explains it. ## Making it predictable - Write the body so that every path returns the same shape. If two shapes are genuinely wanted, that is two operations and should be written as two. - Decide up front which of the three results you are after and say the expected row count out loud — the number of distinct keys, or the input row count — before running the step. - Prefer the tool's own reduction wherever one exists for what you need. Its shape is part of its definition and cannot drift with the data, and the tool can usually execute it without the generality you are not using. - Keep the early-exit paths in view. Degenerate groups are exactly where an inconsistent return sneaks in, and they are also the groups nobody inspects. - When a body must be hand-written, record what one group's return looks like beside it. The next reader cannot recover the intended shape from a result whose shape depends on the input. ## Why this is a senior question A hand-written body is the thing people reach for because it always works: anything expressible can be expressed inside it. The senior observation is what that generality costs. The result's shape stops being a property of the code and becomes a property of the data, while every step downstream was written against a shape somebody assumed once and never wrote down.
- A per-group body returns two values for most groups and one for singletons. What breaks, and when?Nothing breaks immediately. The row count becomes twice the number of groups minus the singletons, so it is stable only while the data is. The first day a key arrives with a single row, every downstream count moves, with no code change to blame it on.
- Why does a tool's own reduction give a row count you can state in advance?Because its result shape is part of its definition: one value per group, on every path, whatever the group holds. There is no branch that could return something else, so the result has one row per distinct key regardless of the data.
- Before running it, how do you know which row count to expect from a hand-written body?Read what the body returns for a single group, on every path through it. One value means one row per distinct key; a value for each row means the input's row count; a verdict means the surviving groups' rows. If the paths disagree, you cannot state a count, and that is the finding.
saying these in an interview costs you the question
- Says a hand-written per-group body always collapses the group
- Assumes every per-group body is invoked once per record
- Returns a different shape on different branches of one body
- Cannot state the expected row count before running the step
- Believes the surface name alone fixes the result shape