skip to content

A pipeline turns each element of a variable-length group in a cell into its own row, repeating the rest of that row alongside it. What does that cost, and when is keeping the group better?

level: seniorimportance: should knowfreq 45%

answer

  1. predict the row count before running it
  2. sum of the group sizes
  3. every other field repeats per element
  4. per-entity averages now weight by group size
  5. expand late and narrow

basics

~20 s

Expanding raises the row count to the total element count and repeats every other value in the row once per element, so per-row figures computed afterwards weight each entity by its group size. Keep the group when the work downstream is per-row.

solid answer

~50 s

The row count becomes the sum of the group sizes, and unlike a surprising multiplication this one is predictable: add the sizes before you run it. Every other value in the row is written once per element, so the repeated fields, not the elements, dominate the result's size when the row is wide. The real damage is arithmetic: any average, sum or count over a row-level measure now weights each entity by how many elements it had, and nothing raises an error. Empty groups are the other trap — designs differ on whether such a row disappears or survives with no element, so check rather than assume. Keep the group when downstream work is addressed to the row, when the row is wide relative to the groups, and when the column is in a packed representation that already answers your element questions without expanding.

go deeper

for a junior

Recall what the operation does: one row per element, with the rest of that row copied alongside each one, so the table gets taller and every other column repeats.

for a middle

Explain the predictable arithmetic — the new row count is the sum of the group sizes — and why a per-entity average computed afterwards weights each entity by how many elements it had.

for a senior

Demonstrate the discipline: predict the count, compute entity-level figures before expanding, narrow the columns first, assert the row count at boundaries, and establish what the design does with an empty group.

for a principal

The angle is where the expansion is allowed to live. A convention that keeps it inside one narrow step, rather than at the top of every pipeline, decides whether entity-level numbers can be trusted anywhere downstream.

## What the expansion actually does Turning each element of a group into its own row is a mechanical operation with three effects, and it is worth being able to state all three before running it. 1. **The row count becomes the total element count.** A table of 500 rows whose groups hold 3,000 elements between them becomes 3,000 rows. This is not a surprise to be discovered afterwards — the sum of the group sizes is computable in advance, and a senior candidate computes it. 2. **Every other value in the row is written once per element.** The identifiers, the timestamps, the measures, the text: all repeated. If the row is forty fields wide and the average group holds eight elements, the expansion writes roughly eight times as many values across thirty-nine columns in order to give one column its own rows. 3. **The element becomes an ordinary column value.** After the expansion the elements are single values in a plain column, so everything that expects one value per cell works again — comparison, sort order, matching against another table, grouping by the element. ## The arithmetic damage, which is the part that ships The failure this causes in production is silent. Consider a table with one row per order, an order total, and a group of item codes: - Before expansion, an average of the order total is an average over orders. - After expansion, that same expression averages the total once per item code, so an order with twelve codes counts twelve times and an order with one counts once. Nothing errors. The number looks plausible and is wrong in a direction that tracks group size, which is usually correlated with something interesting — big orders, busy users, popular products — so the bias is systematic rather than random. The defences are ordinary but have to be deliberate: - compute row-level figures **before** expanding, and carry the results forward; - or reduce back to one row per entity before computing them; - or assert the row count at the boundary, so a pipeline stage that silently received an expanded table fails instead of averaging. ## Empty groups and what varies A row whose group holds no elements has no element to put on a row of its own. Designs differ on what they do: some drop the row entirely, some keep one row with no element in that column. Both are defensible and they give different answers, so a pipeline that must not lose entities has to establish which behaviour it is getting rather than assume. This is one of the two places where an expansion changes the set of entities present, the other being an entity whose group was never populated at all. ## Memory, stated with its precondition It is tempting to say the expansion doubles memory. What is actually true depends on the evaluation model: - Under eager evaluation, the source table and the expanded result are both resident at the moment the result is produced, so the peak is roughly the sum of the two — and the expanded one is the larger, driven by the repeated fields rather than by the elements. - Under a model that records the step and evaluates later, nothing is built at the point the expansion is written, and the cost lands wherever the value is finally demanded. - Under a model that shares untouched runs between holdings, the sharing does not help much here specifically, because an expansion touches every column by repeating it. So the honest sentence names the model: in a design that materialises each step, budget for the source plus a result whose size is set by the row width times the total element count. ## When to keep the group instead | keep the group in the cell | expand to one row per element | |---|---| | downstream work is addressed to the entity | downstream work is addressed to the element | | the row is wide relative to the group size | the row is narrow, or the groups are small | | the column is packed, so counts and element filtering already work | the column is a reference per cell, so element work is per-row regardless | | per-entity figures are computed repeatedly | elements are matched or grouped against other tables | | the ordering of the group matters and no position value exists | a position value is carried, or order does not matter | The third row is the one that separates candidates. If the column is stored as a packed run of child values with boundary offsets, element counts, element access and filtering by element are already whole-column operations, and expanding to get them buys nothing but rows. If the column is a reference per cell, that work is proceeding one row at a time anyway, so expanding once is paying a known cost to convert an unknown number of per-row walks into ordinary column work. ## A workable rule Expand as late as possible and as narrowly as possible: 1. Compute and keep whatever entity-level figures you need first. 2. Select down to the columns the element-level work actually requires, so the repetition applies to three fields rather than forty. 3. Expand. 4. Do the element work. 5. Reduce back to the entity if the output is entity-shaped, and assert the row count you expect on the way out. That sequence keeps the expansion's multiplication confined to a narrow table and keeps every per-entity number out of its reach. ## What an interviewer is listening for That you predict the row count instead of discovering it; that you name the silent averaging failure without prompting; that you qualify the memory claim by evaluation model rather than asserting a multiplier; and that you know the expansion is sometimes unnecessary because the storage already supports the work.

  • How would you catch an accidental expansion before it reaches a report?
    Assert the expected row count, or the uniqueness of the entity identifier, at the boundary between stages. An expansion breaks both immediately and loudly, whereas an average computed over expanded rows breaks neither and simply returns a wrong number that nobody questions.
  • The output must be one row per entity again. What is the risk in reducing back?
    Ordering and multiplicity. Unless a position value was carried through, the original order inside each group is not recoverable, and any element duplicated legitimately in the original group is indistinguishable from one introduced along the way. Reducing back reconstructs a set, not necessarily the group you started with.
  • Does expanding make element-level work faster in every case?
    No. It makes it ordinary column work, which is a large win when the column was stored as a reference per cell. Where the column is already packed as child values with boundary offsets, counts and element filtering are whole-column operations already, and the expansion adds rows without adding capability.

saying these in an interview costs you the question

  • Expands by default before checking what downstream needs
  • Computes a per-entity average after expanding without de-duplicating
  • Says the expanded row count cannot be predicted in advance
  • Thinks the elements, not the repeated fields, dominate the result's size
  • Assumes every design keeps a row whose group was empty
  • States a fixed memory multiplier without naming the evaluation model