When does keeping a variable-length group inside one cell beat repeating the rest of the row once per element?
answer
- which unit is the work addressed to
- consumed whole, or per element
- how wide is the rest of the row
- storage decides how much survives unexpanded
basics
~20 sKeep the group in the cell when downstream work is addressed to the row and the group is consumed whole — counted, tested for membership, passed along. Repeat the row when the element is what you filter, match or aggregate on.
solid answer
~50 sDecide by the unit the downstream work is addressed to, not by taste. Holding the group in one cell wins when the group is read as a whole, when the rest of the row is wide so repeating it would multiply far more values than the elements themselves, when per-row figures must stay honest without de-duplicating, and when the group's length varies so much that a fixed set of columns cannot hold it. Repeating the row wins when elements are filtered, matched, grouped or reported on in their own right, and when downstream consumers assume one value per cell. One caveat that decides real cases: how much element-level work survives without expanding depends on the storage. Where the column is a packed run of child values with boundary offsets, counts, element access and filtering by element stay whole-column operations. Where each cell is a separately allocated object, they do not.
go deeper
Recall the two candidate units — the entity row and the element — and that the layout should make the commoner one direct. Being able to name one cost on each side is enough at this stage.
Explain the mechanics: which figures stay honest without de-duplication, how wide rows multiply values under repetition, and why ordering inside the group needs a position value once it becomes rows.
Demonstrate that you would pick the stored form from the consumption pattern, derive the other where needed, and that you would check which storage the column actually landed in before predicting any cost.
The angle is the standing convention: a codebase where the same relationship is held both ways forces every consumer to discover which shape it received, so the rule is worth fixing once and writing down.
## The question is which unit the work is addressed to The choice is not a style preference. It is a bet about what the next ten transformations will ask for. There are two candidate units — the row entity, and the element inside the group — and the holding that makes the common unit direct is the right one. - If almost everything downstream says *per order*, the group belongs inside the order's row. - If almost everything downstream says *per item code*, the row should be repeated once per code. - If the split is genuinely even, hold the group and expand a copy where the element work happens, because expanding is a cheap deterministic operation and re-collapsing is not. ## When holding the group in one cell wins 1. **The group is consumed whole.** Counting its elements, testing whether one is present, passing it along to a consumer that wants the set — all of these want the group, not its parts. 2. **The rest of the row is wide.** Repeating a forty-field row once per element writes forty values per element while the element itself contributes one. The wider the row and the larger the groups, the more lopsided that becomes. 3. **Per-row figures must stay honest.** With one row per entity, a count of rows is a count of entities and an average over a row-level measure is an average over entities. On the repeated layout the same expressions weight each entity by its group size, and the fix is an extra de-duplication step that is easy to forget. 4. **Length varies wildly.** One entity has a single element and another has hundreds. A fixed set of columns would have to be sized for the largest and left empty for the rest; a group in a cell is indifferent. 5. **The group is the thing being described.** Sometimes the ordered set really is the value — a path, a sequence of steps, a basket — and taking it apart loses the ordering or the identity of the whole. ## When repeating the row wins 1. **The element is a first-class unit downstream.** If elements are filtered, matched against another table, grouped or reported on, they need to be rows. 2. **Consumers assume one value per cell.** A great deal of ordinary tooling — comparison, sort order, a straightforward equality test — is defined over single values and either fails or does something unhelpful against a group. 3. **Element-level checks are needed.** Asserting that every element is valid, or that no element is repeated, is direct on rows and roundabout inside cells. 4. **The storage forces per-row work anyway.** If the design landed the column in a reference-per-cell layout, element work is already proceeding one row at a time. Expanding once pays that cost deliberately and buys whole-column work on the elements afterwards. ## The comparison, side by side | | group inside one cell | one row per element | |---|---|---| | row count | one per entity | total element count | | the entity's other fields | written once per entity | written once per element | | per-entity figures | direct | need de-duplication first | | per-element filtering, matching, grouping | needs storage that reaches inside, or an expansion | direct | | variable group length | costs nothing structural | costs nothing structural | | ordering within the group | preserved inside the cell | needs a position column to survive | | assumption of one value per cell | broken for that column | holds everywhere | ## A record held inside a cell is decided differently A cell can also hold a whole named record — several fields at once — rather than a variable-length group. The comparison there is against a set of flat columns on the same row, and the row count is identical either way. Keeping the fields together helps when they are meaningful only as a unit, when several such records of the same shape appear on one row, or when the whole record is selected, moved and dropped together. It hurts when only one of the fields is ever read, because reaching one field of a record in a cell is more work than reading a column. ## What not to assume about cost The most common bad reason for refusing the holding is a claim that a cell holding many values is inherently slow. That is a statement about storage, not about the idea. Designs in this family genuinely differ: - Where the column is held as one packed run of the child values end to end, plus a run of boundary offsets recording where each cell's group starts and stops, element counts and element-level filtering are arithmetic and whole-column scans. Nothing is visited one row at a time. - Where the column falls back to a reference per cell pointing at a separately allocated object, there is nothing packed to work over, so the same operations do proceed one row at a time — and each of those objects carries its own overhead besides. So the honest sentence is conditional: the holding is cheap under one storage and expensive under another, and which you got depends on the design and on how the column was built. ## What an interviewer is listening for A decision rule rather than a preference, a named cost on both sides, and the conditional about storage. Candidates who answer "nesting is bad practice" or "repeating rows wastes space" have both skipped the only question that matters — what the work downstream is addressed to.
- Both layouts are needed by different consumers. What do you do?Hold the group as the stored form and derive the one-row-per-element form where it is needed. Expanding is deterministic and predictable; reconstructing a group from expanded rows requires re-collecting them and recovering the original ordering, which is more work and easier to get subtly wrong.
- If the group's order matters, which layout preserves it more naturally?The group in the cell, because the elements keep their positions within it. Once the group becomes one row per element, the ordering survives only if a position value is carried alongside, since row order in a table is not a guarantee you should lean on.
- When would you reject a record held inside a cell in favour of flat columns?When only one or two of its fields are ever read, when consumers routinely need to select or compare a single field, or when the set of fields is stable and small. Flat columns make each field a direct, individually typed thing; the record is worth it when the fields are meaningful only together.
saying these in an interview costs you the question
- Says holding a group in a cell is always bad practice
- Says repeating the row is always wasteful
- Assumes a cell holding many values is always slow to work with
- Assumes any element-level work requires expanding the column first
- Chooses by appearance rather than by what downstream consumes