An orders table holds each order's item codes in one cell, not one row per code. What does one row then stand for?
answer
- count the rows, not the codes
- one row still means one order
- the group rides inside the cell
- one value per cell no longer holds
basics
~20 sOne row still stands for one order, so the table's row count is the number of orders, not the number of codes. The variable-length group of codes rides inside a single cell of that row.
solid answer
~50 sOne row stands for one order. Because the codes live inside a single cell as a variable-length group, the row count equals the number of orders, while a one-row-per-code layout would carry as many rows as there are codes across all orders. Two consequences follow at once. First, the order's own fields — customer, date, total — are written once per order rather than once per code, so any per-order count or average is honest without de-duplicating anything. Second, the everyday assumption that a cell holds a single value no longer holds for that column: equality against it compares whole groups, a sort order over it is an order over groups, and asking whether a particular code is present is not a plain comparison at all. Whether element-level work is still cheap depends on how the design physically stores the group.
go deeper
Recall the two layouts and their row counts: one row per order with the codes inside one cell, or one row per code with the order's fields repeated. Be able to say which number answers which question.
Explain what the holding buys and charges: the order's fields written once, per-order figures honest without de-duplication, against the loss of the one-value-per-cell assumption that equality, sorting and comparison all rely on.
Show that you would choose the layout from what downstream consumes, and that you know the cost depends on how the design physically stores the group rather than on the idea of a group in a cell.
The angle is what a team pays for mixed conventions: if some tables carry groups in cells and others repeat rows for the same relationship, every consumer has to discover which one it got before it can count anything.
## Two layouts for one-to-many inside a single table An order has many item codes. Inside one in-process table there are exactly two ways to encode that, and they are not different information — they are different addressing. 1. **One row per element.** The table carries one row for every order-and-code pair. Every field belonging to the order — customer, date, total — is written again on each of those rows. The number of rows is the total number of codes. 2. **One cell holding the whole group.** The table carries one row per order, and that order's codes sit inside a single cell of one column as a variable-length group. A column built this way is a **collection-valued column**: a column whose every cell holds a group of values rather than one single value. The number of rows is the number of orders. The question describes the second. So one row still stands for one order, and the table's row count is the order count. ## Put numbers on it Take 500 orders carrying 3,000 item codes between them. | | one row per code | one cell per order | |---|---|---| | row count | 3,000 | 500 | | where the codes sit | 3,000 cells of one column, one value each | 3,000 values spread over 500 cells | | the order's own fields | written 3,000 times | written 500 times | | counting orders | count the distinct order identifiers | count the rows | | counting codes | count the rows | add up the group sizes | | dropping an order | delete a variable number of rows | delete one row | Neither row count is wrong; they answer different questions. The reliable mistake is reading a row count off the second layout and reporting it as a code count, or the reverse. ## What holding the group in the cell buys - **The group stays attached to its row.** A filter, a sort or a selection acts on rows, and whatever happens to the row happens to the whole group. Nothing can keep half an order's codes by accident. - **The order's own fields are written once per order.** Whether that saves much depends on how wide the row is and how large the groups are, but the direction is fixed. - **Per-order figures need no de-duplication.** An average of order totals over this table is an average over orders, because there is exactly one row per order. On the repeated-row layout the same expression weights each order by how many codes it happens to have. - **Variable length costs nothing structural.** An order with one code and an order with forty each occupy one row. There is no fixed set of columns to size for the largest group, and no empty columns for the smallest. - **The row stays the unit of identity.** Anything that addresses rows — a row label, a row-count assertion, a match against another table — keeps meaning one order. ## What it charges - **One value per cell is no longer true**, and a surprising amount of ordinary behaviour assumes it. Equality against that column compares groups. A sort order over it is an order over groups, and what that order even means is a design decision rather than an obvious one. Testing for one code is not a comparison until something reaches inside the cell. - **Element-level work needs a route in.** Either the design's storage lets whole-column operations reach the elements directly, or the group must be turned into one row per element first. - **Two different senses of "the type" are now live.** The column's representation — how the group is physically stored — is one question. The element type inside the cell is a separate one. "The type changed" can mean either, and they fail differently. - **The cost is not uniform across designs.** Some store the column as one packed run of child values plus a run of boundary offsets, which keeps a good deal of whole-column work available. Others store a reference per cell to a separately allocated object, which does not. The holding is the same idea; the bill is not. ## A record held inside a cell is a different shape A cell can also hold a whole named record — several fields at once, such as buyer, channel and region together — rather than a variable-length group of like values. That is a **record-valued cell**, and the comparison is different: its alternative is a set of flat columns on the same row, not a set of repeated rows. It does not change the row count either way. It groups fields that belong together so they travel, are selected and are dropped as a unit. ## What an interviewer is listening for That you answer the row-count question without hesitating, that you can say what one row means in each layout, and that you volunteer the cost rather than only the benefit. A candidate who says "one row per order, 500 rows, and the price is that this column no longer behaves like a column of single values" has shown the whole idea in one sentence.
- If each order instead carried one cell holding a whole named record — buyer, channel and region together — what is different?That is a record-valued cell: a fixed set of named fields rather than a variable-length group. Its alternative is several flat columns on the same row, not repeated rows, so the row count is unaffected either way. The reason to choose it is that the fields belong together and should be selected, moved and dropped as one unit.
- Does the one-cell layout hold less information than one row per code?No. The same order-to-code links are present in both; only the addressing differs. What changes is which work is cheap. Per-order work is direct on the one-cell layout, while per-code work needs either storage that reaches inside the cell or an expansion to one row per element first.
One contact card listing three phone numbers, against three cards that each repeat the same name and address. The same three numbers are recorded either way; the card layout writes the name once and keeps the numbers together, and the price is that 'how many cards' stops answering 'how many numbers'.
saying these in an interview costs you the question
- Says the row count equals the number of item codes
- Assumes every cell of every column still holds exactly one value
- Claims holding the group in a cell loses information the repeated rows keep
- Calls a cell with many values malformed or invalid by definition
- Sorts or compares the column as though it held single values