In a labelled table typed one column at a time, what is stored per column, and what is a row?
answer
- the holding keeps fields
- one representation per field
- a field's values are kept together
- a record is a cross-section, built on demand
basics
~20 sA labelled table stores fields, not records: each named column is one run of values in a single representation. A row is not stored as a unit — it is a cross-section of every column, assembled on demand.
solid answer
~40 sA **labelled table** — a rectangle whose columns are named and typed one at a time, and whose rows may or may not carry an identity of their own — fixes one **column representation** per field: the single physical form every value in that field is stored in. What gets laid down in memory is therefore the field, one column's values kept together. Nothing lays down a record. Asking for row 500 reads one value out of each field's storage and assembles them into a container, which has to be general enough to hold all those representations at once. So per-column typing is what makes a field cheap and packable, and the record is the awkward unit: it crosses every field's storage and is built fresh each time it is asked for.
go deeper
Recall the two halves: the type is fixed per column, and one column's values are kept together. A record is something the holding builds when asked, not something it has lying ready.
Explain what building a record costs — one read into every field's storage plus a container general enough to hold all of their forms at once — and why that makes the record the awkward unit.
Show where it bites: code that walks records over a wide table pays per field per record, and the repair is to keep the work on whole fields or to convert once at a boundary rather than repeatedly.
The arguable line is where records are allowed. Decide which layers of a codebase may hand records around and which must stay field-oriented, because the assembly cost is paid at every crossing.
## Two facts that get collapsed into one A **labelled table** is a rectangle whose columns are named and typed one at a time, and whose rows may or may not carry an identity of their own. Two separate claims hide inside that description, and keeping them apart is most of this subject. - **The logical fact:** the type is *per field*. Each named column declares one **column representation** — the single physical form every value in that column is stored in. A cell does not carry a type of its own; it is a value already in whatever form its field declared. - **The physical fact:** the values are laid down *field by field*. What is kept together is one column's values, not one record's. Neither implies the other. A record-oriented holding can perfectly well give every field its own type — plenty of structures do. The tables in this family are not that: they are column-major, and the whole cost model follows from it. ## What is actually stored How one field's values are kept varies, and this is the part people assume rather than check. Three holdings are common: 1. **One run per field** — a single allocation holding that field's values back to back. 2. **A chunked column** — the field held as an ordered sequence of separately allocated runs rather than as one buffer. 3. **One slab per representation** — every field sharing a representation packed into a single two-dimensional allocation, so a field is a slice of a slab rather than a thing of its own. All three satisfy "the type is per field". None of them stores a record. ## A row is a cross-section, not a stored thing | | a field | a record | |---|---|---| | stored as a unit? | yes — a run, a sequence of runs, or a slice of a slab | no | | how many representations | one, fixed for the whole field | as many as the fields it crosses | | getting one | hand back the run or the slice | read every field's storage and build a container | | cost as the table widens | unchanged | grows with the number of fields | When you ask for row 500, nothing is looked up in a record store, because there is none. One value is read out of each field's storage at that offset and the results are assembled. Two consequences follow immediately: - the assembled container has to hold several representations at once, so it is a general-purpose holding that stores a reference per value rather than a packed run — **materialising a record throws away the packing that made the fields cheap**; - the work scales with the **number of fields**, not with the size of the table. A record out of a five-field table is five reads; out of a four-hundred-field table it is four hundred. The container is also not kept. Unless you hold on to it, it is built and discarded, and the next request for the same row builds it again. ## What the shape buys, and what it charges - **Cheap:** narrow storage, because each field picks a form that fits only its own values instead of one form covering everything in the table. - **Cheap:** saying something once for a whole field, because that field's values are already together and already in one form. - **Cheap, as a direction:** adding a field, which is one more run or one more slice and does not read the existing fields. - **Expensive:** anything that wants records, because each one crosses every field's storage and is not reusable afterwards. - **Expensive, in a degree that depends on the holding:** appending records, which has to place a value into every field's storage. ## Where designs disagree Flat statements are dangerous here, because this family genuinely diverges. - "A field is one contiguous buffer" is true of the one-run holding only. A chunked field has no single buffer to point at, and a slice of a slab has one but does not own it. - "Adding a field is cheap" holds as a *direction* in every column-major holding, but the *degree* ranges from one extra allocation to rebuilding the whole slab that every same-representation field lives in. - Whether a record is even nameable differs: designs that carry row identity can name one by its row label, while designs without it can only name a position. What does not vary is the shape of the answer: the holding keeps fields, and a record is work it does when asked. ## Reading a cost surprise When something in a table surprises you, ask which of the two facts it came from. "The values came back in the wrong form" is the logical fact — the representation. "This got slower as we widened the table" or "walking records takes minutes" is the physical one. They have different fixes, and mistaking one for the other produces changes that do not help.
- If a record is not stored, what exactly does the holding hand you when you ask for one?A freshly built container holding one value taken from each field's storage. Because those fields may be in different forms, the container has to be general enough to hold all of them at once, so it is not a packed run the way a field is. It is also discarded unless you keep it, so asking again rebuilds it.
- Does per-column typing mean two columns can never share one allocation?No. Some designs pack every field sharing a representation into a single two-dimensional slab, so a field is a slice of a shared allocation rather than an allocation of its own. The logical picture — one form per named field — is unchanged; only who owns the memory differs, and with it the cost of adding another field of that same form.
- Why does widening a table hurt record-at-a-time work more than it hurts field-at-a-time work?Field work touches one field's storage whatever the table's width. Building a record touches every field, so its cost tracks the field count. Going from five fields to four hundred leaves a field operation unchanged and makes each record eighty times more reads, plus a larger container to build and throw away.
A filing cabinet with one drawer per field: all the surnames in one drawer, all the start dates in another. Nobody's complete file exists as a folder. You make one by pulling a single sheet from every drawer, and the more drawers there are, the longer that takes.
saying these in an interview costs you the question
- Describes a table as a collection of record objects, one per row
- Thinks each cell carries its own type, as in a spreadsheet
- Assumes a record sits together in memory and is handed over as-is
- Says every design keeps a field as one contiguous run of values
- Treats per-column typing as a validation rule rather than a storage decision