A seventh measure starts arriving for a dataset kept with one row per measurement — what changes, and what does that arrangement cost?
answer
- rows, not columns
- no schema change, three prices
- identifiers written once per measurement
- one column means one representation
- the measure list became data
basics
~20 sNo column changes: the new measure arrives as more rows, its name becoming another distinct value of the name column. The price is writing the identifying columns once per measurement and forcing every number into one column.
solid answer
~40 sNothing about the columns changes. Rows are appended, the new measure's name becomes another distinct value in the name column, and its numbers join all the others in the value column — no new header, no schema change, no reader that has to be told. That is the attraction, and there are three prices for it. The **identifying columns are written once per measurement** rather than once per subject and moment. **Every measurement shares one column**, so all of them have to be held in one representation. And **no measure has a column of its own**, so nothing can be read without conditioning on the name column. There is a fourth, quieter one: the set of measures is now data, so the table describes what arrived rather than what was expected.
go deeper
Remember the headline: a new measurement shows up as more rows with a new name in the name column, and no column is added to the table.
Give the prices as well as the headline — repeated identifiers, one representation for every number, and every read having to name which rows it means.
Point out that variability moved from the schema, where it is visible and gated, into the data, where a measure can appear or vanish with nothing raised.
The standing decision is where you want change to be visible. Accepting a schema change per measure buys review; accepting rows buys speed and costs you a declared set of measures.
## What changes: rows, not columns When the seventh measure starts arriving, the columns stay exactly as they were. Rows are appended; the new measure's name becomes another distinct value in the name column; its numbers go into the value column alongside everything else. No header is added, no existing row is touched, and code that reads the table by column name keeps working unchanged. Compare one column per measure. There, a seventh measure is a seventh header, and everything that touches the table has to learn about it: the writer, the readers, whatever declares the schema, and any storage that wants the column to exist before a row can be written. That difference is the whole reason this arrangement is chosen for feeds whose set of measures is still moving. ## The three prices 1. **The identifying columns are written once per measurement.** A store-day that carries seven measures carries its store and its date seven times. That is structural, not an accident of how the table was built. Whether it costs anything you would notice depends on how many measures there are and on how repeated values are stored, which varies a great deal between designs — so name the repetition, not a multiplier. 2. **Every measurement lands in one column.** One column means one representation for all of them, which is a real constraint on what can be recorded rather than a formality. 3. **No measure has a column of its own.** Every read has to say which rows it means, because naming the value column names no particular quantity. ## Is this arrangement bigger? Only sometimes It is tempting to say the repetition makes it larger, and that is not true in general. One column per measure materialises the complete grid of subject-and-measure combinations, whether or not the data holds them: every subject gets a cell for every measure, including the measures that were never taken for it. One row per measurement holds only the measurements that exist. - On a **dense** set of measurements, where nearly every subject has nearly every measure, one row per measurement is the larger of the two, because the identifiers repeat. - On a **sparse** one, where most subjects have a few of the measures, it is usually the smaller, because it never stores the combinations that did not happen. Say which of those you are on before claiming either is cheaper. ## What one value column requires One column for every measurement means one representation for every measurement. When the incoming measures do not agree — a count, a decimal, a status word — what happens **varies by design**: some tools give you a column with no single type that holds each value as it arrived, some promote numerically and hand back decimals where you had whole counts, and some refuse outright. Establish which behaviour you are dealing with before you assume that putting six unlike measures in one column is free. The important habit is not to carry a rule from one tool to another: the constraint is universal, the resolution is not. ## The set of measures becomes data This is the consequence people miss, and it is the direct result of naming a measurement by a column rather than by its position: - The distinct values of the name column describe **what arrived**, not what should have. - A measure that silently stopped arriving looks exactly like one that was never expected. - A measure whose name changed spelling is, as far as the table is concerned, a new measure and a missing one. - The expected set has to be known independently of the table, because nothing in the arrangement carries it. So the schema stopped changing, but the *shape of the content* still changes — it just does so without anything raising, being deployed, or being reviewed. That is the trade in one sentence: you moved variability out of the schema, where changes are visible and gated, and into the data, where they are neither. ## How to say it in an interview Lead with the headline — *a new measure is rows, not a schema change* — then immediately give the prices, because the headline alone sounds like advocacy. A strong answer names the repetition of the identifiers, the single representation the value column forces, the fact that every read now has to condition on the name column, and the shift of the measure list from schema into data. A weaker one stops at the headline and calls the arrangement flexible, which is true and tells the interviewer nothing about whether you have paid for it.
- What tells you which measures the table is supposed to contain?Nothing in the table. The distinct values of the name column describe what arrived, so a measure that stopped being sent looks identical to one that was never expected. The expected set has to be known independently of the data, and comparing the two is a deliberate act rather than something the arrangement does for you.
- If no column changes, why does a new measure still break consumers sometimes?Because consumers often assume the set of names is fixed even though the table does not. A report that lists one section per measure, a check expecting a known number of rows per subject, or an aggregate over the whole value column all change behaviour when a new name appears, without anyone editing them.
- How much does repeating the identifying columns actually cost?Structurally, the identifier portion grows with the number of measures per subject and moment. What that occupies in practice depends on how many measures there are and how repeated values are stored, which differs widely between designs, so quote the repetition rather than a figure and measure the specific table if the answer matters.
A pre-printed form has a box for each thing it expects; a logbook has a line per entry, each naming what was recorded. Start recording something new and the logbook simply takes another line, while the form has to be reprinted and everyone holding an old one told about it. The logbook is not free, though: the site and the date get written again on every single line, and everything recorded has to fit the one reading box.
saying these in an interview costs you the question
- Says a new measure needs a new column and a migration
- Calls the arrangement free because nothing had to be rewritten
- Assumes six unlike measures share one column without consequence
- Believes the table declares which measures should be present
- Thinks repeating the identifying columns is the same as duplicating rows