skip to content

The Long Layout

Each row carries one measurement of one subject at one moment, named by a column rather than by its position, so a new measure arrives as rows instead of a schema change.

on this pageshow

questions

4

A table has columns store, date, measure and amount, with a row reading store 7, 2026-03-01, footfall, 812 — what does one row represent?

level: juniorimportance: must knowfreq 68%

answer

  1. an arrangement, not a size
  2. three roles a column can play
  3. the measure's name sits in a cell
  4. identifiers plus name column key the row
  5. the number is payload, never the key

basics

~20 s

One measurement: store and date say which subject and moment, measure names what was counted, amount carries the number. That is the long layout — one row per measurement, with the measure's name in a cell, not in a header.

solid answer

~40 s

Each row is one measurement of one subject at one moment. `store` and `date` are the **identifying columns**: they say what the row is about, and they are written again on every row describing that store on that day. `measure` is the **name column** — it holds the name of the thing measured, so `footfall` is a data value sitting in a cell rather than a column header. `amount` is the **value column**: one column carries every measurement's number, whatever was measured. Together the identifying columns and the name column key the row, and the value column is the payload that key resolves to. Note what the word long does not mean here: it describes the arrangement, not how many rows the table has.

go deeper

for a junior

Be able to look at five rows and say which columns identify the subject and the moment, which one names the measurement, and which one carries the number.

for a middle

Explain that the identifying columns plus the name column key the row, and that the value column is payload whose meaning comes from its row rather than from its header.

for a senior

Say what the arrangement does downstream: every read has to condition on the name column, and the set of measures present is data that arrived rather than schema that was declared.

for a principal

Weigh what a team gains by never touching the schema when a measure appears against what it pays: identifiers written per measurement, and one representation for every number.

## What the row is telling you A row reading `store 7 · 2026-03-01 · footfall · 812` carries exactly one number and three separate pieces of context for it. Two columns say which subject and which moment the number belongs to, one column says what was measured, and one carries the number itself. That arrangement is **the long layout**: one row per measurement, with the measure's name in one column and its number in another. The first thing to fix is what the word does *not* mean. Long here describes an arrangement, not a size. A table in this arrangement can hold twelve rows, and a table with one column per measure can hold forty million and still be the other arrangement. If you find yourself calling a table long because it is tall, you are using the word for something else. ## The three roles a column plays - **The identifying columns** (`store`, `date`) say which subject and which moment the row is about. They are written again for every measurement, not once per store and day. - **The name column** (`measure`) holds the name of the thing measured. Its distinct values — `footfall`, `revenue`, `returns` — are ordinary data sitting in cells. They are not column headers, and nothing in the table declares which of them ought to be present. - **The value column** (`amount`) holds every measurement's number, whatever was measured: one column for all six or sixty measures. These are roles rather than names. A table may call its columns anything; what makes this the arrangement is that the measure's name is a value and every number shares one column. ## The contrast with one column per measure | | one row per measurement | one column per measure | |---|---|---| | a row is | one measurement of one subject at one moment | one subject at one moment, with all its measures | | the measure's name lives | in a cell of the name column | in a column header | | a new measure means | more rows | a new header, so a change of schema | | identifying columns are written | once per measurement | once per subject and moment | | the numbers sit | in one column, sharing one representation | one column per measure, each with its own | | what a cell means | depends on the row's name column | is fixed by the column it sits in | Both arrangements hold the same measurements. Neither is more correct in the abstract: they differ in what is cheap, what is awkward, and what a reader has to know before touching a number. ## What keys the row The combination that makes a row unique is **the identifying columns together with the name column**. Store 7, 2026-03-01, footfall picks out one measurement, and 812 is what that key resolves to. The value column is payload and is never part of the key. If you ever need the number itself to tell two rows apart, the identifying columns are missing a dimension the data genuinely varies over. That key is also the grain, and the grain is a sentence: *one row is one measure, of one store, on one day*. It is a claim about the data and about the thing that produced it, and it is the sentence an interviewer wants to hear before you touch the table. ## Why the arrangement turns up so often Raw feeds tend to arrive this way, because whatever is producing them emits one reading at a time and does not know what else will ever be measured. A sensor, an event log, a survey response, a billing line: each naturally writes a row that names what it recorded. One column per measure is a decision somebody makes later, while one row per measurement is frequently what was already there. ## What is easy to get wrong 1. **Reading the value column as one quantity.** Its cells share a column, not a unit. A number in it means nothing until you have read its row's name column. 2. **Expecting the set of measures to be declared.** The distinct values of the name column describe what arrived, not what was supposed to arrive. 3. **Confusing the arrangement with the size.** A tall table is tall; the arrangement is about where measure names live. 4. **Putting the value column in the key.** Keying on the number makes every row unique by construction and hides real ambiguity in the identifying columns. In an interview, answer in this order: say what one row is, name the three roles, state the key, then state the consequence — a new measure arrives here as more rows rather than as a change of schema, paid for by writing the identifying columns again for every single measurement.

  • Is a table with forty million rows automatically in this arrangement?
    No. Long and wide describe how the measurements are arranged, not how big the table is. One column per measure, with identifying columns written once per subject and moment, stays that arrangement at any row count. The test is whether a measure's name appears as a value in a cell or as a column header.
  • Can the value column ever be part of what identifies a row?
    No, it is the payload. The row is keyed by the identifying columns together with the name column, and the number is what that key resolves to. If you need the number to tell two rows apart, the identifying columns are missing a dimension the data actually varies over, and keying on the value hides that rather than fixing it.
  • Can a table carry more than one naming column?
    Yes. A measurement can be named by more than one attribute — a measure name together with a currency, an instrument or a sensor. Each extra naming column becomes another part of the key, and every consumer has to condition on all of them rather than only the obvious one.

saying these in an interview costs you the question

  • Calls a table long because it has many rows
  • Reads cells of the value column as comparable to each other
  • Thinks the name column's distinct values are declared somewhere in the table
  • Treats the value column as part of what identifies a row
  • Says a row describes a subject rather than one measurement of it
open as a page

A seventh measure starts arriving for a dataset kept with one row per measurement — what changes, and what does that arrangement cost?

level: middleimportance: must knowfreq 58%

basics

~20 s

No column changes: the new measure arrives as more rows, its name becoming another distinct value of the name column. The price is writing the identifying columns once per measurement and forcing every number into one column.

open as a page

In a table holding one measurement per row, which columns together identify a row, and what breaks when one is left out?

level: middleimportance: should knowfreq 52%

basics

~20 s

The identifying columns together with the name column: every column saying which subject, which moment and which measure. Leave one out and two rows become indistinguishable, so nobody can tell a duplicate from a correction or a second reading.

open as a page

A report averages the amount column of a table holding one row per measurement across six measures and returns a meaningless number — why?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Because a cell's meaning comes from its row's name column, not from the column it sits in: one value column holds six quantities in six units. Any aggregate that does not condition on the name column mixes them.

open as a page