skip to content

Long and Wide

Two layouts for the same numbers - one row per measurement, or one column per measure - the operations that convert between them, and what each conversion quietly loses.

on this pageshow

questions

20

A table has columns store, date, measure and amount, with a row reading store 7, 2026-03-01, footfall, 812 — what does one row represent?

level: juniorimportance: must knowfreq 68%

answer

  1. an arrangement, not a size
  2. three roles a column can play
  3. the measure's name sits in a cell
  4. identifiers plus name column key the row
  5. the number is payload, never the key

basics

~20 s

One measurement: store and date say which subject and moment, measure names what was counted, amount carries the number. That is the long layout — one row per measurement, with the measure's name in a cell, not in a header.

solid answer

~40 s

Each row is one measurement of one subject at one moment. `store` and `date` are the **identifying columns**: they say what the row is about, and they are written again on every row describing that store on that day. `measure` is the **name column** — it holds the name of the thing measured, so `footfall` is a data value sitting in a cell rather than a column header. `amount` is the **value column**: one column carries every measurement's number, whatever was measured. Together the identifying columns and the name column key the row, and the value column is the payload that key resolves to. Note what the word long does not mean here: it describes the arrangement, not how many rows the table has.

go deeper

for a junior

Be able to look at five rows and say which columns identify the subject and the moment, which one names the measurement, and which one carries the number.

for a middle

Explain that the identifying columns plus the name column key the row, and that the value column is payload whose meaning comes from its row rather than from its header.

for a senior

Say what the arrangement does downstream: every read has to condition on the name column, and the set of measures present is data that arrived rather than schema that was declared.

for a principal

Weigh what a team gains by never touching the schema when a measure appears against what it pays: identifiers written per measurement, and one representation for every number.

## What the row is telling you A row reading `store 7 · 2026-03-01 · footfall · 812` carries exactly one number and three separate pieces of context for it. Two columns say which subject and which moment the number belongs to, one column says what was measured, and one carries the number itself. That arrangement is **the long layout**: one row per measurement, with the measure's name in one column and its number in another. The first thing to fix is what the word does *not* mean. Long here describes an arrangement, not a size. A table in this arrangement can hold twelve rows, and a table with one column per measure can hold forty million and still be the other arrangement. If you find yourself calling a table long because it is tall, you are using the word for something else. ## The three roles a column plays - **The identifying columns** (`store`, `date`) say which subject and which moment the row is about. They are written again for every measurement, not once per store and day. - **The name column** (`measure`) holds the name of the thing measured. Its distinct values — `footfall`, `revenue`, `returns` — are ordinary data sitting in cells. They are not column headers, and nothing in the table declares which of them ought to be present. - **The value column** (`amount`) holds every measurement's number, whatever was measured: one column for all six or sixty measures. These are roles rather than names. A table may call its columns anything; what makes this the arrangement is that the measure's name is a value and every number shares one column. ## The contrast with one column per measure | | one row per measurement | one column per measure | |---|---|---| | a row is | one measurement of one subject at one moment | one subject at one moment, with all its measures | | the measure's name lives | in a cell of the name column | in a column header | | a new measure means | more rows | a new header, so a change of schema | | identifying columns are written | once per measurement | once per subject and moment | | the numbers sit | in one column, sharing one representation | one column per measure, each with its own | | what a cell means | depends on the row's name column | is fixed by the column it sits in | Both arrangements hold the same measurements. Neither is more correct in the abstract: they differ in what is cheap, what is awkward, and what a reader has to know before touching a number. ## What keys the row The combination that makes a row unique is **the identifying columns together with the name column**. Store 7, 2026-03-01, footfall picks out one measurement, and 812 is what that key resolves to. The value column is payload and is never part of the key. If you ever need the number itself to tell two rows apart, the identifying columns are missing a dimension the data genuinely varies over. That key is also the grain, and the grain is a sentence: *one row is one measure, of one store, on one day*. It is a claim about the data and about the thing that produced it, and it is the sentence an interviewer wants to hear before you touch the table. ## Why the arrangement turns up so often Raw feeds tend to arrive this way, because whatever is producing them emits one reading at a time and does not know what else will ever be measured. A sensor, an event log, a survey response, a billing line: each naturally writes a row that names what it recorded. One column per measure is a decision somebody makes later, while one row per measurement is frequently what was already there. ## What is easy to get wrong 1. **Reading the value column as one quantity.** Its cells share a column, not a unit. A number in it means nothing until you have read its row's name column. 2. **Expecting the set of measures to be declared.** The distinct values of the name column describe what arrived, not what was supposed to arrive. 3. **Confusing the arrangement with the size.** A tall table is tall; the arrangement is about where measure names live. 4. **Putting the value column in the key.** Keying on the number makes every row unique by construction and hides real ambiguity in the identifying columns. In an interview, answer in this order: say what one row is, name the three roles, state the key, then state the consequence — a new measure arrives here as more rows rather than as a change of schema, paid for by writing the identifying columns again for every single measurement.

  • Is a table with forty million rows automatically in this arrangement?
    No. Long and wide describe how the measurements are arranged, not how big the table is. One column per measure, with identifying columns written once per subject and moment, stays that arrangement at any row count. The test is whether a measure's name appears as a value in a cell or as a column header.
  • Can the value column ever be part of what identifies a row?
    No, it is the payload. The row is keyed by the identifying columns together with the name column, and the number is what that key resolves to. If you need the number to tell two rows apart, the identifying columns are missing a dimension the data actually varies over, and keying on the value hides that rather than fixing it.
  • Can a table carry more than one naming column?
    Yes. A measurement can be named by more than one attribute — a measure name together with a currency, an instrument or a sensor. Each extra naming column becomes another part of the key, and every consumer has to condition on all of them rather than only the obvious one.

saying these in an interview costs you the question

  • Calls a table long because it has many rows
  • Reads cells of the value column as comparable to each other
  • Thinks the name column's distinct values are declared somewhere in the table
  • Treats the value column as part of what identifies a row
  • Says a row describes a subject rather than one measurement of it
open as a page

What must be true of a table before one column's distinct values can safely become new headers filled from a second column?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Each combination of row key and new header must occur exactly once. The result has one slot per combination, so two input rows sharing the same identifying values and the same header value leave one cell with two candidate contents.

open as a page

When a table is widened into one column per measure using two name columns instead of one, what do the resulting column labels look like?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Each output column is identified by two name parts, one from each name column, with one column per pair the data holds. Whether those parts stay separately addressable or are composed into a single string depends on the tool's label space.

open as a page

A table with one row per measurement is widened to one column per measure. Which three roles must its columns be assigned?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Three roles: the identifying columns that stay put as the row key, the header source whose distinct values become the new column headers, and the cell source whose entries fill those cells. Folding back reads the same declaration backwards.

open as a page

One row per measurement, or one column per measure — which layout does comparing two measures against each other read more naturally?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The wide layout — one column per measure — puts both measures on the same row, so the comparison is one column-against-column step. In the long layout, one row per measurement, the two numbers sit in different rows and must be brought together first.

open as a page

A seventh measure starts arriving for a dataset kept with one row per measurement — what changes, and what does that arrangement cost?

level: middleimportance: must knowfreq 58%

basics

~20 s

No column changes: the new measure arrives as more rows, its name becoming another distinct value of the name column. The price is writing the identifying columns once per measurement and forcing every number into one column.

open as a page

A widening finds two rows supplying the same cell — what can a tool do about it, and which choice actually tells you?

level: middleimportance: must knowfreq 55%

basics

~10 s

Three kinds of disposition: refuse, because one value per cell is a stated precondition; resolve the cell with a default reduction; or keep both values in the cell. Only the refusal tells you.

open as a page

In a table holding one measurement per row, which columns together identify a row, and what breaks when one is left out?

level: middleimportance: should knowfreq 52%

basics

~20 s

The identifying columns together with the name column: every column saying which subject, which moment and which measure. Leave one out and two rows become indistinguishable, so nobody can tell a duplicate from a correction or a second reading.

open as a page

Before widening a table, which count comparison tells you in advance whether any cell will receive more than one value?

level: middleimportance: should knowfreq 45%

basics

~20 s

Count the distinct combinations of identifying values and header-source value, and compare that with the number of input rows. Equal means one value per cell; fewer distinct pairs than rows means some cell will be resolved rather than filled.

open as a page

Column labels made of two parts can sometimes be addressed one part at a time — what determines whether that is possible at all?

level: middleimportance: should knowfreq 46%

basics

~20 s

The tool's label space decides it, not the data. Where column labels are structured, each part is a level you can name, drop or reorder; where labels are a flat list of strings, the parts were composed at creation and only parsing gets them back.

open as a page

After two name columns produced two-part column headers, you take the block of columns under one outer name — what do the returned labels carry?

level: middleimportance: should knowfreq 38%

basics

~20 s

It varies by design: some hand back single-part labels with the part you selected on removed, others keep both parts with that part now constant. Assuming the first is where code breaks, because a lookup by a plain name then matches nothing.

open as a page

A wide table is folded back to one row per measurement and then widened again. When does that round trip return the original?

level: middleimportance: should knowfreq 50%

basics

~20 s

Only on a complete grid with each key-and-header pair present exactly once, and only for a caller indifferent to the order of columns and rows. Where the grid has holes, the two conversions disagree about cells the data never held.

open as a page

A monthly job widens the same table, and its output columns differ from last month's. Why does unchanged code produce different columns?

level: middleimportance: should knowfreq 55%

basics

~20 s

A widening takes its output headers from the data rather than from your code: the distinct values of the header source become the columns. A new or vanished value therefore changes the result's set of columns, and their order comes from the data too.

open as a page

A pipeline converts between one row per measurement and one column per measure around almost every step — what is wrong with that?

level: middleimportance: should knowfreq 50%

basics

~20 s

Converting per step means nobody decided which arrangement the work reads; the pipeline flips back and forth serving one step at a time. Group the steps by the layout they read and convert once at the boundary between the groups.

open as a page

A report averages the amount column of a table holding one row per measurement across six measures and returns a meaningless number — why?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Because a cell's meaning comes from its row's name column, not from the column it sits in: one value column holds six quantities in six units. Any aggregate that does not condition on the name column mixes them.

open as a page

A monthly figure moved after duplicate records appeared upstream, yet the widening that builds the report raised nothing — how do you find and fix that?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Suspect the layout step. Count the input's rows against its distinct key-and-header pairs for the affected period: a gap means the widening resolved occupied cells with a reduction nobody wrote. Fix it by stating the reduction and deciding the grain deliberately.

open as a page

Two-part column headers were glued into single strings with an underscore, and a later step that split them back mis-assigned one column — how did that happen, and how do you prevent it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

One label part already contained an underscore, so that header split into more pieces than it was built from and the pieces landed in the wrong slots. Gluing is reversible only while no part can contain the separator, so the check belongs at glue time.

open as a page

A monthly fold back names the measure columns to be folded. A new measure column appears upstream. What happens, and how would naming the columns to keep differ?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Naming what moves is a closed list, so the new column is not folded: it joins the row key or is discarded, silently either way. Naming the identifying columns to keep is an open list, which folds the new column in automatically.

open as a page

You widen a table so two measures sit side by side, then hand those two columns to a numeric routine as a plain rectangle of numbers. What must you establish at that boundary?

level: seniorimportance: should knowfreq 40%

basics

~20 s

What lines the two operands up. A labelled table matches them on their row labels; a plain rectangle of numbers matches strictly by position, so crossing that boundary stops the ordering being checked for you and a wrong answer arrives with nothing raised.

open as a page

Steps in your team's pipeline come from tools that assume different layouts — which arrangement do you commit to as the interchange form, and what do you accept by choosing it?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

There is no universally right answer, so make it a stated commitment rather than an inherited default. Commit to the long layout where the interface must be stable, to the wide where the mass of work is measure-against-measure, and put every conversion at the named boundaries.

open as a page