skip to content

Data With More Than Two Axes

Some data is naturally entity by period by measure, and a table can hold it only by repeating a key or stacking headers. What a third axis buys, and what a dense cube costs.

on this pageshow

questions

3

How can a two-axis table hold a figure recorded for every store, every month and every measure?

level: juniorimportance: must knowfreq 58%

answer

  1. three coordinates, two directions
  2. one coordinate has no home
  3. repeat the keys, or fold into headers
  4. a third axis keeps all three

basics

~20 s

A two-axis table cannot hold three coordinates directly, so it flattens one: repeat store and month down the rows beside a measure column, or fold the measure into the column labels. A holding with a third axis keeps all three.

solid answer

~40 s

The figure has three coordinates — store, month, measure — and a table has two directions, down the rows and across the columns. So one coordinate has to be flattened into another. The common move is to repeat the keys: one row per store-month-measure triple with a single value column, so every store identifier is written once per month per measure. The alternative is to fold one coordinate into the column labels, giving one column per measure. A holding with three axes keeps all three as coordinates instead, and its size is then the three axis lengths multiplied together. Which of the three you want depends on what you do next: per-row filtering and matching favour the flattened table, while reducing along one coordinate favours the third axis.

go deeper

for a junior

Recall that a table offers only two directions to address by, and that a third coordinate therefore has to be repeated down the rows or folded into the column labels. Name both ways.

for a middle

Explain what each flattening repeats and what it makes easy: repeated keys give a self-contained observation per row to filter and match on, folded labels give one column per measure and a width that follows the data.

for a senior

Predict the cost before the pipeline is written. Say which coordinate you would fold, why a coordinate that grows with the data is the wrong one to fold or to make an axis, and what the downstream work does with each holding.

for a principal

The tradeoff worth owning is the boundary: which holding leaves a module and which stays inside it, so a team is not converting between three layouts at every seam and paying for the conversion nobody budgeted.

Three-way data is data where finding one value takes three pieces of information. A sales figure might be *store 114, March, units sold*. A sensor reading might be *device, hour, channel*. An accounting balance might be *account, period, measure*. Whatever the domain, the shape is the same: an entity, a period and a measure, with one value where the three meet. ## Three coordinates, two directions A **labelled table** — a rectangle whose columns are named and typed one at a time, and whose rows may or may not carry an identity of their own — offers exactly two directions to address by: down the rows, and across the columns. The data has three coordinates and the holding has two directions. One coordinate has no direction of its own, so the table has to carry it some other way. That is what *flattening* means here, and a table has two ways to do it. This is not a defect in tables. It is the reason the question gets asked at all: an interviewer wants to hear that the holding has a shape, that the data has a shape, and that when the two disagree something has to give. ## Flattening by repeating the keys Give every value its own row, and write the whole address on that row. - The table has one column per coordinate — store, month, measure — plus one column holding the value. - Every address value is written once per combination it takes part in. With sixty months and twenty measures, a single store identifier appears in twelve hundred rows. - Nothing in the layout requires every combination to exist. A combination that never happened is simply a row that is not there, which is why this layout suits data where most combinations never happen. - A row is a self-contained observation, so filtering on one coordinate, matching against another table, and appending tomorrow's arrivals are all straightforward. The repetition is not a modelling error and it is not something to be normalised away; it is the price of holding a third coordinate in a two-direction holding. How much it actually costs depends on how the design stores a column of many repeated values, which is not the same in every design. ## Flattening into the column labels Pick one coordinate and give each of its distinct values a column of its own. The remaining two coordinates then address the rows. - The table becomes as wide as that coordinate has distinct values: twenty measures, twenty columns. - The table's own width now follows the data. A measure that appears for the first time next month adds a column, and code that named columns explicitly has to be told about it. - Fold two coordinates rather than one and a column is named by a pair of parts rather than by a single value. - Fold the coordinate with the fewest and most stable distinct values. Folding the store coordinate of a forty-thousand-store dataset produces a table nobody can work with. ## Keeping the third coordinate as an axis The third option is a holding that genuinely has three axes. Its **axis-length tuple** — the list of how long each axis is, where the number of entries in that list is the number of axes — has three entries rather than two. - A value is addressed by one position on each of the three axes, and nothing about the address is written per value. - The store identifiers live once, as the labels or the ordering of one axis, rather than once per row. - The holding's size is the product of the three axis lengths, whether or not a value landed in each cell. ## The three options side by side | holding | what addresses a value | what repeats | where it hurts | |---|---|---|---| | keys repeated down the rows | the three address columns on that row | every address value, once per observation | the address is most of each row | | one coordinate folded into the labels | two coordinates for the row, the label for the third | the folded coordinate's values become column names | the table's width follows the data | | a genuine third axis | one position on each of the three axes | nothing | size is the product of the axis lengths | ## The shortcut worth correcting It is often said that a third axis means leaving named data behind for a plain rectangle of numbers. That is true of a **uniform-type rectangle** — one buffer of one representation, addressed only by position, with no column names and no row identity — whose axes are numbered rather than named. It is false of a **labelled multi-axis holding**, which carries names on every axis and coordinate values along each one, so a third axis costs no names at all. There are three options in the room, not two, and which one you get is a property of the design you picked rather than of having three coordinates. ## Choosing between them 1. **What runs most often?** Averaging across every month, for each store and each measure, is one operation along one axis in the three-axis holding and a grouping in the flattened one. 2. **How complete is the data?** Combinations that mostly never happened argue for the flattened layout, because it stores only what happened. 3. **What crosses a boundary?** The layout written to a file, handed to another team, or fed to a model is the one worth committing to. Converting between layouts at every seam is a cost of its own, and one nobody budgets for.

  • If you fold one coordinate into the column labels, which of the three should it be, and why?
    The one with the fewest distinct and most stable values, which is usually the measure. Folding creates one column per distinct value, so a coordinate with thousands of values gives a table thousands of columns wide, and a new value arriving tomorrow changes the table's own width. A coordinate whose values are a short, known list does not.
  • Does moving to a third axis mean giving up the names on your data?
    Not necessarily. A purely positional multi-axis holding numbers its axes and carries no names, so the names have to be kept alongside it. Other designs label every axis and carry coordinate values along each one, in which case a third axis costs no names at all. Which you get is a property of the holding, not of having three coordinates.
  • Why does the repeated-key layout store the same store identifier so many times?
    Because a row in that layout is a complete address plus one value: it must say which store, which month and which measure it belongs to. With sixty months and twenty measures, one store identifier appears in twelve hundred rows. How much that repetition costs depends on how the design stores a column of many repeated values, but the repetition itself is inherent to the layout.

saying these in an interview costs you the question

  • Says three-way data simply will not fit in a table
  • Treats the repeated store and month values as a data-quality error to be fixed
  • Assumes a third axis always means giving up the names
  • Knows only the repeated-key flattening and not the folded-label one
  • Treats the flattened and three-axis holdings as interchangeable in cost
open as a page

In a three-axis holding, why must code that reduces over the last dimension change its axis number after an earlier reduction?

level: middleimportance: should knowfreq 51%

basics

~20 s

Axes are named by their place in the axis-length tuple, so a reduction that removes one drops an entry and shifts every dimension after it down a place. The number that meant measures now means something else.

open as a page

A three-axis holding of 40,000 accounts by 60 months by 20 measures has real values in 4% of its cells — how large is it, and when does a repeated-key table cost less?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A dense three-axis holding allocates the product of its axis lengths — 48 million cells here — regardless of how many carry a value. At 4% density the flattened table, storing one address plus one value per real observation, is far smaller.

open as a page