skip to content

A three-axis holding of 40,000 accounts by 60 months by 20 measures has real values in 4% of its cells — how large is it, and when does a repeated-key table cost less?

level: seniorimportance: should knowfreq 44%

answer

  1. size set by the axis lengths
  2. allocated before any value arrives
  3. density against the repeated address
  4. crossover near one over axes-plus-one

basics

~20 s

A dense three-axis holding allocates the product of its axis lengths — 48 million cells here — regardless of how many carry a value. At 4% density the flattened table, storing one address plus one value per real observation, is far smaller.

solid answer

~50 s

The axis lengths multiply: 40,000 by 60 by 20 is 48 million cells, and a dense holding allocates every one of them whether or not an observation landed there. At 4% density only about 1.9 million of those cells mean anything. The flattened alternative stores one row per real observation, and each row carries an address — account, month, measure — plus the value, so roughly 7.7 million entries against 48 million cells. As a rule of thumb the flattened layout starts to win below a density of about one over the number of axes plus one, because that is where the repeated addresses cost less than the empty cells. Two things move that crossover: how narrowly the design stores a column of repeated addresses, and whether a sparse multi-axis holding is available, which stores coordinates beside values and sidesteps the choice.

go deeper

for a junior

Recall the arithmetic: a dense multi-axis holding has as many cells as its axis lengths multiplied together, and that number is fixed before any value is loaded into it.

for a middle

Explain both sides of the comparison — empty cells on one side, repeated address values on the other — and work out the density at which the two meet.

for a senior

Measure the density on real data before committing, and say what will change it: an entity with one month of history, a measure reported quarterly, a backfill that widens an axis.

for a principal

Decide what the team does when density drifts. A holding chosen at 40% density and still in place at 4% is a standing cost with no owner, so the choice needs a trigger to revisit it.

## The arithmetic first Multiply the axis lengths: 40,000 accounts times 60 months times 20 measures is 48,000,000 cells. That number is a property of the three axis lengths and of nothing else. It is fixed before a single value is read, it does not consult the data, and it does not fall when the data turns out to be thin. At 4% density, about 1,920,000 of those cells carry a real observation. The other 46 million or so are cells that were allocated because the coordinates they sit at are legal, not because anything was recorded there. ## Why the empty cells are still there A dense multi-axis holding addresses a value purely by position: so many steps along the first axis, so many along the second, so many along the third. That addressing only works if every combination has a cell, because the position of a cell is computed from the axis lengths rather than looked up. The completeness is not a policy choice; it is what makes the addressing cheap. The consequence is worth saying plainly, because it is the thing candidates get wrong: **occupancy and allocation are unrelated**. Nothing shrinks because a cell stayed empty, and nothing is released later. This is also why an axis that grows with the data is the expensive one. One new account adds a position on the account axis, and that position brings 60 times 20 — twelve hundred — new cells with it, of which perhaps one is real. ## What the flattened layout costs instead The repeated-key layout stores one row per real observation, and each row is a complete address plus the value it addresses: - one entry naming the account, - one entry naming the month, - one entry naming the measure, - one entry holding the value. So roughly four entries per observation, or about 7.7 million entries for 1.9 million observations. That is less than a fifth of what the dense holding allocates, and it is the reason a sparse three-way dataset so often stays flat. | | dense three-axis holding | keys repeated down the rows | |---|---|---| | what sets the size | the three axis lengths multiplied | how many observations exist | | cost per real observation | one cell | about one entry per axis, plus the value | | cost per absent observation | one cell | nothing | | effect of a new account | a whole slab of 60 x 20 cells | as many rows as that account reported | ## Where the crossover sits Compare the two counts. The dense holding stores one entry per cell, so its total is the product of the axis lengths. The flattened layout stores about *(number of axes + 1)* entries per real observation. Setting them equal gives a density of roughly one over the number of axes plus one — about a quarter for three axes. Above that density the dense holding is competitive and often better; well below it, the flattened layout wins, and at 4% it wins by a wide margin. Treat that as a rule of thumb rather than a formula, because the two sides do not store equally cheap entries. Two things move it in practice: 1. **How the design stores a column of repeated addresses.** A column whose values come from a short list of distinct names can be held far more cheaply than a column of distinct values, which pushes the crossover down and favours the flattened layout further. 2. **Whether a sparse multi-axis holding is on the table.** Some designs offer a multi-axis holding that stores only occupied coordinates and their values. That keeps the axis model without paying for the empty cells, and it changes the question from *which of two* to *which of three*. ## Density is not a constant The number that matters is the one the data has now, not the one it had when the holding was chosen. Density moves in one direction most of the time, and it is downward: - new entities arrive with a short history each, adding full slabs of mostly empty cells; - the period axis only ever extends, so old entities that have stopped reporting still occupy every new period; - a measure reported quarterly sits empty in two months out of every three; - a backfill that widens any axis widens the product by the lengths of the other two. A holding chosen when two cells in five were filled and still in place when one in twenty-five is filled is a cost nobody re-examined. If you pick the dense holding, pick a density at which you will revisit it, and measure it rather than assume it. ## What to say in the interview Do the multiplication out loud, state that allocation follows the axis lengths and not the occupancy, compute the flattened layout's entry count including the addresses, and then name the crossover and what moves it. The signal is not the arithmetic itself — it is that you compared both sides honestly, including the address columns that people routinely forget to count, and that you asked what the density will be in a year rather than only what it is today.

  • What happens to the size of the dense holding when one new account appears with a single month of data?
    It grows by a whole slab. One new position on the account axis brings 60 times 20 more cells with it, twelve hundred of them, of which one is real. A dense holding grows by the product of the other axis lengths for every position added to any axis, which is why an axis that grows with the data is the expensive one to have.
  • Does the density crossover mean a sparse dataset should never be held on three axes?
    No. If the work is a reduction along one coordinate — averaging every account across months, say — the three-axis holding expresses that directly while the flattened layout has to group first. The real question is whether the density penalty buys enough of that, and whether a sparse multi-axis holding gives you both.
  • Which of the three axes would you try hardest to keep short?
    The one that grows without bound, usually the entity axis, because every position on it multiplies by the lengths of the other two. Twenty stable measure names are cheap to have as an axis; an account list that grows every week adds twelve hundred cells per arrival in this example.

A wall of pigeonholes built for every combination of floor, corridor and room number. The carpentry bill is set when the wall goes up, by the three counts multiplied together, and it does not fall because most of the holes stay empty. A stack of addressed envelopes costs by the letter instead — one address written on each — which is cheaper only while the letters are few.

saying these in an interview costs you the question

  • Thinks an empty cell in a dense holding costs nothing
  • Assumes a third axis is always more compact than repeated keys
  • Forgets the address columns when sizing the flattened layout
  • Treats every multi-axis holding as dense, so never considers a sparse one
  • Picks the holding on elegance rather than on measured density
  • Assumes today's density is the density the holding will live with