In a three-axis holding, why must code that reduces over the last dimension change its axis number after an earlier reduction?
answer
- position, not name
- the axis-length tuple shrinks
- everything after it shifts down
- keep it at length one instead
basics
~20 sAxes are named by their place in the axis-length tuple, so a reduction that removes one drops an entry and shifts every dimension after it down a place. The number that meant measures now means something else.
solid answer
~50 sIn a purely positional multi-axis holding a dimension has no name — it is identified by where it sits in the axis-length tuple, the list of how long each axis is. Averaging across months removes the month dimension, so the tuple goes from three entries to two and the measure dimension, which was third, becomes second. Later code that still names the third position is out of range, or worse names a dimension that exists and is the wrong one. Two things avoid it: some designs let a reduction keep the reduced dimension at length one, so nothing renumbers; and a holding whose axes carry names lets you name the measure axis, in which case no number is in play. Note that in a two-axis table the same selector usually picks the direction the operation runs rather than a dimension — the same number meaning something different.
code
pseudocode · 17 linesaxis lengths: (40000, 60, 20) # 1st = accounts, 2nd = months, 3rd = measures
reduce over the 2nd dimension # average across months
axis lengths: (40000, 20) # 1st = accounts, 2nd = measures
# the measure dimension was the 3rd; it is now the 2nd.
# a later step that still names the 3rd dimension is out of range here --
# and on a holding that started with four axes it would quietly hit a
# different, legal dimension instead.
#
# whether the first dimension is called 0 or 1 differs between designs.
# what does not differ: removing one shifts every dimension after it.
# the form that avoids it, where a design offers it:
reduce over the 2nd dimension, keeping it at length one
axis lengths: (40000, 1, 20) # the 3rd is still the 3rdgo deeper
Recall that in a positional holding a dimension is identified only by where it sits among the axis lengths, so the identifier is a position, and positions move when the holding changes.
Explain the mechanics: the reduction removes one entry from the axis-length tuple, every later dimension shifts down a place, and a hard-coded position then addresses a different dimension or falls out of range.
Show the habit that prevents it. Reduce in an order that leaves the positions you still need untouched, keep the reduced dimension at length one, or work in a holding whose axes carry names.
Decide whether positional dimension numbers may appear in shared code at all. A number that is correct only given the pipeline's current order is a coupling between two files that nothing verifies.
## A dimension with no name In a positional multi-axis holding, a dimension has no identity of its own. It is identified entirely by where it sits in the **axis-length tuple** — the list of how long each axis is, whose number of entries is the number of axes. A holding of accounts by months by measures might have axis lengths (40,000, 60, 20). The account dimension *is* the first entry, the month dimension *is* the second, the measure dimension *is* the third. Nothing else names them. That is a perfectly workable model, and it is what lets the holding compute a cell's position arithmetically instead of looking anything up. But it has one consequence that surprises people every time: **the identifier is a position, and positions move.** ## What a reduction does to the tuple Averaging across the months collapses that dimension to a single number per remaining combination, so the month dimension is gone. The axis-length tuple goes from three entries to two: (40,000, 20). The remaining lengths close the gap. The measure dimension has not changed in any way a human would notice — it still has twenty positions and the same twenty meanings — but it is no longer the third entry. It is the second. Every later operation that identified it by position was written against the tuple as it was before the reduction, and that tuple no longer exists. ## The two failure modes - **Out of range.** The holding now has two dimensions, and a call naming the third is rejected. This is the good outcome: it is loud, it happens immediately, and it points at the line that is wrong. - **Silently the wrong dimension.** Start with four axes rather than three, remove one in the middle, and a call naming the fourth position now addresses what used to be the third. It is a legal dimension, of a legal length, and the operation succeeds. The result is wrong and nothing says so. The second is the reason this is asked at all. A number that is correct only given the order in which previous steps happened is a coupling between two pieces of code that nothing checks. ## Three ways out 1. **Keep the reduced dimension at length one.** Several designs offer a form of the reduction that writes a single value into the dimension instead of removing it. The tuple keeps the same number of entries, every later position keeps its meaning, and the holding is only marginally larger. This is the cheapest fix when the pipeline is already written against positions. 2. **Recompute the position instead of hard-coding it.** Read the current axis lengths and derive the position you want at the point you need it, so the value tracks whatever the previous steps did. 3. **Use a holding whose axes carry names.** In a **labelled multi-axis holding** — one with three or more axes where each axis carries a name and coordinate values as well as positions — the operation names the measure axis, and removing the month axis renames nothing. The problem does not arise, because no position was ever the identifier. ## The same selector, three different meanings The argument that tells an operation which way to run — **the axis selector** — is the same argument in all three of these holdings, and it names a different kind of thing in each. This is the single commonest confusion in this whole family, and it is worth being explicit about which one you mean before any number appears. | holding | what the selector names | does a reduction renumber anything? | |---|---|---| | two-axis table | the direction the operation runs: down the rows or across the columns | no — there are only two directions, and both always exist | | positional multi-axis holding | a dimension, by its place in the axis-length tuple | yes — removing one shifts every dimension after it | | labelled multi-axis holding | an axis, by its name | no — names do not move | The row that catches people is the first one. In a two-axis table the selector is not identifying a dimension at all; it is choosing between two directions, and which number corresponds to which direction routinely reads as the opposite of what people expect. Importing the positional reading into a two-axis table, or the other way round, produces code that runs and answers the wrong question. One more variation worth knowing: **which number the first dimension carries is not the same across designs.** Some count from zero and some from one. What does not vary is the mechanic underneath — removing a dimension shifts every dimension after it by one place — so reason about the shift rather than about the literal number. ## How to answer it Say that the identifier is a position rather than a name, that a reduction shortens the axis-length tuple, and that everything after the removed dimension therefore moves. Then name the failure that actually costs money — the call that lands on a legal but wrong dimension and produces no error — and give the fix you would actually use: keep the dimension at length one, or work in a holding where the axis has a name.
- How does keeping the reduced dimension at length one avoid the renumbering?The reduction writes a single value into that dimension instead of removing it, so the axis-length tuple keeps the same number of entries and every later dimension keeps its position. The holding is marginally larger and reads a little oddly, but every position still means what it meant, which is what downstream code was depending on.
- Is the selector in a two-axis table the same thing as a dimension number?Usually not. In a two-axis table the selector normally says which way the operation runs — down the rows or across the columns — and which number means which is a routine source of confusion. In a positional multi-axis holding the same argument names a dimension by its place among the axis lengths. Same argument, two different meanings.
- What is the equivalent problem in a holding whose axes are named?It largely does not arise: the operation names the measure axis, and removing the month axis renames nothing. What can still bite is an operation returning a holding with one fewer axis where later code assumed three, so the axis count still needs checking even when the identifiers themselves are stable.
saying these in an interview costs you the question
- Believes the number attached to a dimension is stable across operations
- Confuses a dimension of the holding with the direction an operation runs
- Assumes every design numbers dimensions from the same starting value
- Thinks a reduction leaves the number of axes unchanged
- Reaches for a hard-coded position where the holding offers a name