skip to content

The same numbers are held once with column names and row labels, once as a bare rectangle - what does the labelling cost?

level: middleimportance: should knowfreq 58%

answer

  1. three charges, not one
  2. names cheap, row labels it depends
  3. a default labelling is a rule
  4. reconciled per step, both sides
  5. scales with labels, not values

basics

~20 s

Three charges: the bytes the labels occupy, the metadata every step carries through and reconciles, and a second way of naming a row that the code must keep straight. Field names are cheap; a row labelling materialised from a field is not.

solid answer

~50 s

Field names are a handful of strings per table however many rows there are, so their bytes never matter. Row labels are where bytes appear, and it depends which labelling: the default consecutive-integer labelling is held by several designs as a description of a range - a start, a stop and a step - so it really is close to free, while a labelling set from an existing field materialises a full-length array of those values plus whatever lookup structure the design builds beside it, and a label made of several parts costs one array per part. The second charge is per step: names are carried through and the result's names decided, and where **both** operands carry an identity per row the two label sets are reconciled before any arithmetic - work proportional to the labels, not to the values. The third is human: two ways to name one row. A design with no row identity pays only the first and third.

go deeper

for a junior

Know that labelling is not free and that the field names are the cheap part. Being able to say bytes, work per step and a second way of naming a row is enough at this level.

for a middle

Explain where each charge lives and what it scales with, and distinguish a default consecutive-integer labelling from one set from a field - the first is a rule, the second is an array.

for a senior

Judge when the charges are buying something. A dense uniform numeric computation with no matching pays all three and collects on none; a pipeline whose results must be read by other people collects on all three.

for a principal

The tradeoff to own is the human charge: a second way to name a row is a standing source of ambiguity in a codebase, and whether it pays back depends on how many people read the code that survives.

## The charge is three things, not one 1. **Bytes** - what the labels themselves occupy. 2. **Work per step** - metadata that has to be carried through, decided, and in some designs reconciled, on every operation. 3. **A second coordinate system** - two ways to name a row, which the code and the reader must keep straight. Only the first is what people usually mean by overhead, and it is often the smallest of the three. ## The bytes: names are cheap, row labels depend Field names are one short string per field. A table with forty fields carries forty strings whether it has a thousand rows or a hundred million, so measuring them is a waste of an afternoon. Row labels are per row, and what they cost turns entirely on where they came from: - **The default consecutive-integer labelling is close to free in several designs**, which hold it as a description of a range - a start, a stop and a step - rather than as materialised values. Nothing is allocated per row. - **A labelling set from an existing field materialises a full-length array** of those values, plus whatever lookup structure the design builds beside it so that finding a row by its label beats scanning for it. - **A label made of several parts costs one array per part**, and multiplies the work of bringing two label sets together. - **In a design with no row-identity concept, this whole line of the ledger is zero**, because there is nothing to store. So "the row labels cost nothing, they are just the row numbers" is true of exactly one case - the default labelling - and misleading everywhere else. ## The work per step Every operation on a labelled holding does something with the metadata as well as with the values: - it carries the field names through, and decides what the result's names are; - where an operation combines two operands and **both carry an identity per row**, it reconciles the two label sets before computing anything; - where only one operand carries identity, or neither does, there is nothing to reconcile and **position is the only correspondence** - which is cheap, and silently wrong if an operand has been reordered. The property worth stating is what this charge scales with. It is proportional to how many row labels there are and how complicated each one is, not to the width of the values or the cost of the arithmetic. On a narrow table with an elaborate label it can dominate the step; on a wide numeric computation over a simple label it vanishes into the noise. ## The second coordinate system The third charge is not measured in bytes or seconds. A row in a labelled holding can be named two ways: by where it sits, and by what it is called. The two coincide only while the labelling is the default consecutive integers and nothing has reordered the rows - which is to say at the beginning and rarely afterwards. Everyone reading the code then has to know which of the two a given line means. That cost is paid by the team rather than by the machine, and it produces wrong answers rather than slow ones, which makes it the expensive one. ## Which charge applies where | charge | rectangle by position | table with no row identity | table carrying row identity | |---|---|---|---| | bytes for field names | none | negligible | negligible | | bytes for row labels | none | none | near zero for a default labelling; a full-length array per part for one set from a field | | metadata carried per step | none | field names only | names, plus label reconciliation when both operands carry identity | | second way to name a row | no | no | yes | ## When the charge is worth paying - When the result has to be interpretable by someone who did not write the step that produced it. - When two holdings have to be matched by identity rather than by order. - When fields of unlike kinds must live in one holding. - When a lookup by label is a real part of the work, so the array and the structure beside it are buying something back. And when it is not: a dense numeric computation over values that are all one kind, where every step runs over the whole holding and nothing is matched by identity, pays all three charges and collects on none of them. ## How to say it in an interview Separate the three charges, then attach the precondition to each. Naming the default-versus-set distinction on the bytes, and the both-operands condition on the reconciliation, is what distinguishes someone who has measured this from someone repeating that labels have overhead.

  • When is the reconciliation charge on an element-wise operation actually zero?
    Where neither operand carries an identity per row - a positional rectangle, or a tabular design without the concept - there is nothing to reconcile, and position is the only correspondence: cheap, and silently wrong if one side was reordered. Where identity is carried, two label sets that are already the same is the cheap case; two different sets that must be brought together is the expensive one.
  • Why does the per-step charge grow when a row label has several parts?
    Each part is a full-length array of its own, so a label in three parts is three arrays to hold and three to compare when two holdings are brought together. The charge tracks the labels rather than the values, so multiplying the parts multiplies it.

saying these in an interview costs you the question

  • Says row labels cost nothing because they are just the row numbers
  • Assumes the bytes of field names are worth measuring
  • Thinks reconciliation work scales with the values, not the labels
  • Believes every design pays these charges, including ones without row identity
  • Ignores that a row label in several parts costs one array per part