skip to content

A dataset sits in a table with named typed columns, or in a rectangle addressed only by position - what do the names buy?

level: juniorimportance: must knowfreq 80%

answer

  1. metadata beside the values
  2. names describe, positions do not
  3. one representation covers a bare rectangle
  4. identity is what you match on
  5. descriptive, never enforcing

basics

~20 s

Names and per-field types make a rectangle self-describing: every field says what it is, and each field can hold a different kind of value. A rectangle addressed only by position carries one representation for everything and offsets to address it by.

solid answer

~50 s

A **labelled table** is a rectangle whose columns are named and typed one at a time, and whose rows may or may not carry an identity of their own. A **uniform-type rectangle** is one buffer of a single representation, addressed only by position: no field names, no row identity. The names buy interpretability - a result still says what each field is after five transformations, and a reviewer can check it without re-deriving anything - and they let fields of different kinds sit in one holding, which a rectangle with one representation for every cell cannot do. Where a design also carries row identity, that identity is something two holdings can be matched on, so correspondence survives a reordering. What none of it buys is a guarantee: labels describe, and nothing in the holding checks them.

go deeper

for a junior

Be able to state the difference in one breath: named fields typed one at a time, and possibly an identity per row, on one side; one representation and positions on the other. Say what each buys before discussing cost.

for a middle

Explain the mechanics: where the metadata lives, what a step has to carry through, and why a holding with a single representation for every cell cannot hold fields of different kinds.

for a senior

Show you choose on evidence and know what survives a pipeline. Correspondence by position breaks the first time a step filters or reorders, and nothing reports it.

for a principal

The angle worth arguing is how much meaning your codebase is willing to keep outside the data: positions push it into the reader, names keep it with the values, and neither enforces anything.

## Two holdings for the same numbers A **labelled table** is a rectangle whose columns are named and typed one at a time, and whose rows may or may not carry an identity of their own. A **uniform-type rectangle** - a positional rectangle - is one buffer of a single representation, addressed only by position: no field names, no row identity. The values inside the two can be identical. What differs is the metadata sitting beside them, and what that metadata makes possible. The question underneath the question is whether you know which one you are holding, and can state the ledger in both directions: what the extra metadata buys, and what it charges. ## What the names buy - **A result that is still readable five steps later.** Every operation that produces a new holding also produces names for it. The reader, the reviewer and the next step can all see which field is which without reconstructing it from the code that built it. - **Fields of different kinds side by side.** Each field carries its own representation, so timestamps, text and counts coexist in one holding. A rectangle with a single representation for every cell can only hold fields that share one. - **Addressing that survives a change of layout.** A field referred to by name still resolves after a step inserts, drops or reorders fields. An offset does not; it silently means something else. - **Something a human can check.** A field name is a claim about content. It is a weak claim, because nothing verifies it, but a weak checkable claim beats no claim at all. ## What row identity buys, where a design carries it Row identity is the per-row key a design stores beside the values. Where it exists, it is the thing that survives a reordering: filter the rows, sort them, or keep every third one, and each surviving row still answers to the same key. That is what lets two holdings be brought back together by *which record a row is* rather than by *where it sits*. This is the sentence about tables that is most often stated too broadly, so state the precondition with it: - Some designs materialise the row-label set as a first-class structure beside the values, and a lookup by label is a real operation on the holding. - Some tabular designs have **no row-identity concept at all**. Rows are addressed by position, and any identity has to be an ordinary field you carry and match on explicitly. That is a deliberate choice with its own advantages, not an omission. - Some keep row names but discourage them - a single text vector, dropped by most operations. So "you can always look a row up by its identity" is true of the first group and false of the other two. ## What positions buy The bare rectangle is not the poor relation; it is the specialised one. - Nothing has to be carried through an operation, because there is no per-field or per-row metadata to carry, and nothing has to be reconciled between two operands. - Every cell is the same kind and the same width, so a step applied to the whole holding is uniform. - Correspondence between two holdings is position alone, which is cheap - and silently wrong the moment one of them is reordered. ## The ledger, side by side | | labelled table | rectangle addressed by position | |---|---|---| | addressing a field | by name | by offset | | addressing a row | by position, and by label where the design carries one | by position | | kinds of value | one representation per field | one representation for the whole holding | | identity of a row | depends on the design | none | | metadata per operation | names carried; labels reconciled where both sides carry them | none | | what any of it enforces | nothing | nothing | ## What neither of them gives you Labels describe; they do not constrain. Naming a field does not type-check it, and carrying row labels does not make them unique - nothing in an in-process holding checks either one. This surprises people arriving from a storage engine, where a declared key is enforced by the engine itself. Here the metadata is documentation that travels with the data: genuinely valuable, and not a guarantee. ## How to answer it out loud 1. Name both holdings and what each one is made of, in one sentence each. 2. Give the ledger in both directions - names buy interpretability and mixed kinds; identity, where the design carries it, buys correspondence that survives reordering; positions buy uniformity and no metadata to keep straight. 3. Close on the limit: none of it enforces anything. A candidate who says only "a table is more convenient" has described a preference. A candidate who says what each buys, what each charges and where designs disagree has described the object.

  • Does a rectangle addressed only by position have any advantage over a labelled table holding the same numbers?
    Yes. With one representation for every cell and no per-field or per-row metadata, nothing has to be carried or reconciled through a step, and every cell is the same kind and width. The cost is that what each position means lives outside the holding - in code, in a comment, or in someone's head - so it cannot be checked by reading the data.
  • Can you always look a row up by its identity in a labelled table?
    No, and that is a design difference rather than a detail. Some designs materialise a row-label structure beside the values, so a lookup by label is a real operation. Other tabular designs have no row-identity concept at all: rows are addressed by position and any identity has to be an ordinary field you match on explicitly. A third group keeps row names but drops them across most operations.
  • If the field names are only descriptions, why not keep the numbers positionally and document the ordering?
    Because the document is not attached to the data and the names are. A name travels with the values through every step; a written-down ordering has to be re-established by hand at each step and fails silently the first time a step drops or reorders a field. The name still enforces nothing - it just stops the mapping drifting away from the values.

saying these in an interview costs you the question

  • Says a labelled table is faster than a bare rectangle of numbers
  • Thinks naming a field constrains what values it may hold
  • Assumes every labelled table carries row identity you can look up by
  • Claims a rectangle addressed by position can mix text and counts freely
  • Says position is just as readable once the ordering is documented