skip to content

A labelled table stores a per-row key beside its values. What does that key let you do, and what does nothing check about it?

level: juniorimportance: must knowfreq 72%

answer

  1. a name for a row, not a count
  2. stored beside the values
  3. key in name only
  4. nothing checks it for repeats
  5. not every design has one

basics

~20 s

A row label is a per-row key stored beside the values, so a row can be named rather than counted. Nothing checks it for uniqueness, duplicates are legal, and some tabular designs carry no row identity at all.

solid answer

~50 s

A row label is an identity a holding stores for each row, alongside the columns, so you can ask for the row named `ACC-1042` instead of row 37. It buys a lookup by name, a selection over a range of labels when the labels are ordered, and a name that still points at the same row after other rows are filtered away. What it does not buy is any guarantee: in every in-process design in this family the labels sit beside the values and nothing validates them, so two rows may carry the same label and the holding will store both without complaint. It is a key in name only, unlike a key a storage engine declares and enforces. And the feature is not universal — some tabular designs have no row-identity concept, so identity there is an ordinary column you carry yourself.

go deeper

for a junior

Be able to say in one sentence what a row label is — a per-row key stored beside the values so a row can be named — and that nothing checks it for repeats.

for a middle

Explain the mechanism: the labels are a full-length sequence carried alongside the columns, a design builds a lookup over them, and a retrieval returns however many rows carry the label asked for.

for a senior

Show that you treat unenforced identity as a property of the data you are reading, not a guarantee, and that you can name what in a pipeline would break the first day a label repeats.

for a principal

Frame it as a design split rather than a fact: some holdings carry identity and some expect it to be a column, and any rule you set for a codebase has to hold in both.

## What a row label actually is A **labelled table** — a rectangle whose columns are named and typed one at a time, and whose rows may or may not carry an identity of their own — can hold one thing beyond the values themselves: a **row label** for each row. That per-row key is what lets a reader of your code say *the row for account `ACC-1042`* rather than *row 37*. Designs that have the feature materialise the whole set of labels as a structure beside the columns; the everyday word for that structure, when you meet it in an interview, is *the index*, but the word is a name and the mechanism is the thing: a full-length sequence of keys, one per row, carried alongside the values and updated as rows come and go. It is worth separating three things that are easy to run together: - the **row label** — a key the holding stores for the row; - the **row position** — the integer offset of the row in the holding, which exists whether or not labels do; - an **identifier column** — an ordinary field of the table whose values happen to name the entity. The third is always available. The first is a design feature that some holdings have and others do not. ## What carrying a label buys - **Naming a row.** You can retrieve a row by the key it carries rather than by counting to it, which keeps code readable and keeps it correct when rows move. - **A lookup.** Where the labels are materialised, a design usually builds a structure beside them so that finding the row for a given label is much cheaper than walking every row. - **A range over ordered labels.** When the labels are in ascending order, a selection from one label to another can be resolved by finding two boundaries rather than by testing every row. - **Stability under filtering and reordering.** Drop half the rows or put them in a different order and `ACC-1042` still names the same row; a position does not. ## What nothing checks This is the part interviewers are actually probing. **No in-process holding in this family validates a row label.** There is no uniqueness enforcement, no rejection of a repeat, no renaming. The labels are data stored beside other data. Consequences: 1. Two rows may carry the same label, and the holding stores both. 2. A retrieval by label therefore returns however many rows carry that label — which may be one, and may be four. 3. Because of (2), the number of rows a correct-looking line of code returns is a property of **the data**, not of the code. 4. Nothing announces any of this. There is no warning at the point the duplicate arrives and none at the point the lookup changes behaviour. The expectation that duplicates are rejected is imported from storage engines, where a declared key really is enforced and an offending insert really is refused. Carrying that expectation into an in-process holding is the single most common wrong belief on this subject. ## Designs differ, and the difference matters | holding | row identity | how you name a row | what enforces the label | |---|---|---|---| | labelled table with row identity | a materialised set of labels beside the values | by its label | nothing | | labelled table without row identity | none | by position, or by matching an identifier column | not applicable | | labelled table where row names exist but are discouraged | a restricted single vector, dropped by many operations | mostly by matching a column | nothing | | uniform-type rectangle | none | by position only | not applicable | A **uniform-type rectangle** is one buffer of one representation, addressed only by position, with no column names and no row identity — so the question of labels does not arise there at all. And among label-carrying tabular designs there is a real split: one family treats the row-label structure as a first-class object that operations consult, another has no such concept and expects identity to be an ordinary column, and a third has row names but treats them as a legacy nicety that most operations discard. A candidate who says "every table has row labels" has described one design and called it the class. ## How to answer it out loud Define it in a sentence — a per-row key stored beside the values — then say the two things that make it interesting: it buys a lookup by name, and it guarantees nothing. Then add the design caveat, because that is what separates someone who has used one tool from someone who understands the mechanism: *in designs that carry row identity, the label is a real object with a lookup behind it; in designs that do not, identity is a column, and there is nothing wrong with that.*

  • If nothing enforces a row label, in what sense is it a key at all?
    It is a key in the addressing sense only: it is what a retrieval names a row by. It carries no constraint, no rejection of repeats and no declaration anywhere. The word is borrowed from storage engines, where a declared key is enforced and an offending row is refused; here the same word describes data stored beside data.
  • What happens to the row labels when you filter rows out of a table?
    In designs that carry row identity, the surviving rows keep the labels they had, so the label set is no longer consecutive and no longer matches position. That is the point of the feature: the label names the row, not its place. Code that assumed label and position agree after a filter is relying on a coincidence that filtering removes.
  • How is a row label different from an identifier column holding the same values?
    The values can be identical; the difference is where they live and what they are for. A label is carried by the holding and is what a by-name retrieval addresses, so a design can build a lookup structure over it. An identifier column is an ordinary field of the table, visible in output and available in every design, including those with no row identity.

A handwritten name tag on a coat at a cloakroom. It is useful — you ask for the coat labelled Petrov instead of counting hooks — but nobody at the counter ever checks whether two coats say Petrov, so one day the same request hands you two coats and nothing has gone wrong from the cloakroom's point of view.

saying these in an interview costs you the question

  • Says the holding rejects a duplicate row label the way a storage engine rejects a duplicate key
  • Treats the row label and the row's position as the same thing
  • Assumes every tabular design carries row identity
  • Calls the row labels a primary key without saying nothing enforces them
  • Believes the row labels are always the consecutive row numbers