skip to content

A lookup of one row label returns twelve rows — what does that tell you, and what would the same request by offset return?

level: middleimportance: should knowfreq 49%

answer

  1. the row count is telling you something
  2. a match, not a coordinate
  3. labels can repeat, positions cannot
  4. the result's kind moved, not just its size

basics

~20 s

It tells you the row labels are not unique, so they are not a key. Label addressing matches and returns every row carrying that label. An offset cannot do this: one offset names exactly one row, whatever the data looks like.

solid answer

~40 s

The two schemes differ in kind here, not just in convention. Addressing by label is a **match**: the design compares the value you gave against the row labels and hands back every row that carries it, so twelve rows means twelve rows are labelled the same and the labels are not a key. Addressing by offset is a **coordinate**: position 3 is one place in the current layout and can only ever designate one row. The practical consequence is a result shape that depends on the data. With unique labels, a single-label lookup often collapses to a bare record; with a duplicate, the same expression hands back several rows. Code written against the first shape breaks on the second, and it breaks on input rather than on deployment, which is why it survives review.

go deeper

for a junior

Recall that a row label can repeat while a position cannot, so asking for one label may return many rows while asking for one offset always returns exactly one.

for a middle

Explain the mechanism: label addressing compares a value against the row labels and returns every match, so the result's shape depends on the data, while an offset is a coordinate with a fixed result shape.

for a senior

Show what breaks downstream — a unique match that collapses to a bare record instead of a one-row table, and an aggregate that silently sums twelve rows — and where you would place the uniqueness check so it fails at the step responsible.

for a principal

Frame the trade: identity-bearing addresses have data-dependent result shapes, coordinates have fixed ones and no identity. Decide as a standard whether labels are asserted to be a key in this codebase, and where that assertion lives.

## A label is a match, not a coordinate The two addressing schemes are usually described as "by name" against "by position", which makes them sound symmetrical. They are not. An offset is a coordinate into the current layout: there is exactly one row at position 3, always, by construction. A label is a **value compared against the row labels**, and comparison has no such guarantee. If three rows are labelled `A-19`, asking for `A-19` gets three rows, because three rows answer to that name. So a lookup returning twelve rows is not a malfunction. It is a measurement, and what it measures is the data: **the row labels are not unique, therefore they are not a key.** That is worth knowing on its own, because a great deal of downstream code silently assumes they are. ## What comes back | the request | labels are unique | labels repeat | |---|---|---| | one row label | one row, often collapsed to a bare record | several rows, as a table | | a range of labels | the rows between the two endpoints inclusive | ambiguous — designs differ, some reject it outright | | one offset | exactly one row | exactly one row | | a list of offsets | exactly that many rows | exactly that many rows | The offset rows are the boring ones, and that is the point: the offset scheme's result shape is a function of the request alone. The label scheme's result shape is a function of the request **and the data**. ## The shape change is the real bug The returned row count is visible and easy to check. The change of *kind* is not, and it is what actually breaks code: - In several designs a single-label lookup that matches exactly one row **collapses to a bare record instead of a one-row table**. The next step, written to read a column out of a table, now gets something with a different surface. - The same expression against duplicated labels stays a table, so the next step works — for that input. - Which branch you get is decided by data the author never inspected, so the failure is input-triggered. It passes review, passes the tests built from a clean fixture, and appears the first time a source system emits the same identifier twice. There is a second, quieter consequence. An aggregate computed over "the row for customer 4471" is correct when the label is unique and silently sums twelve rows when it is not. Nothing raises; the number is simply too big. ## Why offsets cannot behave this way A coordinate cannot match twice. Ask for offset 3 and you get one row or an error, never both and never twelve. That makes the offset scheme's shape predictable, and it is a genuine advantage — but it buys the predictability by giving up identity, which is the thing you usually wanted. The trade is the whole subject: the scheme that can name a record is the scheme whose result depends on the data; the scheme with the fixed result shape cannot name a record at all. ## Guarding it 1. **Check uniqueness once, where the labels are established**, not at every use. A single assertion that the labels are unique converts every downstream assumption from hope into an invariant, and it fails at the step that broke it. 2. **Assert the row count after a single-label lookup.** One line, and it turns "the total is mysteriously high" into a failure at the lookup. 3. **Do not rely on the collapse.** If the next step wants a table, ask for a result that is always a table, so the code takes one branch on every input. 4. **When labels are genuinely not unique, stop treating them as an address.** They are a grouping value at that point, and the right operation is one that expects several matches rather than one that tolerates them. ## What an interviewer is checking They want to hear you say *not unique* immediately, from the row count alone, before looking at any data — and then say what that implies for the two schemes. The weak answer treats twelve rows as a bug in the lookup. The strong answer treats it as information about the labels, notes that offsets are immune because a coordinate cannot match twice, and then goes to the part that hurts: the result's kind, not just its size, moved under the code.

  • Why is the change of result kind worse than the change of row count?
    Because the count is visible and the kind is not. A step that reads a column out of a table keeps working when the count moves from one to twelve, and stops working when a unique match collapses to a bare record instead. The second failure depends on input, so it survives review and clean fixtures and appears in production.
  • The labels are legitimately not unique. What should the lookup become?
    An operation that expects several matches rather than one that tolerates them. Treat the repeated value as a grouping value, not an address: ask for all rows carrying it and then say explicitly which one you want — the earliest, the largest, all of them summed. The bug is not the duplication, it is code that addressed as though there could only be one.

saying these in an interview costs you the question

  • Calls the twelve-row result a bug in the lookup itself.
  • Assumes row labels are unique because they usually have been.
  • Believes an offset can also match several rows at once.
  • Checks only the row count and misses the change of result kind.
  • Adds a take-the-first step instead of deciding which row is wanted.