A lookup by row label returned one record in March and returns four today, with the same code. What happened?
answer
- uniqueness was an accident, not a rule
- the label now repeats
- one match and several match differ in kind
- result shape follows the data
- silent until something downstream unpacks
basics
~20 sThe label now appears on four rows. Nothing ever checked row labels for uniqueness, so a retrieval returns however many rows carry the label asked for — and the shape of the result is decided by the data, not by the code.
solid answer
~50 sRow labels are stored beside the values and never validated, so uniqueness was an accident of March's data rather than a property of the holding. Once the same label lands on four rows, the same retrieval hands back four rows instead of one, and in most designs it hands back a different kind of thing as well: a small table rather than a single record. Nothing raised, because nothing was violated. What breaks is downstream: an arithmetic step that expected one number now works over four, a write-back touches four rows, and a step that unpacked fields from a record now unpacks from a group. The general lesson is the one to say out loud — **a retrieval keyed on an unenforced label has a result shape that moves with the input**, which is why code that looks stable can change behaviour without being edited.
code
pseudocode · 8 lines// March: exactly one row carries "ACC-1042"
row = lookup_by_row_label(accounts, "ACC-1042")
new_balance = row.balance * 1.05 // one number in, one number out
// Today: four rows carry "ACC-1042" after a second delivery
row = lookup_by_row_label(accounts, "ACC-1042")
// the retrieval now hands back four rows, not one record
new_balance = row.balance * 1.05 // four numbers, and no error was raisedgo deeper
Remember that a lookup by row label can return more than one row, because the same label may sit on several rows and nothing prevents that.
Explain the mechanism and the kind change: one match may come back as a record and several as a group, so downstream arithmetic and unpacking see a different object.
Trace the blast radius — the inflated aggregate, the write-back that touched four rows — and say what in the pipeline should stop depending on a retrieval's arity.
Treat it as a grain question: decide whether identity at this stage is one row per entity or something finer, and make the code state which rather than inherit it from the data.
## The claim under the bug The code did not change, so something about the holding's guarantees was misread. The misreading is specific: someone treated the **row label** — the per-row key a holding stores beside the values — as if something validated it. Nothing does. In March each label happened to occur once; that was a property of March's data. Today one label occurs four times, which is equally legal, and the retrieval faithfully returns all four rows. ## Why the result changed kind, not just size This is the part that turns a data change into a program failure. Many label-carrying designs give a by-label retrieval a **polymorphic** result: - when exactly one row carries the label, you get a single record — one value per field; - when several rows carry it, you get a group of rows — a small holding of the same width. So the failure is not only "four instead of one". Downstream code receives an object of a different shape, and what happens next depends on what that code does with it: 1. **Arithmetic on a field** that was a single number now runs over four values, and the result is four values rather than one. 2. **Unpacking or formatting** a record breaks outright, or quietly formats the first of four. 3. **A write back through the retrieval** touches four rows where it used to touch one. 4. **A row count check further down** sees a number that is right for the data and wrong for the assumption, if anyone is looking at it at all. 5. **An aggregation** downstream silently absorbs the extra rows and returns a plausible, wrong number — the worst outcome, because it raises nothing and looks like a result. ## Why nothing raised Because nothing was violated. There is no declaration anywhere in the program saying this label is unique; there is no constraint attached to the holding; there is no validation step at the point rows are added. The expectation lived only in the author's head and in the shape of the code that consumed the result. This is the honest difference between an in-process holding and a storage engine that enforces a declared key: in the engine the offending row is refused at the door, and in the holding it is simply stored. ## Where the duplicates came from You do not need to know, to answer the question — but naming the usual arrivals shows you have lived with it: | arrival | why the label repeats | |---|---| | a second delivery of the same period | the same entity appears in both batches, each with its own row | | a labelling set from a column that was not unique | the column was unique in the sample and is not in production | | a corrected record shipped alongside the original | two rows now describe one entity, deliberately | | a widened grain | rows used to be one per entity and are now one per entity per period | The last one is the most instructive: the data did not become wrong, it became finer, and the code's assumption about identity silently became a statement about a grain that no longer holds. ## Designs differ It is worth saying explicitly that the shape-shifting retrieval is a property of designs that carry row identity. Where a design has **no row-identity concept**, identity is an ordinary column and you reach the entity's rows by testing that column — which always yields a set of rows, possibly of size one, possibly empty. That is more verbose and it never changes kind under you: the shape of the result is fixed by the operation, and only its row count follows the data. Neither model is better in the abstract; they trade convenience against a failure mode. ## What to say in an interview Give the mechanism, not the incident: *row labels are unenforced, so a by-label retrieval returns however many rows carry the label, and in several designs a single match and a multiple match come back as different kinds of object*. Then name the consequence that matters — **the shape of a correct-looking result depends on the input** — and finish with what you would do about it in the code you own: stop depending on a retrieval's arity, or carry identity somewhere the operation's result shape does not depend on the data.
- Would the same code have failed loudly if the identifier were an ordinary column instead of the row labels?It would have failed differently, and usually more visibly. Reaching rows by testing a column yields a set of rows whatever the data does, so the code has to decide up front what to do with more than one. The arity still follows the data, but the kind of the result does not, so there is no silent switch from record to group.
- The duplicate labels are legitimate — the grain really is finer now. What changes?Then the bug is the assumption, not the data. The step needs to say which row it wants — the latest, the corrected one, the one for a given period — and that choice has to be explicit, because the holding will never make it for you. Reading the arity as an error would throw away real records.
- Why is a downstream aggregation the most dangerous consumer of this?Because it absorbs the extra rows without complaint and returns a number of the right type and a plausible magnitude. A sum is inflated, a mean is dragged toward whichever entity duplicated, and nothing in the output looks wrong. Failures that change a number rather than raise are the ones that reach a report.
saying these in an interview costs you the question
- Says the retrieval must have been written against positions rather than labels
- Assumes the holding would have rejected the duplicate when it arrived
- Treats the extra rows as corruption rather than as legal data
- Believes the failure would have raised somewhere before the report
- Claims a by-label retrieval always returns exactly one row