Why must each row planted in a shared test dataset carry a label that survives a rebuild?
answer
- Failure reports should name situations
- A count mismatch explains nothing
- Rebuilds renumber rows
- Ask by label, never by identifier
basics
~20 sA label ties a failing test to the exact planted row it hit, turning a bare count mismatch into a named situation. Because the dataset is rebuilt, the label must come from the build definition, not from a store-assigned identifier.
solid answer
~40 sGive every planted row a stable, meaningful name — `customer.surname_absent`, `order.amount_at_maximum` — and carry that name as data: a dedicated column, a manifest keyed by label, or a reserved identifier range. Tests then ask for a row by label rather than by a number someone read off a screen. Two things follow. A failure reports which situation broke, so the report is diagnostic without anyone opening the dataset. And when the dataset is rebuilt on a new machine with fresh numbering, the label still resolves, because the build assigned it, while a stored identifier will not. The working rule: a test may hard-code a label, never a row identifier.
code
pseudocode · 11 linesplant(label: "customer.surname_absent", record: { surname: EMPTY })
plant(label: "order.amount_at_maximum", record: { amount: agreed_maximum })
# a test resolves by label, never by a stored identifier
row = dataset.by_label("customer.surname_absent")
assert renders_without_error(display(row))
else fail("planted row customer.surname_absent broke the listing")
# after a rebuild the numbering moves and the label does not
before: label "customer.surname_absent" -> id 41827
after: label "customer.surname_absent" -> id 90114go deeper
Be ready to say why a test should ask for a planted row by a name you chose rather than by a number the store assigned, and what the failure message looks like in each case.
Explain that a rebuild renumbers rows, so a hard-coded identifier either breaks or silently points at an ordinary record. The label has to be written by the dataset build and read back by the test.
Show where the label lives — a column, a manifest, a reserved range — what each choice costs, and why the build should fail on a duplicate label rather than resolving it arbitrarily.
Own the convention across teams: one naming scheme for planted rows, one place it is defined, and a review rule that no test hard-codes a stored row identifier.
Planting a rare row is half the work. The other half is making sure that when a test fails on it, the failure says so. ## What a failing test should tell you Compare two reports from the same broken build. The first says `expected 3 results, got 2`. The second says `planted row customer.surname_absent did not appear in the listing`. Both describe the same defect. Only one of them can be triaged without opening the dataset, and only one of them survives being read by somebody who did not write the test. The difference is that the second test asked for a row **by a name that describes the situation**, so the name was available to put in the message. That name is the label, and it is the whole subject here. A good label describes the situation, not the implementation: `invoice.expired_yesterday`, not `invoice_test_4`. It is stable across builds, because it is written by the dataset build rather than assigned by the store. And it is unique, because resolution has to be deterministic. ## Why a stored identifier cannot be the handle The obvious shortcut is to note the identifier of the planted row once and hard-code it. It breaks for three separate reasons, and all three are ordinary: - **A rebuild renumbers.** Load order, retries, parallel inserts and any sequence the store owns will hand out different numbers next time. - **Environments differ.** The same dataset definition loaded into two places produces two numberings, so a test pinned to one number passes in one place and fails in the other. - **Failures become mute.** Even when the number happens to still be right, `row 41827 was missing` tells a reader nothing about which situation is broken. The failure mode is worse than a simple break. A hard-coded identifier that goes stale usually does not disappear — it now points at *some other row*, one the generator happened to place there. The test then exercises an ordinary background record while claiming to exercise a boundary situation, and it passes. That is a silently disabled check, which is the most expensive kind. ## Where the label lives | Placement | How it works | What it costs | |---|---|---| | A dedicated column on the row | the build writes the label alongside the data | ships a field the product itself does not need | | A manifest beside the dataset | label to identifier, written by the build, read by tests | one lookup, and a file that must stay in step | | A reserved identifier range | planted rows occupy a fixed, agreed block | brittle if the store allocates numbers itself | All three work. The choice is mostly about whether the schema is allowed to carry a test-only field, and how much the team minds tests performing a lookup before they assert. What is not optional is that **the build writes it and the test reads it** — a label recorded only in a code comment, a wiki page or somebody's memory is not a label, because nothing verifies it. ## Rules that keep labels useful 1. **One label, one row.** If two planted rows share a label, resolution returns whichever came back first and the test passes or fails on ordering. Fail the dataset build on a duplicate; it is a one-line check and the alternative is an intermittent failure nobody can reproduce. 2. **Name the situation, not the sequence.** A label that says what is unusual about the row lets the next reader decide whether the row is still needed. 3. **Never hard-code an identifier in a test.** Make it a review rule. A numeric literal standing for a row is the defect that hides behind a green build. 4. **Keep the reason next to the label.** One line saying which branch the row exercises turns a mysterious row into a maintainable one. ## What a label does not do It is worth being precise about the limit. A label makes a planted row **addressable and diagnosable**. It does not make the row **present**: if the build changed order, if the load step ran before the reset, if a later cleanup filled the field in, the label simply resolves to nothing, or resolves to a row that is no longer unusual. Presence and rarity are properties of the dataset that have to be asserted separately, after every build. Labelling is the prerequisite for that check — you cannot verify a set of rows you have no way to name — but it is not the check itself.
- Should the label live on the row itself or in a manifest beside the dataset?Either works, and the choice is about whether the product schema may carry a test-only field. A dedicated column is simplest and makes the label visible to any query, but it ships a field the product does not need. A manifest mapping label to identifier keeps the schema clean at the cost of a lookup and a file that must stay in step. What matters is that the build writes it and the tests read it.
- What breaks if two planted rows accidentally share a label?Resolution stops being deterministic. A test asks for the label and gets whichever row the lookup returned first, so it passes or fails on ordering rather than on the product. Make labels unique and fail the dataset build on a duplicate — it is a one-line check, and the alternative is an intermittent failure that nobody can reproduce on demand.
saying these in an interview costs you the question
- Referring to a planted row by a stored identifier
- Assuming identifiers stay the same after a rebuild
- Labelling only the rows that have already failed
- Keeping the label in a comment rather than in data
- Letting two planted rows share the same label