After a test dataset is rebuilt, how do you prove the planted rare rows are still there and still exceptional?
answer
- Assume nothing after a rebuild
- Two failures: absent and diluted
- Resolve every label after building
- Fail the build, not the suite
basics
~20 sAssert it in the build. A check run after every dataset build resolves each planted label, confirms the row still has the characteristic it was planted for, and fails the build when one is missing or has quietly become ordinary.
solid answer
~50 sTwo failures hide here and only one is loud. The loud one is **absence**: a build step moved, the plant step ran before the reset, and the row is simply gone. Resolving every planted label after the build catches that. The quiet one is **dilution**: the row is present but no longer exceptional, because the bulk volume grew and now holds forty other records with an empty surname, or because a later step filled the field in. Write the guard to assert both — each label resolves to exactly one row, that row still satisfies the condition it was planted for, and the number of rows across the whole dataset matching that condition is still what the tests assume. Fail the dataset build rather than the suite; a broken dataset should never reach a test run.
code
pseudocode · 15 linesafter dataset_build:
for entry in planted_manifest:
rows = dataset.matching(entry.condition)
require(dataset.by_label(entry.label) exists,
"planted row missing after rebuild: " + entry.label)
require(evaluate(entry.condition, dataset.by_label(entry.label)),
"planted row no longer satisfies its condition: " + entry.label)
require(rows.count == entry.expected_count,
"planted row diluted: " + entry.label +
" expected " + entry.expected_count + " got " + rows.count)
# a dataset failing this is never published to a test rungo deeper
Be ready to say that a rebuilt dataset is not automatically identical to the last one, and that a test can fail because a row it depends on never arrived rather than because the product is wrong.
Explain the two ways a planted row goes wrong after a rebuild — it is missing, or it is no longer unusual — and that only a check written against the built dataset catches the second one.
Show the guard: resolve every label, re-evaluate the condition the row was planted for, compare the population count, and fail the dataset build rather than letting a suite discover it one confusing failure at a time.
Own where this sits in the pipeline and what it blocks. A dataset failing its own manifest should not be publishable, and teams should not be able to opt out of the check row by row.
A planted row is an assumption, and assumptions that nothing checks quietly stop being true. The dataset gets rebuilt on every pipeline run, on every developer machine, and after every change to the generation rules — and each rebuild is a chance for a planted row to go wrong in one of two ways. ## The two failure modes | Failure | What happened | How it shows up | |---|---|---| | **Absence** | the plant step ran before the reset, a step was reordered, a load partially failed, someone edited the plant list | a loud test failure, or an empty result nobody notices | | **Dilution** | bulk volume grew and now contains many rows with the same characteristic, or a later step populated the field | tests that count or filter fail, and are usually "fixed" by loosening them | Absence is the one people design for. Dilution is the one that actually erodes a dataset, because its symptom looks like a flaky test rather than a broken dataset. A test asserting "exactly one customer has no recorded surname" is really an assertion about the shape of the whole dataset. When the generator's rules change and forty more such customers appear, the test fails for a reason that has nothing to do with the product — and the usual reaction under deadline is to relax the assertion to "at least one", which makes it pass forever and prove nothing. There is a third, quieter variant worth naming: the row is present and still unique, but a later step **changed it**. A cleanup normalises empty fields, a fixture patch sets a default, and the row that was planted to have nothing in a field now has something. It resolves, it is unique, and it no longer tests what it was planted for. ## Writing the guard The check is short and it belongs to the dataset, not to any test: 1. **Resolve every label.** For each entry in the planted manifest, look the label up. Zero rows or more than one row fails immediately, with the label in the message. 2. **Re-assert the characteristic.** The manifest records why the row was planted as a condition — field is empty, amount equals the agreed maximum, expiry date is in the past. Evaluate that condition against the row that came back. This is what catches a row that was silently normalised. 3. **Count the population.** Evaluate the same condition across the whole dataset and compare against the expected count. This is what catches dilution, and it is the step most teams leave out. 4. **Fail the build.** Not a warning, not a report. A dataset that does not satisfy its own manifest is not publishable. ## Why it fails in the build rather than the suite Putting the check in the test suite looks cheaper and is worse in four ways. The same broken dataset is reported once per affected test, so the cause is buried under a wall of unrelated failures. Tests that ran before the affected one already used the bad dataset and their results are now meaningless. The failure appears to belong to whichever feature the test covered, so it is triaged by the wrong person. And nothing stops the dataset being reused by the next run. Failing in the build gives one message, at one place, naming one label, pointing at the step that broke. That is the whole argument, and it is the same argument for validating any artefact where it is produced rather than where it is consumed. ## What to do when the guard trips The interesting case is when the guard is *right* and the dataset changed for a legitimate reason — the generation rules were updated and a shape that used to be rare is now common. There are two honest options and one bad one: - **Restore the rarity.** Adjust the generation rules so the shape returns to the far tail, keeping the planted row exceptional. Right when tests genuinely depend on the row being unique, and when the new proportions were not deliberate. - **Update the expected count and the tests that depend on it.** Right when the new proportions reflect the live store better than the old ones. It is more work, and it is the honest answer when the generator got more accurate. - **Loosen the assertion** to "at least one" and move on. This is the bad one. It converts a precise check into a check that nothing can fail, and the next person has no way to tell that it was ever meant to be exact. ## Keeping it cheap The guard costs one pass over the dataset per condition, which is why it should read counts rather than rows and run against the built dataset once, not per test. Its real cost is keeping the manifest honest: every planted row must record the condition it exists for, in a form the guard can evaluate. That is a small discipline, and it is the same discipline that makes the planted set reviewable at all.
- Why check that a planted row is still rare, and not only that it is still present?Because a test asserting that exactly one record has an empty surname is asserting the shape of the whole dataset, not the existence of one row. If the generation rules change and forty more appear, the test fails for a reason unrelated to the product, and the usual reaction is to loosen it to at least one — which makes it pass forever and prove nothing.
- Should this guard run in the dataset build or in the test suite?In the build, because that is where the fault is. A guard inside the suite reports one broken dataset once per affected test, buries the cause under unrelated failures, and lets everything that ran earlier use the bad data anyway. Failing the build stops the dataset being published and points at the single step that broke it.
saying these in an interview costs you the question
- Assuming a rebuild reproduces the previous dataset exactly
- Checking that a planted row exists but never that it is rare
- Loosening an exact assertion when the dataset changes
- Running the check only after a test has already failed
- Treating a missing planted row as a flaky test