Why plant rows the current product can no longer create into a manufactured test dataset?
answer
- Today's code did not write everything
- Removed features leave rows behind
- Read paths still meet old shapes
- Profile the store, plant representatives
basics
~20 sA manufactured dataset contains only states today's code can produce, but the live store also holds residue from removed features, skipped migrations and manual corrections. Read paths still meet those rows, so a representative of each must be planted deliberately.
solid answer
~40 sGeneration rules are written from the product as it stands today, so a fresh dataset quietly claims that every stored row was written by current code. The live store is not like that. It holds records from features that were removed, columns a migration back-filled for some rows and not others, values a support engineer corrected by hand, and shapes that were legal under a validation rule since tightened. Read paths still meet those rows and often break on them: a display assuming a field the old writer never set, a total assuming a unit the product stopped supporting. Discover the shapes by profiling the live store for rows current write paths could not have produced, then plant one labelled representative of each so every build exercises the read path.
code
pseudocode · 13 lines# 1. profile the live store for shapes, never rows
for shape in shapes_of(live_store): # counts only, no values
if not producible_by(current_write_paths, shape):
residue.add(shape, count: shape.row_count)
# 2. plant one authored representative of each surviving shape
plant("customer.region_never_set", region: EMPTY) # writer removed
plant("account.backfill_skipped", tier: EMPTY) # migration missed these
plant("order.retired_unit", unit: "unit-withdrawn")
plant("invoice.hand_corrected", total: adjusted, correction_note: EMPTY)
# 3. keep the reason with the row
note("customer.region_never_set", "withdrawn regional flow, ~40k rows remain")go deeper
Be ready to say that a stored record can be older than the code reading it, and that a dataset built only from what today's product writes will not contain those older shapes at all.
Explain where such rows come from — removed features, migrations that skipped rows, manual corrections, rules since tightened — and that read paths still meet them, so a representative of each belongs in the manufactured dataset.
Show that you would profile the live store for shapes current write paths cannot produce, plant one labelled representative of each, and treat a failure on such a row as a defect rather than as bad test data.
Own the trade-off between carrying historical shapes in the dataset indefinitely and funding a cleanup of the live store, and decide when a planted shape is allowed to retire.
The most dangerous thing a manufactured dataset does is agree with the code that generated it. ## What a generated dataset quietly assumes Generation rules describe the fields the current writers set, the enumerated values current validation allows, and the combinations the current flows can reach. That makes the dataset an unexamined claim: **every row in the real store was written by the code that is running now.** A store that has been in service for any length of time is sedimentary instead. It holds rows written by three versions of the writer, rows a migration touched and rows the same migration skipped, rows a support engineer corrected by hand at two in the morning, and rows that were perfectly legal under a rule tightened two years ago. None of those can be produced by today's write paths, and all of them are read by today's read paths. ## Where the residue comes from | Origin | The shape it leaves behind | What still reads it | |---|---|---| | A feature that was removed | a field only that feature ever set, or never set | listings, exports, any read older than the removal | | A back-fill migration | rows the back-fill skipped, failed on, or never saw | anything assuming the column is now always populated | | A manual correction | a value no validation would accept, a state reached out of order | state machines, reconciliation, audit views | | A validation rule that was tightened | values legal under the old rule, rejected by the new one | validation on read, re-save flows, bulk edits | | A bulk import | lengths, encodings and precisions the interactive path never produces | display, search, comparison | | A partial cleanup | a child row whose parent record was removed | joins, cascade handling, totals | The common thread is that these are all **read-path defects**. Nothing fails while the row sits there. Something fails when the product is asked to display it, total it, export it, or migrate it again. And a read-path defect found in production is expensive precisely because the row cannot be re-created — you cannot reproduce the bug by using the product, which is why it is often filed as unreproducible and closed. ## Finding the shapes without copying the data 1. **Profile, do not extract.** Run shape queries over the live store that count rows by whether a field is populated, by enumerated value, by combination, and by length band. What comes back is a table of shapes and counts, holding no personal data. 2. **Compare against what today's writers can emit.** For each shape, ask whether any current write path could produce it. Where the answer is no, the shape is residue, and its count tells you how much of it is still out there. 3. **Author a representative.** Write one row that has that shape, with invented values. It is manufactured data describing a real shape, not a copy of a real row, so nothing leaves the production boundary. 4. **Record why it exists.** A one-line note beside the label — "written by the withdrawn regional pricing flow, roughly forty thousand rows still present" — is what lets a future reader decide whether the row still earns its place. ## Then decide what a failure on one means The point of carrying these rows is that a test will eventually fail on one, and there are exactly two honest responses. Choosing between them is the judgement this subject is really about. - **The shape still exists in the live store.** The failure is a genuine defect. Real users, or a nightly job, will hit it. Fix the code; the planted row stays. - **The shape no longer exists** because a cleanup ran or the last rows were finally migrated. The planted row now describes a past that is gone. Delete it and record the date and the reason. This is the only legitimate way such a row leaves the dataset. What must not happen is the third option people reach for under deadline: deleting the row because it is inconvenient, without asking which situation applies. That converts a known defect into an unknown one and silently restores the assumption the planted row existed to break. ## How much of it to carry Not every residue shape earns a permanent row. Rank them by how many rows of that shape survive and how many read paths touch them. A shape with four rows left that no current screen displays is a better candidate for cleaning out of the live store than for planting — a shape removed from production is a shape nobody has to carry in the dataset forever. A shape with hundreds of thousands of rows is not going anywhere, and belongs in the manufactured dataset until it does.
- How do you find these shapes without copying production data into a test environment?Profile rather than extract. Run queries over the live store that count rows by nullability, by enumerated value, by combination and by length band, and compare the result against what current write paths can emit. What comes back is a list of shapes with counts, not personal data. You then author a representative row from the shape description, so no real record ever crosses the production boundary.
- A planted legacy row makes a test fail. How do you decide whether to fix the code or delete the row?Ask whether the live store still holds rows of that shape. If it does, the failure is a defect waiting for a real user and the code must handle it. If the last such rows were cleaned up months ago, the planted row describes a past that no longer exists: delete it and record why. The dataset should track the store's real shapes, not accumulate history.
A generated dataset is a photograph of the product as it is today. The live store is a sediment layer holding fossils from every version that ever wrote to it.
saying these in an interview costs you the question
- Assuming every stored row was written by current code
- Deleting a planted legacy row whenever a test fails
- Believing a back-fill migration reached every affected row
- Copying production records instead of profiling their shapes
- Treating unreachable states as impossible rather than historical