A bulk generator filled a test dataset with realistic records: which rows will it essentially never produce, and why plant those by hand?
answer
- The distribution decides what appears
- Common middle full, tails empty
- Rare rows arrive by luck
- One named row per situation
basics
~20 sRealistic generation reproduces the common middle, not the rare tails. Absent optional fields, exact boundary amounts, right-to-left names and rows flagged deleted appear far too rarely to arrive by chance, so each must be planted deliberately rather than waited for.
solid answer
~50 sA generator that matches real proportions is *by construction* dominated by the common case. If one customer in fifty thousand has no recorded surname, a hundred-thousand-row dataset holds two of them on a good build and none on a bad one. The rows that break code are exactly the improbable ones: an optional field left empty, an amount sitting on the agreed maximum and one unit past it, a name written right to left, a value at the maximum stored length, a record already past its expiry date, a child row whose parent is gone, a row flagged deleted but still stored. Plant one of each explicitly, as named rows, and treat the generated records as background volume. Otherwise which exceptional rows exist changes from build to build and no test can depend on them.
code
pseudocode · 15 linesdataset_build:
background = generate(100000) # realistic volume, nothing more
planted = [
{ label: "customer.surname_absent", surname: EMPTY },
{ label: "order.amount_at_maximum", amount: agreed_maximum },
{ label: "order.amount_over_maximum", amount: agreed_maximum + 1 },
{ label: "customer.name_right_to_left", display_name: RTL_SAMPLE },
{ label: "address.line_at_max_length", line1: repeat("x", max_length) },
{ label: "invoice.expired_yesterday", expires_on: today - 1 },
{ label: "line.parent_removed", parent_id: id_no_longer_present },
{ label: "customer.flagged_deleted", deleted_flag: true }
]
load(background + planted)go deeper
Be ready to say why a hundred thousand realistic records still miss the rows that break code: rarity. Name two you would plant by hand, such as an empty optional field and a value sitting exactly on an agreed maximum.
Explain the mechanics. Generation follows proportions, so a shape occurring once in fifty thousand real rows can appear zero times in a build. Show that you would author a fixed row per situation rather than raise the generator's frequency.
Show judgement about which situations earn a permanent planted row in a shared dataset, what that adds to maintenance, and how you stop the planted set becoming a museum of everything anyone once found interesting.
Own the policy: who may add a planted row, what evidence justifies one, and how the set is reviewed so it tracks the branches the product actually has rather than growing without limit.
A manufactured dataset has two jobs at once, and a bulk generator only does one of them. ## Realism and rarity pull in opposite directions The first job is to be **plausible**: volumes, distributions and combinations close enough to the live store that a paging screen, a report and a query plan all behave roughly the way they will behave for real. The second job is to **contain the situations the code has branches for**. A generator fitted to real proportions serves the first well and the second badly, for a reason no amount of tuning removes: it emits rows in proportion to how often they occur. If one customer in fifty thousand has no recorded surname, a hundred thousand generated rows contain about two of them, and "about two" means some builds contain none. If one order in a million sits exactly on the agreed maximum amount, no realistic dataset of a size a test suite can load will reliably contain one. The rows that break code are, almost by definition, the rows a faithful distribution barely produces. So the dataset splits into two parts governed by different rules. **Background volume** comes from the generator and exists to make the shape realistic. **Planted rows** are authored one at a time, each because some branch of the product exists to handle it. ## What earns a planted row The catalogue is short, and it comes from reading the code rather than from imagination. Walk the read and write paths and ask what each one special-cases. | Situation to plant | Why generation misses it | What it breaks | |---|---|---| | An optional field left empty | optional fields are usually filled in | a display or export that assumes a value | | A value exactly on an agreed limit, and one past it | limits sit in the far tail | a comparison written with the wrong strictness | | A name written right to left | the fitted sample is regional | layout, truncation, ordering of names | | A field at its maximum stored length | lengths cluster well below the cap | overflow, silent truncation, broken alignment | | A record already past its expiry date | generated dates cluster near now | renewal logic, filters that assume validity | | A child row whose parent record is gone | generation builds parents first | joins, cascade handling, reported totals | | A row flagged deleted but still stored | deletion is rare and recent | listings that forget to filter, counts | | A state the running product can no longer create | today's writers cannot emit it | read paths older than the writers | Two of those deserve a note. Text written right to left is not an exotic curiosity — it exercises the entire display stack, and it is the single cheapest planted row most teams are missing. And a child row whose parent is gone is a real state in most stores that have ever had a partial cleanup, however firmly the schema claims otherwise. ## How to plant 1. **One row per situation, authored explicitly.** Not "make the generator emit more of these": a raised rate is still a random draw, and it distorts the very proportions the background volume exists to provide. 2. **Fixed and named.** The row is referenced by a name you chose, so a test asks for the situation rather than for a number somebody read off a screen. 3. **Loaded alongside the volume, not instead of it.** Planted rows sitting in a miniature dataset of their own only prove the code works on ten rows; the point is that they are present while a realistic amount of everything else is present too. 4. **Justified by a branch.** A planted row earns its place because some code path handles it. A row that is merely strange adds maintenance and proves nothing when it breaks. ## Why not just tune the generator The tempting shortcut is to raise the frequency of the rare shapes until they show up naturally. It fails on both jobs. The dataset stops being realistic, so anything that reads an aggregate, depends on selectivity or pages through results now sees a population that does not exist. And it still gives no guarantee about the *specific* row: raising a rate to one in a hundred means a build can still miss it, and whichever row does arrive carries whatever other values the draw produced, so a test cannot assert against it precisely. The opposite failure is a team that plants nothing and simply generates a very large dataset, on the theory that enough rows contain everything. Enough rows contain everything that is merely *uncommon*. They contain nothing the distribution assigns essentially zero probability — which is exactly where the impossible-looking states that do reach production live. ## The cost side Planted rows are permanent, and each one is something a future reader must understand. A set that only grows becomes a museum. Keep the justification next to the row — which branch it exercises — so the set can be pruned when that branch disappears, and review the planted set like code rather than letting it accumulate like sediment.
- A colleague suggests telling the generator to emit these rare shapes more often instead. What does that cost?Two things. Raising the frequency of a rare shape makes it common, so the dataset stops resembling the live store and anything reading proportions, selectivity or aggregates becomes wrong. And it still gives no guarantee of the specific row: a higher rate is still a random draw, so a build can miss it and the row that arrives carries unpredictable values. Frequency skew is not a substitute for a fixed, named row you can assert against.
- Which rows would you refuse to plant, even though they are genuinely rare?Rows no code path can encounter, and rows that exist only because they look interesting. A planted row earns its place by exercising a branch the product actually has: an unhandled empty field, a limit comparison, a length check, a display path for other scripts. Anything else is permanent maintenance that proves nothing when it breaks, and it makes the planted set harder to review.
A faithful sample of a city's traffic gives you thousands of cars and no ambulances. If you need to test how the junction handles an ambulance, you have to drive one there yourself.
saying these in an interview costs you the question
- Assuming a large enough dataset contains every edge row
- Leaving rare rows to the generator's randomness
- Raising a rare shape's frequency and calling it planted
- Blaming generator capability rather than its proportions
- Planting strange rows no code path actually handles