How do hand-built fixtures, generated records and production extracts differ as test-data sources?
answer
- Three places test data can come from
- Written, invented, or copied
- Readable versus realistic versus safe
- Copied data must be cut down and transformed
- Generators only emit what the model allows
basics
~20 sHand-built fixtures state only the few records a case needs, so they read clearly and carry no privacy risk. Generated records buy volume and variety cheaply. Production-derived extracts show real shapes and skew, but must be cut down and masked first.
solid answer
~50 sThree sources, three trade-offs. A **hand-built fixture** is written by the test itself: a handful of records holding only what the case depends on. It is readable, repeatable and free of personal data, but it encodes what the team *believes* real data looks like. **Generated records** come from rules or a schema description; they are cheap at volume and useful for boundary sweeps and size-sensitive cases, but a generator only emits data the model already allows, so genuinely strange legacy rows never appear. A **production-derived extract** is copied out of the live store and transformed before use: cut down to a workable size, with personal fields removed or replaced by stable surrogates. It is the only source that shows real skew, and the only one that drags privacy obligations into a weaker environment. Most suites default to hand-built, generate for volume, and reserve extracts for the few cases that need real shapes.
code
pseudocode · 14 lines// hand-built: only what the assertion needs
teacherA = createTeacher(name: "T-A")
slot1 = createSlot(teacher: teacherA, day: MON, period: 3)
slot2 = createSlot(teacher: teacherA, day: MON, period: 3)
assert detectClashes([slot1, slot2]).count == 1
// generated: the count and the shape are the point
slots = generateSlots(count: 61214, teachers: 187, days: 5, periods: 8)
assert rebuildTimetable(slots).durationMs < 4500
// extracted: real shapes, transformed before it lands
extract = subsetFromProduction(rootTable: "school", rootIds: [3], followForeignKeys: true)
extract = maskPersonalFields(extract, strategy: STABLE_SURROGATE)
load(extract)go deeper
Be ready to name the three sources and one honest advantage and disadvantage of each, without claiming any is universally best. Knowing that copied real records must be cut down and transformed before they enter a test environment is the part interviewers listen for.
Explain the mechanics behind each trade-off: why a generator structurally cannot produce a record that violates today's invariants, why a hand-built fixture makes a failure readable, and what a transform on an extract has to preserve for the data to stay useful.
Show the combination in practice: a masked background plus per-case records, with no assertion naming data the case did not create. Be able to say which defects on a real system each source actually caught for you, and which it never would have.
Own the policy. Decide when the organisation accepts copied personal data at all, what the transform pipeline must prove before an extract is usable, and how much realism the suite is buying for that risk. State it as a standing rule so eleven engineers do not each answer it differently.
A **fixture**, in this sense, is the data that must already exist before a test runs — the records the behaviour under test reads, updates or reports on. Deciding where that data comes from is a design choice, and interviewers ask about it early because the three available answers trade the same three things against each other: readability, realism and privacy. ### Hand-built minimal fixtures The test constructs the records it needs, and nothing else. Two teachers, one room, three timetable slots. The properties that matter are that the data is **visible** — a reader of the test can see exactly what state produced the result — and that it is **owned** by the case, so nobody else's edit changes it. A failure is diagnosable because the input is three lines above the assertion rather than in a file someone loaded years ago. The cost is that hand-built data is an act of imagination. It contains what the author thought was possible. A field that is optional in the schema but empty in ninety per cent of real rows will usually be populated in a hand-built fixture, and the code path for the empty case never runs in the suite. ### Generated records A generator produces records from rules: a name pattern, a value range, a distribution, a count. This is how you get a thousand of something without typing a thousand of something, and it is the right tool for size-sensitive behaviour (a report that must page, a scheduler that must stay inside a time budget), for sweeping a parameter space, and for filling in the many fields a case never asserts on. The limitation is structural rather than accidental: a generator emits data the model permits. It writes the schema as the team currently understands it. The row inserted four years ago by a since-deleted import path, whose end date precedes its start date, is not in the generator's grammar. How much realism a generator can be pushed to reach is genuinely contested — teams that invest in distribution-shaped generation report catching far more than teams that do not — but the class of defect it structurally cannot produce is the record that violates today's invariants and exists anyway. ### Production-derived extracts Here records are copied out of the live store and transformed before they land anywhere a test can reach: **subset** (take a referentially consistent slice rather than everything), **masked** (personal values removed or replaced), and often **pinned** to a fixed point so the data does not shift under the suite. This is the only source that carries real skew — the one account with four hundred children where every other has two, the encoding oddities, the historical rows. The cost is that the moment real personal data lands in a test environment, it brings its obligations with it into a place with weaker access control, more copies, and more people. That is why extracts are transformed rather than restored, and why a masking step that only half-ran must be treated as no masking at all. ### A worked comparison A school timetable planner, maintained by an 11-person team. Production holds 4,213 students, 187 teachers, 1,046 courses and 61,214 timetable slots. - *Clash detection*: two teachers, three slots, one deliberate overlap. Hand-built. The whole fixture fits in one screen and the assertion names the clash. - *Whole-term rebuild inside a time budget*: 60,000 generated slots across generated courses. Nobody reads them; the shape and the count are the point. - *A defect only real data reproduces*: one teacher is registered at two schools and appears twice in the staff list with the same identifier. No hand-built fixture invented that, and no generator was told it was possible. A masked extract of the staff and enrolment tables reproduces it on the first run. ### How to answer the choice Default to hand-built, because readable data is a permanent asset and invented data is a one-time cost. Generate when the case is about size, spread or a sweep rather than about a specific record. Reach for an extract when you need shapes you cannot invent — and then take the smallest referentially valid slice you can, transform the personal fields, and be able to rebuild the whole extract from the pipeline rather than treating it as a precious artefact. The three are not exclusive. A common shape is a masked extract as the background the environment starts from, with each case still creating the specific records it asserts on, so the assertions never depend on data the case did not put there.
- You need one case to run against real shapes and the rest of the suite to stay hand-built. How do you combine them without the hand-built cases becoming dependent on the extract?Treat the extract as background, never as an assertion target. The environment may start from a masked slice, but every case still creates the records it asserts on and asserts only on those. Nothing in the suite may name a record it did not create — no fixed identifier, no seeded row. Then the extract can be rebuilt, resized or replaced without touching a single test, and the one case that genuinely needs real shapes states that dependency explicitly rather than inheriting it.
- A generator produces a thousand valid records and the suite passes. Why might that still be weak evidence?Because the generator was written from the same understanding of the domain as the code under test, so it tends to produce exactly the inputs the code already handles. Everything is well-formed, distributions are flat rather than skewed, and nothing violates an invariant that real history violates. The suite proves the code handles the model the team believes in. Adding deliberately hostile values — extremes, empty optional fields, unusual encodings — buys back more than raising the record count.
- What would make you refuse a production-derived extract outright?When the transform cannot be verified. If nobody can demonstrate which fields were rewritten, or the masking step is a manual run rather than a repeatable pipeline, the extract has to be treated as live personal data and it does not belong in a test environment. The same applies when the fields are so entangled that masking destroys the very realism the extract was wanted for — at that point a generator shaped from aggregate statistics is the safer trade.
It is the difference between drawing a diagram of a building, rendering one from a floor-plan generator, and photographing the real building: each is clearer, cheaper or truer than the others, and only the photograph shows you the pipe someone added in 1994.
saying these in an interview costs you the question
- Says real data is always better because it is realistic
- Assumes generated records will surface legacy or malformed rows
- Treats a copy of live records as safe once it leaves production
- Thinks a bigger fixture is a better fixture
- Cannot name a case where hand-built data is the wrong choice
- Believes one source must be chosen for the whole suite