Why is restoring a full production copy usually a poor source of test data?
answer
- Everything you copy, you must protect
- Restore time is paid on every refresh
- Asserting on rows a real user owns
- Production holds what happened, not what you need
- Subset, transform, verify, rebuild
basics
~20 sA whole-database copy drags every record's privacy obligation into a weaker environment, grows until the restore is the bottleneck, and gives an unstable target: rows a case asserts on change between refreshes. A masked subset costs far less.
solid answer
~50 sFour reasons, and they compound. **Privacy**: the whole personal-data estate now lives somewhere with weaker access control, and every field must be transformed correctly forever. **Volume**: restore time, storage and the cost of every rebuild scale with the entire dataset, while the suite touches a tiny fraction of it — teams end up refreshing rarely, which defeats the point of using real data. **An unstable target**: a case that asserts on a record it did not create is asserting on something a real user can edit or delete, so the suite fails for reasons unrelated to the code. **Missing cases**: the copy contains what happened, not what you need — the expired, the malformed, the boundary record often is not there, and you must construct it anyway. The usual answer is a referentially consistent subset, transformed, pinned and rebuildable, with each case still creating the records it asserts on.
code
pseudocode · 13 linessubset = traverseFrom(root: "school", ids: [3], follow: FOREIGN_KEYS, depth: ALL)
subset = includeWhole(subset, tables: ["period", "room_type", "term"])
subset = maskPersonalFields(subset, strategy: STABLE_SURROGATE)
checks = [
allForeignKeysResolve(subset),
noFieldMatchesRealContactShape(subset),
rowCountsWithinExpectedRange(subset)
]
if any(check.failed for check in checks):
discard(subset)
fail("extract not publishable")
publish(subset, version: buildId, expiresInDays: 14)go deeper
Be able to give two concrete reasons a whole-database copy is not the default: it brings real personal data somewhere less protected, and its records change under the suite so tests fail for reasons unrelated to the code.
Explain what replaces it end to end — a referentially consistent slice around a chosen root, personal fields transformed on the way out, and checks run against the produced data before anything uses it — and name what typically breaks in the slice.
Demonstrate judgement about the oracle: describe why cases must create the records they assert on even when a realistic background exists, and how you would diagnose a suite whose failures track the refresh schedule rather than the commits.
Frame it as a standing cost. Decide what data may leave production at all, what the extract pipeline must prove before publishing, how long a copy may live, and what the team gives up if it treats a full copy as the foundation instead of an experiment.
"Just restore production" is the most natural-sounding test-data plan there is, and it degrades in a predictable order. Being able to walk an interviewer through that order — and through what replaces it — is the senior part of this topic. ### What goes wrong, in the order it usually bites **1. Privacy scope.** A full copy imports the complete personal-data estate into an environment with more accounts, more copies and fewer controls. Every field with a person behind it must be transformed correctly on every rebuild, forever, and the failure mode is silent: an extract that is *mostly* transformed looks exactly like one that is fully transformed. **2. Volume you do not use.** A school timetable planner run by an 11-person team holds 4,213 students, 187 teachers, 1,046 courses and 61,214 timetable slots — a 6.8 GB dump that takes about 23 minutes to restore. The suite's assertions touch perhaps a few hundred rows. Every one of those 23 minutes is paid on every refresh, every environment and every rebuild after a corrupting run, and the predictable consequence is that the team stops refreshing. Data that is nine months stale is no longer 'realistic', it is merely large. **3. The oracle stops being stable.** An **oracle** is whatever decides the observed behaviour is right. If a case asserts that a particular student appears in a particular timetable, the oracle now depends on a record a real user owns. Someone withdraws that student in production; the next refresh, the case fails, and an engineer spends a morning on a defect that does not exist. Suites built on a full copy accumulate these until people stop believing red results. **4. The cases you need are absent.** Production holds what happened. The boundary case — the enrolment that expires today, the teacher with no assigned room, the course at exactly its capacity limit — is often not present, or is present once and disappears next month. You end up constructing those records anyway, which means the full copy bought volume and cost you a stable target. **5. It hides nothing about correctness.** A full copy makes a suite feel thorough while the assertions still only cover the paths someone wrote a case for. Volume is not coverage. ### What replaces it The standard answer is a pipeline, not an artefact: **Subset.** Choose a root entity and take a slice around it, following relationships outward: one school, its 187 teachers, its enrolled students, the courses those enrolments point at, plus every reference and lookup table copied whole because they are small and everything depends on them. The output must be **referentially consistent** — no row pointing at a parent the subset omitted. The classic failures are dangling references from rows the traversal reached but whose parents it did not, cyclic relationships that need a two-pass insert, and rows whose validity depends on a sibling table nobody thought to include. **Transform.** Personal fields removed or replaced by consistent surrogates on the way out, never after landing. **Verify.** Checks against the produced subset before anything is allowed to use it: every foreign key resolves, no field still matches the shape of a real contact value, per-entity aggregate counts match expectations. This matters because these pipelines fail partway. On one such run, the transform rewrote the student table, then failed on the guardian table; the step rolled back only its own partial work and exited, leaving 3,912 guardian rows with real contact values and references that no longer matched the rewritten student identifiers. The exit status said 'failed'; the extract was loaded anyway, because nothing was gating on the data itself. **Pin and rebuild.** The subset is a build output, reproducible from the pipeline, versioned and disposable — never a hand-tended artefact that accumulates manual fixes nobody can reproduce. **Keep assertions off it.** The subset is background, giving realistic shapes, volume and neighbours. Each case still creates the specific records it asserts on. That single discipline removes the unstable-oracle problem entirely, and it is what lets the subset be resized or regenerated without touching a test. ### When a full copy is genuinely right Be fair to it, because interviewers listen for nuance. Rehearsing a migration on true volume; reproducing a defect that only appears at production scale or with production skew; capacity work where the whole distribution matters. Those are **short-lived, purpose-built, tightly controlled** copies with an owner and an expiry — not the standing source a regression suite runs against every day. The distinction is between a copy used as an experiment and a copy used as a foundation.
- Your subset loads but a handful of rows point at parents that were not copied. How do you decide between widening the traversal and dropping those rows?Ask what the dangling rows represent. If they are a real relationship the suite exercises, widen the traversal to include the parent table — usually it is a lookup or reference table small enough to copy whole. If they are historical leftovers pointing at entities outside the chosen root, drop them, because keeping them means loading data with references the application will never resolve either. What you must not do is load them and let the constraint failures be discovered by whichever test happens to touch them first.
- The team wants a nightly refresh so the data stays current. What do you push back on?Currency is rarely the property that matters, and a nightly rebuild makes the target move under the suite every day. Ask which case needs data newer than the last release; usually none does. A subset rebuilt on demand and pinned for a period is better: reproducible, cheap to recreate after a corrupting run, and stable enough that a red result means the code changed. If a specific case genuinely needs fresh shapes, give that case a fresh slice rather than moving the whole environment.
- How would you size the subset?By what the suite actually needs, measured rather than guessed. Start from the entities the cases touch, keep the reference tables whole, and add volume only where a case asserts on behaviour that depends on it — paging, a time budget, a report aggregate. Then check the two numbers that matter: how long a rebuild takes, and how long the suite takes against it. If either has grown without a case demanding it, the extra rows are cost with no evidence attached.
saying these in an interview costs you the question
- Says a full copy is fine once names are replaced
- Assumes realistic volume implies better coverage
- Asserts on records the test did not create
- Treats the extract as a hand-tended artefact, not a build output
- Ignores that copied data must be re-transformed on every refresh
- Cannot name a legitimate use for a full-volume copy