skip to content

Test Data & Seeding Strategy

You will learn where an end-to-end test's data comes from and how it gets cleaned up, which is what makes a suite safe to run in parallel against a shared environment. Interviewers ask because most candidates have only ever tested against a hand-seeded dev database.

on this pageshow

questions

5

An end-to-end suite registers with the hard-coded address [email protected] and asserts on a fixture order numbered 1001. What breaks once that suite runs on several parallel workers, or simply runs twice in a row, and what changes if each test builds its data from a factory instead?

level: middleimportance: must knowfreq 65%

answer

  1. hard-coded values behave like a lock
  2. two runs, one identifier
  3. overrides say what actually matters
  4. prefix tells you who made this row
  5. random is not the same as traceable

basics

~20 s

Fixed identifiers collide: a second worker or a second run is rejected because the address is already taken, or two tests mutate the same order and see each other's changes. A factory mints fresh unique values per test, so each test owns data nothing else touches.

solid answer

~50 s

Any value the system treats as unique — an email, a slug, an external reference — becomes a shared lock when it is hard-coded. Two workers registering `[email protected]` at once means one of them fails on a duplicate, and a rerun on a non-reset environment fails the same way. A fixed record number is worse in a quieter way: two tests edit order 1001 concurrently and each sees the other's writes, which shows up as a flake that only appears under load. A factory is a function that returns a valid object with unique values baked into the fields that must be unique, plus overrides for the one field the test actually cares about — `makeUser({ plan: 'pro' })`. That also documents intent: everything not overridden is irrelevant to this test. Namespace the unique parts with a run identifier so leftovers are traceable, and log generated values so a failure is reproducible.

code

javascript · 12 lines
javascript
function makeUser(overrides = {}) {
  const tag = `${Date.now()}-${Math.random().toString(36).slice(2, 8)}`;
  return {
    email: `e2e-${tag}@example.test`,
    name: `E2E User ${tag}`,
    plan: 'free',
    ...overrides,
  };
}

console.log(makeUser());
console.log(makeUser({ plan: 'pro' }));

go deeper

for a junior

Know that hard-coded emails and record numbers collide when tests run at the same time or twice in a row, and that each test should create its own data with fresh unique values.

for a middle

Explain both failure modes — outright duplicate rejection and silent cross-talk between tests sharing a record — and describe a factory: complete valid defaults, uniqueness on constrained fields, overrides that document what the test depends on.

for a senior

Show the operational side: namespacing generated values with a run or worker identifier so records are traceable and reapable, logging generated inputs so a CI-only failure can be reproduced, and keeping assertions scoped to the data the test owns.

for a principal

Own the convention across teams — a shared factory layer that tracks the domain model, an agreed namespacing scheme readable by both cleanup jobs and on-call engineers, and a policy on what may remain fixed shared reference data.

## The two ways fixed data fails **Collision.** The system enforces uniqueness on some fields — email address, username, slug, order reference, external id. A hard-coded value turns that constraint into a mutual-exclusion lock across every test that uses it. Two workers registering the same address at the same moment means one gets "already taken". Running the same suite twice against an environment that is not wiped in between fails on the first run's leftovers. The suite passes only in the single, sequential, freshly-reset case — which is exactly the case that disappears the moment the suite becomes useful. **Cross-talk.** This one is quieter and worse. Two tests share order 1001: one cancels it, the other asserts it is shippable. Neither is wrong on its own; together they interleave and one fails, seemingly at random, only when the machine is busy. Because nothing in the failing test refers to the other test, the natural diagnosis is "timing", and someone adds a wait that fixes nothing. ## What a factory is A factory is a plain function that returns a complete, valid object with fresh values every call, and accepts overrides: ```javascript function makeUser(overrides = {}) { const tag = `${Date.now()}-${Math.random().toString(36).slice(2, 8)}`; return { email: `e2e-${tag}@example.test`, name: `E2E User ${tag}`, plan: 'free', ...overrides }; } ``` Three properties matter: 1. **Completeness.** The default object is valid on its own, so adding a required field to the model is a one-line change in the factory rather than an edit to forty tests. 2. **Uniqueness where uniqueness is enforced.** Only the constrained fields need to vary. A display name may repeat harmlessly; an email may not. 3. **Overrides as documentation.** `makeUser({ plan: 'pro' })` says loudly that the plan is the only thing this scenario depends on. A test that instead loads a fixture file gives the reader no way to tell which of its thirty fields matters. ## Namespacing, not just randomness Random values stop collisions but leave you with an environment full of anonymous junk. Prefix the unique part with something meaningful: the suite name, the CI run id, the worker index. `[email protected]` tells you at a glance who created a record, lets a cleanup job select everything from a run, and lets an on-call engineer distinguish test traffic from a real customer in a shared environment. Many mail systems also accept a plus-suffix (`[email protected]`) so each test can own a distinct address that still routes to one mailbox. ## Where fixtures still belong Factories are for the data a test **mutates or owns**. Reference data that everything reads and nothing changes — a country list, product catalogue entries, a tax table, the standard set of role definitions — is legitimately fixed, seeded once with the environment. The dividing line is writes: shared read-only data is a resource, shared writable data is a race. ## The pitfalls of over-randomising - **Irreproducibility.** If a run fails on a name containing an emoji and nothing recorded that name, you cannot reproduce it. Log what you generated, or derive values from a seed that the report prints. - **Accidental invalid input.** A generator that occasionally produces a 300-character name or a string the app legitimately rejects makes the suite fail for reasons unrelated to the feature. Generate values inside the valid domain; if you want to test the edges, write an explicit test for the edge. - **Assertions on generated values.** Asserting that the page shows exactly the string you generated is fine and correct. Asserting on a value you did *not* generate — "the list contains three rows" — reintroduces the shared-world assumption the factory just removed. Assert on your own data, scoped to it. ## The shape of a well-behaved test Build the data with a factory, create it through the app's API, remember the identifiers you got back, act, then assert only on those identifiers. Everything the test touches is then something it created, which is what makes running eight copies at once safe.

  • Which fields in a generated object actually need to be unique?
    Only the ones the system constrains or that a test looks up by: email, username, slug, external reference, invoice number. Display names, addresses and descriptions can repeat harmlessly. Varying everything adds noise and makes failures harder to read, so vary deliberately rather than by default.
  • Is there anything a shared fixed fixture is still right for?
    Yes — reference data that every test reads and none of them writes: a country list, catalogue entries, tax rates, role definitions. Seed it once with the environment. The dividing line is mutation: shared read-only data is a resource, shared writable data is a race waiting to happen.
  • A test fails only in CI, on a randomly generated name. How do you make that reproducible?
    Record the generated values in the test output or attach them to the report, so the failing input is visible in the run that failed. Deriving values from a printed seed works too. Without that, random data buys isolation at the price of never being able to re-run the exact failure.
  • Why is asserting "the orders list shows three rows" a problem even with a factory?
    Because the count depends on the whole environment, not on your data. Another worker's order or last week's leftovers change it without any bug existing. Assert on the specific record the factory created — that it appears, with the values you generated — and scope the query to it.

saying these in an interview costs you the question

  • Says collisions are a flakiness problem and adds a wait
  • Puts one shared QA account in every test and reuses its records
  • Randomises every field, then cannot reproduce the failing input
  • Assumes the environment is wiped before every run
  • Asserts on total row counts in a shared environment

context

open as a page

An end-to-end suite creates real records through the application's API. A test crashes halfway through and its cleanup step never runs; over weeks the shared test environment fills with orphaned data and the suite starts failing. How would you design the data lifecycle so that any rerun is safe?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Do not rely on cleanup running. Make tests assert only on data they created, tag every record with a run identifier so a scheduled reaper can delete leftovers, register teardown at creation time so it deletes in reverse, and treat resetting the environment as the real backstop.

open as a page

An end-to-end test for an order-detail page needs the signed-in account to already have one completed order. Why do teams create that prerequisite data with an API call or a database setup step instead of driving the application's own UI to create it first, and when is UI-driven setup still the right choice?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Creating prerequisite data through the product's API or a database setup step is faster and far less brittle, because the test then depends on no UI it is not asserting about. Drive the UI for setup only when that creation flow is itself the behaviour under test.

open as a page

An end-to-end suite runs eight workers in parallel against one shared staging environment. Giving each test its own freshly created records fixed most of the interference, but a handful of tests still fail only when the whole suite runs together. What kinds of state cannot be made unique per test, and what do you do about them?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Singleton state resists per-test uniqueness: application-wide feature flags and settings, limited inventory, rate limits, a shared mailbox, background jobs, caches and the clock. Partition what you can into a per-worker tenant or account, and run the few tests that must mutate genuinely global state serially.

open as a page

A team proposes running the end-to-end suite against a nightly anonymized copy of the production database instead of against data the tests create for themselves. How would you evaluate that proposal, and what would you put in place either way?

level: principalimportance: should knowfreq 40%

basics

~20 s

Judge it on what it adds versus what it destabilises: real volume and shapes catch bugs synthetic data never will, but a refreshing snapshot makes assertions on specific records unrepeatable and carries personal data into a lower-trust environment. The usual answer is a hybrid.

open as a page