skip to content

What does a checked-in fixture file trade away versus building test data in the test?

level: middleimportance: nice to knowfreq 32%

answer

  1. Where the data lives, not where it came from
  2. Bulk and awkward payloads favour a file
  3. Nobody owns a shared file's rows
  4. Asserted-on data belongs to the case
  5. Treat the file as regenerable output

basics

~20 s

A checked-in file of records is compact and reusable, but it separates data from the case that needs it, couples every case that loads it, and rots as the schema moves. Data built in the case stays owned and visible.

solid answer

~50 s

A shared fixture file wins on bulk and on realism-per-line: a large, awkward payload is far more readable as a file than as construction code. It loses on ownership. Nothing records which case needs which row, so no row can safely be deleted, edits made for one case silently change what another asserts on, and the file drifts from the schema until half of it exists only because removing it breaks something. Building the data inside the case inverts all of that: the reader sees exactly what state produced the result, the case owns its records, and a change touches one place. The practical rule is that data a case asserts on is constructed by that case, while files are reserved for bulk background, large realistic payloads and binary assets — and even then the file is treated as an input to be regenerated, not a hand-tended artefact.

code

pseudocode · 12 lines
pseudocode
// loaded once: reference data nothing asserts on
loadBackground("reference/periods", "reference/room_types", "reference/terms")

test "two lessons in one room at one period clash":
    room    = createRoom(name: "R-114")
    lessonA = createLesson(room: room, day: TUE, period: 4)
    lessonB = createLesson(room: room, day: TUE, period: 4)

    clashes = detectClashes([lessonA, lessonB])

    assert clashes.count == 1
    assert clashes[0].room == room

go deeper

for a junior

Know that data can either sit in a file loaded before the run or be created by the case itself, and that the second makes a failing case much easier to read because the input is right next to the assertion.

for a middle

Explain the coupling a shared file creates and how rot sets in: no row has an owner, edits for one case change another's starting state, and the file drifts from the schema until nobody dares delete anything.

for a senior

Show the split you would enforce — anything asserted on is built by the case, background may be a file — and describe how you would unwind an existing shared file without breaking a suite the team depends on daily.

for a principal

Treat fixture files as build outputs with a regeneration path and a review process, and set the rule that keeps a suite maintainable across a growing team, so ownership of test data does not quietly become nobody's job.

This is a shape question rather than a source question: wherever the records came from — invented, generated or extracted — they can either sit in a file that is loaded before the run, or be constructed by the case that needs them. Both are used, and interviewers like the question because the answer reveals whether someone has maintained a suite for more than a few months. ### What a file buys **Bulk without noise.** Fifty records expressed as construction calls is a wall of code; as a data file it is a table you can scan. **Awkward content stated exactly.** A deeply nested payload, a document with a real-world encoding quirk, an image or an archive — these are far better as literal artefacts than as code that assembles them. **Shared background.** Reference data that everything needs and nothing asserts on — the period table, the room types, the term calendar of a school timetable planner — belongs somewhere central. Constructing it in every case is pure duplication. **Reuse across levels.** The same recorded payload can feed a unit-level case, a contract check and a manual reproduction. ### What a file costs **Ownership disappears.** Once a file is loaded before every case, nothing records which row exists for which reason. That is the mechanism behind **fixture rot**: rows accumulate, none can be safely deleted, and after a year a meaningful share of the file exists only because removing it turns something red for reasons nobody can explain. **Cases couple through it.** Adjusting a record so one case passes changes the state every other case starts from. On an 11-person team this is a steady tax: someone edits a row, an unrelated suite goes red, and the two changes are not obviously connected. **The dependency is invisible.** Reading a case that asserts a clash on a particular timetable slot tells you nothing about why that clash exists. The reason is in a file, in a row nobody points at. Diagnosis time goes up sharply. **Drift.** The file is written against the schema of the day it was created. A new required field, a renamed relationship, a changed default — each divergence gets patched into the file rather than fixed at the source, and the file gradually describes a product that no longer exists. **A false sense of realism.** A file that was extracted from real data once is realistic exactly once. Two years later it is a historical document, and cases written against it are asserting on shapes the product no longer produces. ### Building in the case The alternative states the records inside the case, right above the assertion. The reader sees the whole story in one place; the case owns its data and can change it without consulting anyone; a failure points at input that is visible in the same screen. The costs are real too — construction code repeats, a case with a long prerequisite chain gets noisy, and bulk becomes impractical past a few dozen records. ### The rule that resolves it Split by whether the case *asserts on* the data. - **Asserted on → constructed by the case.** If a value appears in an assertion, the case that asserts it creates it. No exceptions; this is what keeps the suite readable and keeps cases from coupling. - **Background → a file is fine.** Reference and lookup data, bulk volume, large literal payloads. Nothing asserts on it, so nobody edits it for one case's benefit. Then treat the file as a **build output rather than an heirloom**. It should be regenerable from a described source — a pipeline, a recorded interaction, a generator run — so refreshing it is a command rather than an archaeology project. Keep it in the repository beside the suite so a change to it is reviewed like code. And give rot a way to surface: keep files small enough to read, split them so a case's needs are visible rather than pooled, and periodically delete rows to see what actually fails. A row nobody can justify is a row that will be maintained forever. ### The failure this prevents The worst version is the file that is loaded before everything, edited by everyone and understood by nobody — the one where a partially applied change leaves half its rows describing the new schema and half the old, and every case that touches it fails for a different reason. Nobody set out to build that. It is what a shared file becomes when no case owns its rows.

  • How would you tell whether a shared fixture file has rotted?
    Look for rows nobody can justify. Delete a row and run the suite: if something fails, the failure tells you which case depended on it and that dependency belongs in the case instead; if nothing fails, the row was dead weight. Other symptoms are edits to the file made to fix one case, patches applied to satisfy a schema change rather than regenerating from the source, and cases whose assertions name identifiers that appear nowhere in the test code.
  • A suite has one large file loaded before every case. What is the first change you would make?
    Stop cases asserting on it. Leave the file as background and move every value that appears in an assertion into the case that asserts it, one case at a time. That breaks the coupling first, which is what makes every other change safe: once no assertion depends on the file, rows can be split out, deleted or regenerated without a suite-wide risk. Splitting the file before decoupling the assertions just redistributes the same problem.
  • When is a large literal file clearly the right answer?
    When the artefact itself is the input and constructing it in code would obscure it: a deeply nested real-world payload, a document with an encoding or formatting quirk that matters, a binary asset, or a recorded response whose exact bytes are the point of the case. Keep it beside the case that uses it, name it after that case, and record where it came from so it can be regenerated rather than hand-edited when the format moves.

A shared fixture file is a communal kitchen cupboard: convenient at first, and eventually full of jars nobody will throw out because someone might be using them.

saying these in an interview costs you the question

  • Says files are always cleaner because they avoid duplication
  • Keeps assertion targets inside a shared background file
  • Patches a fixture file by hand after a schema change
  • Cannot explain why rows in an old fixture file exist
  • Treats an extracted file as realistic years after extraction
  • Splits a shared file without first decoupling assertions

context